NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-01-09ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Show HN: EuConform – Offline-first EU AI Act compliance tool (open source)

EuConform is an open-source project developed to offer an offline-first solution for complying with the EU AI Act. Initiated as a personal endeavor, its core purpose is to transform the EU AI Act's intricate requirements into tangible, inspectable technical checks for AI systems. The tool emphasizes a local-first compliance methodology, ensuring heightened data privacy and operational autonomy. Its functionalities encompass comprehensive risk classification, aligning with Articles 5–15 of the Act, which includes identifying prohibited use cases. EuConform also integrates bias evaluation mechanisms, specifically leveraging the CrowS-Pairs method to assess potential biases in AI models. Additionally, it automates the generation of Annex IV-oriented PDF reports, streamlining documentation for regulatory adherence. A key technical differentiator is its complete independence from cloud services or external APIs, operating entirely browser-based and integrating with local AI models via Ollama. The developer is seeking feedback on the practical relevance and effectiveness of this technical interpretation of AI regulation in real-world scenarios.

02

Show HN: Scroll Wikipedia like TikTok

This Hacker News submission introduces a novel application demonstrating fully generative user interfaces (UIs), where HTML and Canvas elements are generated just-in-time, powered by Large Language Models like Gemini 3 Flash. The platform simulates a TikTok-like scrolling feed, with each post being dynamically streamed and rendered. A key technical aspect involves the use of Cloudflare Workers Durable Objects to facilitate fast, bidirectional communication for features such as comments and direct messages. Additionally, generated content is stored in a Durable Object SQLite database, optimizing subsequent delivery to user feeds. The developer highlights inspiration from previous projects, including a VSCode extension called Wikitok and other generative UI experiments, aiming to create an immersive, dynamic content consumption experience akin to popular short-form video platforms but applied to informational content like Wikipedia.

03

Show HN: Executable Markdown files with Unix pipes

An innovative open-source tool has been unveiled, designed to transform Markdown files into executable scripts using a shebang line. This functionality allows users to pipe Markdown content through "Claude Code," an AI-powered execution environment that supports full stdin/stdout capabilities, akin to traditional Unix programs. This enables Markdown files to not only contain descriptive text but also to execute complex tasks, run shell commands, generate scripts, read files, and make API calls, effectively turning them into dynamic command orchestrators. The tool facilitates powerful workflows, such as automating test execution and summarizing results directly within a Markdown document. Its seamless integration with Unix pipes allows for chaining multiple operations, enhancing its utility for data processing and complex automation scenarios. This development represents a significant step towards bridging the gap between documentation and execution, offering a flexible and powerful way to manage and run code-related tasks within a familiar Markdown format, thereby streamlining development and operational processes.

04

Anthropic blocks third-party use of Claude Code subscriptions

Anthropic, a prominent artificial intelligence research company known for its Claude family of large language models, has reportedly implemented stringent measures to block the third-party use of its specialized Claude Code subscriptions. This significant development, initially brought to light through discussions on GitHub, specifically within the anomalyco/opencode repository, indicates a strategic shift in how AI service providers manage access and licensing for their advanced generative AI tools dedicated to code generation. The action suggests Anthropic is asserting tighter control over the distribution and utilization of its powerful code-generating capabilities, potentially aiming to safeguard intellectual property, ensure strict compliance with its terms of service, or to optimize resource allocation specifically for its direct subscribers. This decision could carry substantial implications for developers, AI agents, and integrated platforms that have historically relied on Claude Code through indirect or aggregated third-party channels, potentially necessitating a swift re-evaluation of existing integration strategies and impacting the broader ecosystem of third-party AI tool access and availability. The incident highlights the complex and rapidly evolving landscape of commercial AI deployment and the inherent challenges in managing widespread AI model access within a dynamic and expanding global developer community.

05

AI Zealotry

The concept of 'AI Zealotry' highlights a growing concern within the technology community regarding an overly enthusiastic and often uncritical approach to artificial intelligence development and adoption. This phenomenon, characterized by an unquestioning belief in AI's transformative power and a tendency to overlook its limitations or potential pitfalls, warrants careful examination. Critics argue that such zealotry can lead to unrealistic expectations, misallocation of resources, and a neglect of ethical considerations and societal impacts. Rather than fostering genuine innovation, an unbridled enthusiasm might inadvertently stifle critical discourse and hinder the development of robust, equitable, and responsible AI systems. The discourse around AI zealotry calls for a more balanced and pragmatic perspective, encouraging developers, policymakers, and the public alike to temper excitement with a healthy dose of skepticism, rigorous evaluation, and a commitment to addressing the complex challenges inherent in advanced AI technologies. It emphasizes the importance of understanding AI as a tool, with inherent strengths and weaknesses, rather than an infallible solution or an object of blind faith, advocating for cautious optimism.

06

How to store a chess position in 26 bytes (2022)

This article presents an ingenious method for compactly storing a complete chess board position using only 26 bytes, a remarkable feat of data compression. The technique relies on sophisticated bit-level manipulation and optimized encoding strategies to represent all critical game elements. These include the precise placement of each piece, the availability of castling rights for both players, the potential en passant target square, and the halfmove clock, which tracks moves since the last capture or pawn push. By drastically minimizing the memory footprint for each unique position, this approach offers substantial advantages for the development of high-performance chess engines, extensive chess databases, and advanced analysis software. Such efficiency enables significantly faster state lookups, reduces overall memory overhead, and improves cache utilization. The methodology vividly illustrates how deep understanding of game rules combined with clever bit packing can lead to profound data compression, providing invaluable practical insights for engineers developing resource-constrained applications or systems demanding rapid processing and storage of complex game states. This low-level optimization showcases the profound impact on storage and retrieval performance.

huggingface

6 stories
01

Agent-as-a-Judge

LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and multi-step, the reliability of LLM-as-a-Judge has become constrained by inherent biases, shallow single-pass reasoning, and the inability to verify assessments against real-world observations. This has catalyzed the transition to Agent-as-a-Judge, where agentic judges employ planning, tool-augmented verification, multi-agent collaboration, and persistent memory to enable more robust, verifiable, and nuanced evaluations. Despite the rapid proliferation of agentic evaluation systems, the field lacks a unified framework to navigate this shifting landscape. To bridge this gap, we present the first comprehensive survey tracing this evolution. Specifically, we identify key dimensions that characterize this paradigm shift and establish a developmental taxonomy. We organize core methodologies and survey applications across general and professional domains. Furthermore, we analyze frontier challenges and identify promising research directions, ultimately providing a clear roadmap for the next generation of agentic evaluation.

02

Plenoptic Video Generation

Camera-controlled generative video re-rendering methods, such as ReCamMaster, have achieved remarkable progress. However, despite their success in single-view setting, these works often struggle to maintain consistency across multi-view scenarios. Ensuring spatio-temporal coherence in hallucinated regions remains challenging due to the inherent stochasticity of generative models. To address it, we introduce PlenopticDreamer, a framework that synchronizes generative hallucinations to maintain spatio-temporal memory. The core idea is to train a multi-in-single-out video-conditioned model in an autoregressive manner, aided by a camera-guided video retrieval strategy that adaptively selects salient videos from previous generations as conditional inputs. In addition, Our training incorporates progressive context-scaling to improve convergence, self-conditioning to enhance robustness against long-range visual degradation caused by error accumulation, and a long-video conditioning mechanism to support extended video generation. Extensive experiments on the Basic and Agibot benchmarks demonstrate that PlenopticDreamer achieves state-of-the-art video re-rendering, delivering superior view synchronization, high-fidelity visuals, accurate camera control, and diverse view transformations (e.g., third-person to third-person, and head-view to gripper-view in robotic manipulation). Project page: https://research.nvidia.com/labs/dir/plenopticdreamer/

03

AgentDevel: Reframing Self-Evolving LLM Agents as Release Engineering

Recent progress in large language model (LLM) agents has largely focused on embedding self-improvement mechanisms inside the agent or searching over many concurrent variants. While these approaches can raise aggregate scores, they often yield unstable and hard-to-audit improvement trajectories, making it difficult to guarantee non-regression or to reason about failures across versions. We reframe agent improvement as release engineering: agents are treated as shippable artifacts, and improvement is externalized into a regression-aware release pipeline. We introduce AgentDevel, a release engineering pipeline that iteratively runs the current agent, produces implementation-blind, symptom-level quality signals from execution traces, synthesizes a single release candidate (RC) via executable diagnosis, and promotes it under flip-centered gating. AgentDevel features three core designs: (i) an implementation-blind LLM critic that characterizes failure appearances without accessing agent internals, (ii) script-based executable diagnosis that aggregates dominant symptom patterns and produces auditable engineering specifications, and (iii) flip-centered gating that prioritizes pass to fail regressions and fail to pass fixes as first-class evidence. Unlike population-based search or in-agent self-refinement, AgentDevel maintains a single canonical version line and emphasizes non-regression as a primary objective. Experiments on execution-heavy benchmarks demonstrate that AgentDevel yields stable improvements with significantly fewer regressions while producing reproducible, auditable artifacts. Overall, AgentDevel provides a practical development discipline for building, debugging, and releasing LLM agents as software development.

04

AT^2PO: Agentic Turn-based Policy Optimization via Tree Search

LLM agents have emerged as powerful systems for tackling multi-turn tasks by interleaving internal reasoning and external tool interactions. Agentic Reinforcement Learning has recently drawn significant research attention as a critical post-training paradigm to further refine these capabilities. In this paper, we present AT^2PO (Agentic Turn-based Policy Optimization via Tree Search), a unified framework for multi-turn agentic RL that addresses three core challenges: limited exploration diversity, sparse credit assignment, and misaligned policy optimization. AT^2PO introduces a turn-level tree structure that jointly enables Entropy-Guided Tree Expansion for strategic exploration and Turn-wise Credit Assignment for fine-grained reward propagation from sparse outcomes. Complementing this, we propose Agentic Turn-based Policy Optimization, a turn-level learning objective that aligns policy updates with the natural decision granularity of agentic interactions. ATPO is orthogonal to tree search and can be readily integrated into any multi-turn RL pipeline. Experiments across seven benchmarks demonstrate consistent improvements over the state-of-the-art baseline by up to 1.84 percentage points in average, with ablation studies validating the effectiveness of each component. Our code is available at https://github.com/zzfoutofspace/ATPO.

05

One Sample to Rule Them All: Extreme Data Efficiency in RL Scaling

The reasoning ability of large language models (LLMs) can be unleashed with reinforcement learning (RL) (OpenAI, 2024; DeepSeek-AI et al., 2025a; Zeng et al., 2025). The success of existing RL attempts in LLMs usually relies on high-quality samples of thousands or beyond. In this paper, we challenge fundamental assumptions about data requirements in RL for LLMs by demonstrating the remarkable effectiveness of one-shot learning. Specifically, we introduce polymath learning, a framework for designing one training sample that elicits multidisciplinary impact. We present three key findings: (1) A single, strategically selected math reasoning sample can produce significant performance improvements across multiple domains, including physics, chemistry, and biology with RL; (2) The math skills salient to reasoning suggest the characteristics of the optimal polymath sample; and (3) An engineered synthetic sample that integrates multidiscipline elements outperforms training with individual samples that naturally occur. Our approach achieves superior performance to training with larger datasets across various reasoning benchmarks, demonstrating that sample quality and design, rather than quantity, may be the key to unlock enhanced reasoning capabilities in language models. Our results suggest a shift, dubbed as sample engineering, toward precision engineering of training samples rather than simply increasing data volume.

06

VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control

Video world models aim to simulate dynamic, real-world environments, yet existing methods struggle to provide unified and precise control over camera and multi-object motion, as videos inherently operate dynamics in the projected 2D image plane. To bridge this gap, we introduce VerseCrafter, a 4D-aware video world model that enables explicit and coherent control over both camera and object dynamics within a unified 4D geometric world state. Our approach is centered on a novel 4D Geometric Control representation, which encodes the world state through a static background point cloud and per-object 3D Gaussian trajectories. This representation captures not only an object's path but also its probabilistic 3D occupancy over time, offering a flexible, category-agnostic alternative to rigid bounding boxes or parametric models. These 4D controls are rendered into conditioning signals for a pretrained video diffusion model, enabling the generation of high-fidelity, view-consistent videos that precisely adhere to the specified dynamics. Unfortunately, another major challenge lies in the scarcity of large-scale training data with explicit 4D annotations. We address this by developing an automatic data engine that extracts the required 4D controls from in-the-wild videos, allowing us to train our model on a massive and diverse dataset.