NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-06DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Launch HN: Freestyle: Sandboxes for AI Coding Agents

Freestyle, founded by Ben and Jacob, is spearheading the development of a dedicated cloud infrastructure tailored for advanced AI Coding Agents. Their core offering involves providing sophisticated 'sandboxes' that function as comprehensive computing environments, designed to be seamlessly interchangeable with services like AWS EC2 from the perspective of an AI agent. This approach represents a significant leap from the previous generation of AI agents, which were often confined to limited toolsets or basic serverless application deployments, such as SQL script generation or simple web app building. A cornerstone technical innovation highlighted by Freestyle is the capability to horizontally fork a running sandbox, replicating its entire memory state in less than 400 milliseconds. This feature is crucial for enabling highly efficient, scalable, and concurrent operation of AI agents, allowing them to fully exploit computational resources without incurring substantial performance penalties or operational delays. By offering these advanced sandboxes, Freestyle aims to empower the latest generation of AI agents to tackle more complex and resource-intensive coding and development tasks.

02

Show HN: Gemma Gem – AI model embedded in a browser – no API keys, no cloud

Gemma Gem is a innovative Chrome extension that integrates Google's Gemma 4 (2B) AI model directly within the browser environment. This client-side approach, powered by WebGPU in an offscreen document, eliminates the need for external API keys or cloud infrastructure. The extension empowers the Gemma model with a comprehensive set of tools, enabling it to interact extensively with webpages by reading content, taking screenshots, clicking elements, typing text, scrolling, and executing JavaScript. Users interact through a chat overlay, where the AI attempts to understand queries and utilize appropriate tools, often displaying a chain-of-thought reasoning process. While effective for simple page questions and JavaScript execution, the 2B model exhibits limitations in complex, multi-step tool chains. Notably, its self-contained agent loop has zero external dependencies and can be extracted as a standalone library.

03

Claude Code is unusable for complex engineering tasks with the Feb updates

Recent user feedback and reports suggest that Anthropic's Claude Code has experienced a significant performance degradation following its February updates, rendering it "unusable" for complex engineering tasks. This regression, highlighted on platforms like GitHub, indicates a critical issue in the AI's ability to effectively handle intricate coding challenges and demanding development workflows. The reported setback raises serious concerns among developers who rely on Claude Code for advanced programming assistance, potentially impacting productivity and project timelines. The incident underscores the challenges in maintaining consistent model efficacy and reliability across successive updates, especially for specialized AI applications such as code generation and complex problem-solving within software engineering environments. The community is now seeking prompt resolution and clarification regarding these performance shortcomings, which directly affect the practical utility of the AI tool in professional settings.

04

Show HN: I built a tiny LLM to demystify how language models work

A developer has unveiled 'guppylm', a lightweight Large Language Model (LLM) specifically engineered to demystify the internal workings of complex language models. Shared on Hacker News, this project features a ~9 million parameter LLM built entirely from scratch, employing a straightforward vanilla transformer architecture. The model's training regimen involved 60,000 synthetic conversations, with its entire PyTorch implementation condensed into approximately 130 lines of code. A key highlight is its remarkable efficiency, capable of completing training in just five minutes on a free Google Colab T4 GPU, which significantly lowers the barrier to entry for experimentation and learning. The creator positions guppylm as an accessible educational tool, inviting users to fork the repository and explore customizing the model's personality for various character-driven applications. This initiative provides a tangible, hands-on resource for developers and enthusiasts seeking a clearer understanding of modern AI language model principles without extensive computational resources.

05

Agent Reading Test

This Hacker News story, titled 'Agent Reading Test', introduces a dedicated initiative focused on developing and implementing a standardized evaluation methodology for artificial intelligence agents. Although the provided content is concise, primarily comprising the title, the accompanying URLs—agentreadingtest.com and a specific blog post detailing 'designing-agent-reading-test' from dacharycarey.com—strongly indicate a comprehensive project. This project aims to establish robust metrics and frameworks for assessing the efficacy with which AI agents can comprehend, interpret, and process textual information. The overarching objective appears to be the creation of a definitive benchmark to objectively measure agent reading capabilities across various sophisticated language understanding tasks, which may include intricate question answering, precise summarization, and efficient information extraction. Such a testing paradigm is critical for accelerating AI research and development, providing a quantifiable standard for researchers and developers to effectively gauge and improve the linguistic intelligence and performance of their advanced AI models.

06

Wikipedia's AI agent row likely just the beginning of the bot-ocalypse

The recent controversy surrounding AI agents operating on Wikipedia is indicative of a nascent but rapidly escalating challenge to the integrity of online information platforms. Dubbed potentially 'the beginning of the bot-ocalypse,' this situation highlights the growing complexities introduced by autonomous AI entities in curating, generating, and potentially manipulating content at scale. While specific details of the Wikipedia 'row' are not provided, the title strongly implies instances where AI agents have either disrupted established editorial processes, contributed problematic information, or raised significant concerns about source veracity and human oversight. Experts suggest that such incidents on a platform as globally influential as Wikipedia serve as a critical early warning. As AI technology, particularly large language models and autonomous agents, becomes more sophisticated and accessible, the ability for malicious or even poorly-designed bots to spread misinformation, erode trust, and overwhelm human moderation efforts will exponentially increase. This development necessitates urgent strategic responses from platform operators, policymakers, and the AI community to safeguard the digital commons against an impending wave of automated content challenges, ensuring the continued reliability and authenticity of public knowledge repositories.

huggingface

6 stories
01

Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

Distilling video generation models to extremely low inference budgets (e.g., 2--4 NFEs) is crucial for real-time deployment, yet remains challenging. Trajectory-style consistency distillation often becomes conservative under complex video dynamics, yielding an over-smoothed appearance and weak motion. Distribution matching distillation (DMD) can recover sharp, mode-seeking samples, but its local training signals do not explicitly regularize how denoising updates compose across timesteps, making composed rollouts prone to drift. To overcome this challenge, we propose Self-Consistent Distribution Matching Distillation (SC-DMD), which explicitly regularizes the endpoint-consistent composition of consecutive denoising updates. For real-time autoregressive video generation, we further treat the KV cache as a quality parameterized condition and propose Cache-Distribution-Aware training. This training scheme applies SC-DMD over multi-step rollouts and introduces a cache-conditioned feature alignment objective that steers low-quality outputs toward high-quality references. Across extensive experiments on both non-autoregressive backbones (e.g., Wan~2.1) and autoregressive real-time paradigms (e.g., Self Forcing), our method, dubbed Salt, consistently improves low-NFE video generation quality while remaining compatible with diverse KV-cache memory mechanisms. Source code will be released at https://github.com/XingtongGe/Salt.

02

Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall short: they lack flexible tool integration, test visual and search tools separately, and evaluate primarily by final answers. Consequently, they cannot verify if tools were actually invoked, applied correctly, or used efficiently. To address this, we introduce Agentic-MME, a process-verified benchmark for Multimodal Agentic Capabilities. It contains 418 real-world tasks across 6 domains and 3 difficulty levels to evaluate capability synergy, featuring over 2,000 stepwise checkpoints that average 10+ person-hours of manual annotation per task. Each task includes a unified evaluation framework supporting sandboxed code and APIs, alongside a human reference trajectory annotated with stepwise checkpoints along dual-axis: S-axis and V-axis. To enable true process-level verification, we audit fine-grained intermediate states rather than just final answers, and quantify efficiency via an overthinking metric relative to human trajectories. Experimental results show the best model, Gemini3-pro, achieves 56.3% overall accuracy, which falls significantly to 23.0% on Level-3 tasks, underscoring the difficulty of real-world multimodal agentic problem solving.

03

AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents

Computer-use agents extend language models from text generation to persistent action over tools, files, and execution environments. Unlike chat systems, they maintain state across interactions and translate intermediate outputs into concrete actions. This creates a distinct safety challenge in that harmful behavior may emerge through sequences of individually plausible steps, including intermediate actions that appear locally acceptable but collectively lead to unauthorized actions. We present AgentHazard, a benchmark for evaluating harmful behavior in computer-use agents. AgentHazard contains 2,653 instances spanning diverse risk categories and attack strategies. Each instance pairs a harmful objective with a sequence of operational steps that are locally legitimate but jointly induce unsafe behavior. The benchmark evaluates whether agents can recognize and interrupt harm arising from accumulated context, repeated tool use, intermediate actions, and dependencies across steps. We evaluate AgentHazard on Claude Code, OpenClaw, and IFlow using mostly open or openly deployable models from the Qwen3, Kimi, GLM, and DeepSeek families. Our experimental results indicate that current systems remain highly vulnerable. In particular, when powered by Qwen3-Coder, Claude Code exhibits an attack success rate of 73.63%, suggesting that model alignment alone does not reliably guarantee the safety of autonomous agents.

04

Self-Distilled RLVR

On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals for each sampled trajectory, in contrast to reinforcement learning with verifiable rewards (RLVR), which only obtains sparse signals from verifiable outcomes in the environment. Recently, the community has explored on-policy self-distillation (OPSD), where the same model serves as both teacher and student, with the teacher receiving additional privileged information such as reference answers to enable self-evolution. This paper demonstrates that learning signals solely derived from the privileged teacher result in severe information leakage and unstable long-term training. Accordingly, we identify the optimal niche for self-distillation and propose RLSD (RLVR with Self-Distillation). Specifically, we leverage self-distillation to obtain token-level policy differences for determining fine-grained update magnitudes, while continuing to use RLVR to derive reliable update directions from environmental feedback (e.g., response correctness). This enables RLSD to simultaneously harness the strengths of both RLVR and OPSD, achieving a higher convergence ceiling and superior training stability.

05

Test-Time Scaling Makes Overtraining Compute-Optimal

Modern LLMs scale at test-time, e.g. via repeated sampling, where inference cost grows with model size and the number of samples. This creates a trade-off that pretraining scaling laws, such as Chinchilla, do not address. We present Train-to-Test (T^2) scaling laws that jointly optimize model size, training tokens, and number of inference samples under fixed end-to-end budgets. T^2 modernizes pretraining scaling laws with pass@k modeling used for test-time scaling, then jointly optimizes pretraining and test-time decisions. Forecasts from T^2 are robust over distinct modeling approaches: measuring joint scaling effect on the task loss and modeling impact on task accuracy. Across eight downstream tasks, we find that when accounting for inference cost, optimal pretraining decisions shift radically into the overtraining regime, well-outside of the range of standard pretraining scaling suites. We validate our results by pretraining heavily overtrained models in the optimal region that T^2 scaling forecasts, confirming their substantially stronger performance compared to pretraining scaling alone. Finally, as frontier LLMs are post-trained, we show that our findings survive the post-training stage, making T^2 scaling meaningful in modern deployments.

06

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cognition. Existing frameworks suffer from narrow domain coverage, reliance on generalist tasks, or self-evaluation biases. To bridge this gap, we present XpertBench, a high-fidelity benchmark engineered to assess LLMs across authentic professional domains. XpertBench consists of 1,346 meticulously curated tasks across 80 categories, spanning finance, healthcare, legal services, education, and dual-track research (STEM and Humanities). These tasks are derived from over 1,000 submissions by domain experts--including researchers from elite institutions and practitioners with extensive clinical or industrial experience--ensuring superior ecological validity. Each task uses detailed rubrics with mostly 15-40 weighted checkpoints to assess professional rigor. To facilitate scalable yet human-aligned assessment, we introduce ShotJudge, a novel evaluation paradigm that employs LLM judges calibrated with expert few-shot exemplars to mitigate self-rewarding biases. Our empirical evaluation of state-of-the-art LLMs reveals a pronounced performance ceiling: even leading models achieve a peak success rate of only ~66%, with a mean score around 55%. Models also exhibit domain-specific divergence, showing non-overlapping strengths in quantitative reasoning versus linguistic synthesis.. These findings underscore a significant "expert-gap" in current AI systems and establish XpertBench as a critical instrument for navigating the transition from general-purpose assistants to specialized professional collaborators.