NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-07DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Agents need control flow, not more prompts

The article, "Agents need control flow, not more prompts," posits that the current paradigm of AI agent development, heavily reliant on elaborate prompting strategies, is fundamentally limited for achieving robust and autonomous behavior. It argues that simply providing more detailed or structured prompts does not solve the underlying architectural deficiencies. Instead, the piece advocates for integrating explicit control flow mechanisms, similar to those found in traditional programming. This would enable AI agents to execute sequential tasks, implement conditional logic, manage state, and iterate through processes more effectively. By moving beyond a purely reactive, prompt-driven model to one incorporating structured execution paths, agents could exhibit greater reliability, reduce instances of "hallucination," and tackle more complex, multi-step problems with enhanced efficiency and predictability. This paradigm shift suggests a blend of large language model capabilities with conventional software engineering principles to create more capable and dependable AI systems.

02

Natural Language Autoencoders: Turning Claude's Thoughts into Text

Anthropic's latest research introduces Natural Language Autoencoders (NLAEs), a groundbreaking method aimed at enhancing the interpretability of large language models (LLMs) such as Claude. NLAEs function by translating the complex, high-dimensional internal activations, or "thoughts," of an LLM into concise, human-understandable natural language explanations. This innovative approach allows researchers to peer into the model's reasoning process, offering unprecedented insights into how it processes information, forms concepts, and arrives at specific outputs. By effectively decoding the latent space of LLMs, NLAEs represent a significant leap forward in mechanistic interpretability. This capability is crucial for identifying biases, debugging errors, and verifying the safety of advanced AI systems. Ultimately, this research not only deepens our understanding of current LLM architectures but also paves the way for developing more transparent, controllable, and trustworthy artificial intelligence.

03

AlphaEvolve: Gemini-powered coding agent scaling impact across fields

AlphaEvolve is introduced as an innovative coding agent, leveraging the capabilities of Google's Gemini large language model. This advanced AI system is engineered to enhance and automate various aspects of the software development lifecycle, from code generation and optimization to debugging and deployment across diverse fields. Its primary objective is to scale the impact of AI in coding, enabling developers and researchers to accelerate progress in complex domains. By integrating powerful generative AI functionalities, AlphaEvolve aims to significantly improve efficiency, reduce development time, and foster breakthroughs in areas that require intricate computational solutions. The agent's design emphasizes adaptability, allowing it to be applied to a wide array of programming challenges and scientific endeavors, thereby expanding the practical applications of AI-driven code generation and intelligent automation. This initiative highlights the growing trend of AI agents augmenting human expertise in specialized technical fields.

04

Agent-harness-kit scaffolding for multi-agent workflows (MCP, provider-agnostic)

The Agent-harness-kit offers a comprehensive scaffolding solution tailored for the development and deployment of sophisticated multi-agent workflows. This toolkit is specifically engineered to be provider-agnostic, enabling developers to seamlessly integrate and interchange various underlying AI model providers, such as large language models, into their agent-based applications. It empowers the creation of complex systems where multiple AI agents can collaborate efficiently, likely guided by a formal Multi-Agent Coordination Protocol (MCP) to ensure structured interaction, task delegation, and conflict resolution. The primary objective of this kit is to streamline the orchestration, communication, and execution of intricate tasks across disparate agent components, thereby providing a robust and flexible framework for building scalable AI agent solutions. By offering standardized tooling and promoting interoperability, the Agent-harness-kit significantly reduces the development overhead and complexity associated with designing and implementing advanced multi-agent AI systems, addressing a critical need in the evolving AI landscape.

05

ZAYA1-8B matches DeepSeek-R1 on math with less than 1B active parameters

A new language model, ZAYA1-8B, has achieved a notable milestone in mathematical reasoning, demonstrating performance comparable to the more established DeepSeek-R1 model, despite employing significantly fewer active parameters. This breakthrough is particularly significant as ZAYA1-8B operates with less than one billion active parameters, implying a much more efficient architecture for both training and inference. Such a reduction in active parameters typically leads to lower computational resource requirements, making advanced AI capabilities more accessible and cost-effective. The model's ability to match the mathematical prowess of DeepSeek-R1, which is a prominent model in the domain, highlights advancements in model design and optimization for specialized tasks. This development is crucial for applications that demand strong mathematical problem-solving within resource-constrained environments and for researchers focused on developing high-performing yet compact large language models. Furthermore, ZAYA1-8B's open-source availability is expected to foster innovation and broader adoption of efficient AI solutions in areas like mathematics and coding. This benchmark underscores the continuous progress in creating powerful and resource-efficient AI systems.

06

ProgramBench: Can language models rebuild programs from scratch?

ProgramBench introduces a novel and comprehensive benchmark specifically designed to evaluate the capability of large language models (LLMs) to rebuild executable programs from scratch. This benchmark addresses a crucial limitation in existing evaluation frameworks, which often focus on narrower tasks like code completion or minor debugging, rather than the more challenging problem of holistic program synthesis. ProgramBench challenges LLMs to generate complete, functional programs based solely on high-level problem descriptions or specifications, simulating complex software development scenarios. By requiring models to demonstrate a deep understanding of programming logic, algorithmic design, and intricate problem-solving, the benchmark aims to provide a more accurate assessment of their true generative potential in software engineering. Initial applications of ProgramBench indicate that while LLMs show considerable progress in various coding aspects, their ability to autonomously construct entire, sophisticated programs from conceptual stages still presents a significant frontier for ongoing research and development in AI-driven code generation. This initiative offers a standardized platform for tracking advancements in this critical area.

huggingface

6 stories
01

Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation

Distillation-based acceleration has become foundational for making autoregressive streaming video diffusion models practical, with distribution matching distillation (DMD) as the de facto choice. Existing methods, however, train the student to match the teacher's output indiscriminately, treating every rollout, frame, and pixel as equally reliable supervision. We argue that this caps distilled quality, since it overlooks two complementary axes of variance in DMD supervision: Inter-Reliability across student rollouts whose supervision varies in reliability, and Intra-Perplexity across spatial regions and temporal frames that contribute unequally to where quality can still be improved. The objective thus conflates two questions under a uniform weight: whether to learn from each rollout, and where to concentrate optimization within it. To address this, we propose Stream-R1, a Reliability-Perplexity Aware Reward Distillation framework that adaptively reweights the distillation objective at both rollout and spatiotemporal-element levels through a single shared reward-guided mechanism. At the Inter-Reliability level, Stream-R1 rescales each rollout's loss by an exponential of a pretrained video reward score, so that rollouts with reliable supervision dominate optimization. At the Intra-Perplexity level, it back-propagates the same reward model to extract per-pixel gradient saliency, which is factored into spatial and temporal weights that concentrate optimization pressure on regions and frames where refinement yields the largest expected gain. An adaptive balancing mechanism prevents any single quality axis from dominating across visual quality, motion quality, and text alignment. Stream-R1 attains consistent improvements on all three dimensions over distillation baselines on standard streaming video generation benchmarks, without architectural modification or additional inference cost.

02

OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi-step reasoning. Despite rapid progress, top-tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high-quality training data, transparent trajectory synthesis pipelines, or detailed training recipes. To this end, we introduce OpenSearch-VL, a fully open-source recipe for training frontier multimodal deep search agents with agentic reinforcement learning. First, we curated a dedicated pipeline to construct high-quality training data through Wikipedia path sampling, fuzzy entity rewriting, and source-anchor visual grounding, which jointly reduce shortcuts and one-step retrieval collapse. Based on this pipeline, we curate two training datasets, SearchVL-SFT-36k for SFT and SearchVL-RL-8k for RL. Besides, we design a diverse tool environment that unifies text search, image search, OCR, cropping, sharpening, super-resolution, and perspective correction, enabling agents to combine active perception with external knowledge acquisition. Finally, we propose a multi-turn fatal-aware GRPO training algorithm that handles cascading tool failures by masking post-failure tokens while preserving useful pre-failure reasoning through one-sided advantage clamping. Built on this recipe, OpenSearch-VL delivers substantial performance gains, with over 10-point average improvements across seven benchmarks, and achieves results comparable to proprietary commercial models on several tasks. We will release all data, code, and models to support open research on multimodal deep search agents.

03

MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Recent progress in multimodal large language models (MLLMs) has brought AI capabilities from static offline data processing to real-time streaming interaction, yet they still remain far from human-level multimodal interaction. The key bottlenecks are no longer modality coverage or latency alone, but the interaction paradigm itself. First, perception and response are still separated into alternating phases, preventing models from incorporating new inputs for timely adjustment during generation. Second, most current models remain reactive, responding only to explicit user requests instead of acting proactively in the evolving multimodal environment. We present MiniCPM-o 4.5, our latest effort towards human-like multimodal interaction, which mitigates these gaps by real-time full-duplex omni-modal interaction. It can see, listen, and speak simultaneously in real-time, while also exhibiting proactive behaviors such as issuing reminders or comments based on its continuous understanding of the live scene. The key technique behind MiniCPM-o 4.5 is Omni-Flow, a unified streaming framework that aligns omni-modal inputs and outputs along a shared temporal axis. This formulation converts conventional turn-based interaction into a full-duplex, time-aligned process, enabling simultaneous perception and response and allowing proactive behavior to arise within the same framework. With a total of 9B parameters, MiniCPM-o 4.5 approaches Gemini 2.5 Flash in vision-language capabilities, delivering state-of-the-art open-source performance at its scale. It also surpasses Qwen3-Omni-30B-A3B in omni-modal understanding and delivers better speech generation, with significantly higher computation efficiency. Driven by its efficient architecture design and inference optimization, the model can perform real-time full-duplex omni-modal interaction on edge devices with less than 12GB RAM cost.

04

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning

Reinforcement Learning with Verifiable Rewards (RLVR) enhances reasoning of Large Language Models (LLMs) but usually exhibits limited generation diversity due to the over-incentivization of positive rewards. Although methods like Negative Sample Reinforcement (NSR) mitigate this issue by upweighting penalty from negative samples, they may suppress the semantic distributions shared between positive and negative responses. To boost reasoning ability without losing diversity, this paper proposes negative sample projection Residual Reinforcement Learning (ResRL) that decouples similar semantic distributions among positive and negative responses. We theoretically link Lazy Likelihood Displacement (LLD) to negative-positive head-gradient interference and derive a single-forward proxy that upper-bounds representation alignment to guide conservative advantage reweighting. ResRL then projects negative-token hidden representations onto an SVD-based low-rank positive subspace and uses projection residuals to modulate negative gradients, improving reasoning while preserving diversity and outperforming strong baselines on average across twelve benchmarks spanning Mathematics, Code, Agent Tasks, and Function Calling. Notably, ResRL surpasses NSR on mathematical reasoning by 9.4% in Avg@16 and 7.0% in Pass@128. Code is available at https://github.com/1229095296/ResRL.git.

05

SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies

The emergence of "vibe coding" platforms, where users describe applications in natural language and AI agents autonomously generate full-stack software, has created a need for rigorous evaluation beyond code-level benchmarks. In order to assess them as virtual software development agencies on understanding business requirements, making architectural decisions, writing production code, handling iterative modifications, and maintaining business readiness, we introduce SWE-WebDev Bench, a 68-metric evaluation framework spanning 25 primary and 43 diagnostic metrics across seven groups, organized along three dimensions: Interaction Mode (App Creation Request (ACR) vs. App Modification Request (AMR)), Agency Angle (Product Manager (PM), Engineering, Ops), and Complexity Tier (T4 multi-role SaaS, T5 AI-native). Our evaluation (six platforms, three domains, 18 evaluation cells) reveals four recurring shortcomings in the current generation of AI app builders: (1) A specification bottleneck, where platforms compress rich business requirements into oversimplified technical plans, (2) A pervasive frontend-backend decoupling, where visually polished UIs mask absent or broken backend infrastructure, (3) A steep production-readiness cliff, where no platform scores above 60% on engineering quality and post-generation human effort varies substantially across platforms and (4) Widespread security and infrastructure failures, with no platform exceeding 65% Security Score against a 90% target and concurrency handling as low as 6%. These observations are descriptive of our sample and require larger-scale replication to establish generality. We release SWE-WebDev Bench as a community benchmark to enable such replication and help platform builders identify and address these gaps. Code and benchmark resources are available at: https://github.com/snowmountainAi/webdevbench and https://webdevbench.com/.

06

CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing

Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the lens of creative tool use, where a model repurposes available objects by reasoning about their affordances and attributes rather than relying on canonical usage. As a first step, we introduce CreativityBench, a benchmark for evaluating affordance-based creativity in LLMs. To this end, we build a large-scale affordance knowledge base (KB) with 4K entities and 150K+ affordance annotations, explicitly linking objects, parts, attributes, and actionable uses. Building on this KB, we generate 14K grounded tasks that require identifying non-obvious yet physically plausible solutions under constraints. Evaluations across 10 state-of-the-art LLMs, including closed and open-source models, show that models can often select a plausible object, but fail to identify the correct parts, their affordances, and the underlying physical mechanism needed to solve the task, leading to a significant drop in performance. Furthermore, improvements from model scaling quickly saturate, strong general reasoning does not reliably translate to creative affordance discovery, and common inference-time strategies such as Chain-of-Thought yield limited gains. These results suggest that creative tool use remains a major challenge for current models, and that CreativityBench provides a useful testbed for studying this missing dimension of intelligence, with potential implications for planning and reasoning modules in future agents.