NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-01DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

TurboQuant KV Compression and SSD Expert Streaming for M5 Pro and IOS

A novel system, referred to as TurboQuant KV Compression and SSD Expert Streaming, is introduced, specifically engineered to optimize the deployment and performance of artificial intelligence workloads on Apple's M5 Pro processors and the iOS ecosystem. This innovative technology primarily addresses the persistent challenges of memory and storage constraints inherent in mobile and edge devices. TurboQuant KV Compression focuses on efficiently reducing the memory footprint of key-value stores, a critical component in modern transformer-based models, by applying advanced quantization techniques. Concurrently, SSD Expert Streaming is designed to enable high-throughput and low-latency data access from Solid State Drives, effectively allowing larger models or extensive datasets to be streamed and processed than would typically fit into device RAM. This integrated approach is poised to significantly enhance the capability of running sophisticated AI applications, particularly large language models, directly on M5 Pro-equipped devices. The aim is to improve user experience through faster on-device inference, reduce reliance on cloud infrastructure, and unlock new possibilities for AI applications in mobile environments.

02

StepFun 3.5 Flash is #1 cost-effective model for OpenClaw tasks (300 battles)

A recent evaluation within the OpenClaw task environment has identified StepFun 3.5 Flash as the leading model in terms of cost-effectiveness, following a rigorous assessment involving 300 simulated battles. This benchmark highlights the model's optimized performance, demonstrating its ability to achieve superior results while minimizing operational costs, a critical factor in the deployment and scaling of AI solutions. The OpenClaw tasks likely represent a challenging domain requiring strategic decision-making and efficient resource utilization. The consistently strong performance of StepFun 3.5 Flash across a substantial number of trials suggests robust design and effective training methodologies, allowing it to deliver high-value outcomes without incurring excessive computational or resource expenditures. This achievement positions StepFun 3.5 Flash as a significant development for practitioners and researchers focused on developing and deploying economically viable AI agents. The findings are particularly relevant for applications where efficiency is paramount, offering a compelling case for its adoption in resource-constrained environments or large-scale operations. Further analysis will likely delve into the specific architectural or algorithmic innovations contributing to this cost-efficiency breakthrough.

03

Show HN: Real-time dashboard for Claude Code agent teams

Agents Observe is an open-source project introducing a real-time dashboard designed for monitoring Claude Code agent teams. Initiated as an exploration into building automation harnesses for Claude Code, the platform allows users to observe agent activities, filter, and search their outputs in real-time. Key technical insights gained during development include the significant performance degradation caused by blocking Claude code hooks when numerous plugins are active, contrasting with the superior information provided by hooks compared to OTEL data. The project also highlights that Claude's JSONL files offer a comprehensive view of agent operations. A crucial finding was the substantial performance improvement achieved by transitioning to background (fire and forget) hooks and deactivating other plugins, underscoring the often-overlooked impact of plugin configurations on Claude's efficiency. The "Agents Observe" plugin leverages Docker for deploying its API and dashboard services. This tool offers crucial visibility into complex AI agent behaviors, facilitating better understanding and optimization of multi-agent systems.

04

The OpenAI Graveyard: All the Deals and Products That Haven't Happened

An analysis titled 'The OpenAI Graveyard' critically examines various deals, partnerships, and product initiatives from OpenAI that have failed to materialize or were ultimately discontinued. This report delves into the intricate challenges faced by the leading AI research and deployment company, highlighting instances where strategic collaborations did not proceed as anticipated or where internal product development efforts were shelved. The article suggests that despite OpenAI's significant successes in advancing AI technology, such as the GPT series, the competitive landscape and inherent complexities of cutting-edge innovation lead to numerous ventures that never reach fruition. It likely explores the business implications of these unfulfilled projects, offering insights into the dynamic and often unpredictable nature of the rapidly evolving artificial intelligence industry and the operational realities behind a high-profile tech enterprise, acknowledging both its breakthroughs and its strategic missteps or unrealized ambitions.

05

Claude Wrote a Full FreeBSD Remote Kernel RCE with Root Shell (CVE-2026-4747)

A recent report highlights a significant cybersecurity development where an artificial intelligence model, specifically Claude, is credited with autonomously generating a full Remote Kernel Code Execution (RCE) exploit for the FreeBSD operating system, identified as CVE-2026-4747. This exploit reportedly grants root shell access, demonstrating a highly sophisticated capability in automated vulnerability discovery and exploit generation. The incident underscores the rapidly evolving landscape of AI in cybersecurity, posing both unprecedented opportunities for defensive measures and substantial challenges in mitigating AI-assisted threats. The ability of a large language model like Claude to independently identify, analyze, and weaponize a complex kernel-level vulnerability, ultimately achieving root privileges on a robust operating system, signifies a critical and concerning milestone in AI's offensive capabilities. This event necessitates urgent consideration of AI safety protocols, responsible AI development, and advanced defensive strategies to counteract the potential misuse of such powerful AI tools, particularly in the realm of national security and critical infrastructure protection, prompting a reevaluation of current cyber threat models.

06

Claude Code Unpacked : A visual guide

"Claude Code Unpacked: A Visual Guide" presents itself as a crucial resource designed to dissect and explain the internal mechanics of Claude, an advanced AI model, in the wake of its recent source code leak. The guide aims to provide clear, visual explanations of the complex components and functionalities exposed by the leak, which originated from a map file within Claude's NPM registry in March 2026. This security incident ignited extensive community discourse, revealing intriguing elements such as "fake tools," "frustration regexes," and an "undercover mode" embedded within the AI's architecture. The visual guide will likely offer unprecedented transparency, demystifying proprietary methods and design choices that were previously confidential. It serves as an invaluable educational tool for developers, AI researchers, and enthusiasts seeking to understand the intricate technical construction and operational behavior of a leading large language model.

huggingface

6 stories
01

CutClaw: Agentic Hours-Long Video Editing via Music Synchronization

Editing the video content with audio alignment forms a digital human-made art in current social media. However, the time-consuming and repetitive nature of manual video editing has long been a challenge for filmmakers and professional content creators alike. In this paper, we introduce CutClaw, an autonomous multi-agent framework designed to edit hours-long raw footage into meaningful short videos that leverages the capabilities of multiple Multimodal Language Models~(MLLMs) as an agent system. It produces videos with synchronized music, followed by instructions, and a visually appealing appearance. In detail, our approach begins by employing a hierarchical multimodal decomposition that captures both fine-grained details and global structures across visual and audio footage. Then, to ensure narrative consistency, a Playwriter Agent orchestrates the whole storytelling flow and structures the long-term narrative, anchoring visual scenes to musical shifts. Finally, to construct a short edited video, Editor and Reviewer Agents collaboratively optimize the final cut via selecting fine-grained visual content based on rigorous aesthetic and semantic criteria. We conduct detailed experiments to demonstrate that CutClaw significantly outperforms state-of-the-art baselines in generating high-quality, rhythm-aligned videos. The code is available at: https://github.com/GVCLab/CutClaw.

02

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen parametric knowledge, which makes them struggle with real-world image generation involving long-tail and knowledge-intensive concepts. Inspired by the broad success of agents on real-world tasks, we explore agentic modeling to address this limitation. Specifically, we present Unify-Agent, a unified multimodal agent for world-grounded image synthesis, which reframes image generation as an agentic pipeline consisting of prompt understanding, multimodal evidence searching, grounded recaptioning, and final synthesis. To train our model, we construct a tailored multimodal data pipeline and curate 143K high-quality agent trajectories for world-grounded image synthesis, enabling effective supervision over the full agentic generation process. We further introduce FactIP, a benchmark covering 12 categories of culturally significant and long-tail factual concepts that explicitly requires external knowledge grounding. Extensive experiments show that our proposed Unify-Agent substantially improves over its base unified model across diverse benchmarks and real world generation tasks, while approaching the world knowledge capabilities of the strongest closed-source models. As an early exploration of agent-based modeling for world-grounded image synthesis, our work highlights the value of tightly coupling reasoning, searching, and generation for reliable open-world agentic image synthesis.

03

Think Anywhere in Code Generation

Recent advances in reasoning Large Language Models (LLMs) have primarily relied on upfront thinking, where reasoning occurs before final answer. However, this approach suffers from critical limitations in code generation, where upfront thinking is often insufficient as problems' full complexity only reveals itself during code implementation. Moreover, it cannot adaptively allocate reasoning effort throughout the code generation process where difficulty varies significantly. In this paper, we propose Think-Anywhere, a novel reasoning mechanism that enables LLMs to invoke thinking on-demand at any token position during code generation. We achieve Think-Anywhere by first teaching LLMs to imitate the reasoning patterns through cold-start training, then leveraging outcome-based RL rewards to drive the model's autonomous exploration of when and where to invoke reasoning. Extensive experiments on four mainstream code generation benchmarks (i.e., LeetCode, LiveCodeBench, HumanEval, and MBPP) show that Think-Anywhere achieves state-of-the-art performance over both existing reasoning methods and recent post-training approaches, while demonstrating consistent generalization across diverse LLMs. Our analysis further reveals that Think-Anywhere enables the model to adaptively invoke reasoning at high-entropy positions, providing enhanced interpretability.

04

FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration

Scientific idea generation (SIG) is critical to AI-driven autonomous research, yet existing approaches are often constrained by a static retrieval-then-generation paradigm, leading to homogeneous and insufficiently divergent ideas. In this work, we propose FlowPIE, a tightly coupled retrieval-generation framework that treats literature exploration and idea generation as a co-evolving process. FlowPIE expands literature trajectories via a flow-guided Monte Carlo Tree Search (MCTS) inspired by GFlowNets, using the quality of current ideas assessed by an LLM-based generative reward model (GRM) as a supervised signal to guide adaptive retrieval and construct a diverse, high-quality initial population. Based on this population, FlowPIE models idea generation as a test-time idea evolution process, applying selection, crossover, and mutation with the isolation island paradigm and GRM-based fitness computation to incorporate cross-domain knowledge. It effectively mitigates the information cocoons arising from over-reliance on parametric knowledge and static literature. Extensive evaluations demonstrate that FlowPIE consistently produces ideas with higher novelty, feasibility and diversity compared to strong LLM-based and agent-based frameworks, while enabling reward scaling during test time.

05

VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or applying geometry-aware alignment. However, architectural modifications can compromise the generalization of internet-scale pretrained models, while existing alignment methods are limited to static scenes and rely on RGB-space rewards that require repeated VAE decoding, incurring substantial compute overhead and failing to generalize to highly dynamic real-world scenes. To preserve the pretrained capacity while improving geometric consistency, we propose VGGRPO (Visual Geometry GRPO), a latent geometry-guided framework for geometry-aware video post-training. VGGRPO introduces a Latent Geometry Model (LGM) that stitches video diffusion latents to geometry foundation models, enabling direct decoding of scene geometry from the latent space. By constructing LGM from a geometry model with 4D reconstruction capability, VGGRPO naturally extends to dynamic scenes, overcoming the static-scene limitations of prior methods. Building on this, we perform latent-space Group Relative Policy Optimization with two complementary rewards: a camera motion smoothness reward that penalizes jittery trajectories, and a geometry reprojection consistency reward that enforces cross-view geometric coherence. Experiments on both static and dynamic benchmarks show that VGGRPO improves camera stability, geometry consistency, and overall quality while eliminating costly VAE decoding, making latent-space geometry-guided reinforcement an efficient and flexible approach to world-consistent video generation.

06

GEMS: Agent-Native Multimodal Generation with Memory and Skills

Recent multimodal generation models have achieved remarkable progress on general-purpose generation tasks, yet continue to struggle with complex instructions and specialized downstream tasks. Inspired by the success of advanced agent frameworks such as Claude Code, we propose GEMS (Agent-Native Multimodal GEneration with Memory and Skills), a framework that pushes beyond the inherent limitations of foundational models on both general and downstream tasks. GEMS is built upon three core components. Agent Loop introduces a structured multi-agent framework that iteratively improves generation quality through closed-loop optimization. Agent Memory provides a persistent, trajectory-level memory that hierarchically stores both factual states and compressed experiential summaries, enabling a global view of the optimization process while reducing redundancy. Agent Skill offers an extensible collection of domain-specific expertise with on-demand loading, allowing the system to effectively handle diverse downstream applications. Across five mainstream tasks and four downstream tasks, evaluated on multiple generative backends, GEMS consistently achieves significant performance gains. Most notably, it enables the lightweight 6B model Z-Image-Turbo to surpass the state-of-the-art Nano Banana 2 on GenEval2, demonstrating the effectiveness of agent harness in extending model capabilities beyond their original limits.