NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-02-24ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Show HN: Steerling-8B, a language model that can explain any token it generates

Guidelabs.ai has officially released Steerling-8B, an innovative language model that introduces a novel capability: the ability to explain the generation of any token it produces. This development, detailed on their blog, represents a significant advancement in the quest for more transparent and interpretable artificial intelligence systems. Unlike many black-box models, Steerling-8B aims to demystify its decision-making process by providing a clear rationale for each output token. This feature holds substantial promise for researchers and developers seeking to understand, debug, and improve large language models. Enhanced interpretability is crucial for building trust, ensuring accountability, and enabling the safe and ethical deployment of AI in critical applications. The release of Steerling-8B underscores the industry's increasing focus on Explainable AI (XAI) and its practical integration into generative models, moving beyond mere performance metrics to embrace clarity and insight into model behavior.

02

HuggingFace Agent Skills

HuggingFace Agent Skills refers to a significant initiative or framework developed by Hugging Face, focused on enabling the creation and management of modular capabilities for sophisticated AI agents. This project likely provides a structured and open-source approach for defining, developing, and integrating diverse 'skills' that an AI agent can leverage to perform complex tasks and interact effectively with various environments. Given Hugging Face's profound expertise in machine learning and natural language processing, it is anticipated that these skills would encompass advanced NLP functionalities, integration with external tools and APIs, and mechanisms for task-specific automation and reasoning. The framework aims to serve as a crucial resource for developers and researchers looking to build more autonomous, adaptable, and versatile AI agents by offering a standardized and collaborative way to extend their functionalities, pushing the boundaries of what AI agents can achieve in real-world applications.

03

OpenAI resets spending expectations, from $1.4T to $600B

OpenAI has substantially revised its long-term spending projections, cutting its future expenditure target from an estimated $1.4 trillion to approximately $600 billion. This significant recalibration reflects a strategic reassessment of the resources required to achieve its ambitious goals, including the development of advanced artificial general intelligence (AGI). The updated forecast, reportedly set for completion by 2030, suggests a renewed emphasis on efficiency gains in AI research and development, potentially driven by advancements in algorithmic optimization, more efficient hardware utilization, or a more refined understanding of scaling costs. Industry observers speculate that this adjustment could indicate a shift towards smarter, rather than simply larger, investment in computational infrastructure and talent. This move might also be a response to evolving market dynamics, investor expectations for fiscal prudence, or a refinement of the technological roadmap for future AI systems. The revised outlook carries implications for the broader AI industry, potentially setting a precedent for more disciplined financial planning within highly capital-intensive sectors.

04

Show HN: Emdash – Open-source agentic development environment

Emdash is an innovative, open-source, and provider-agnostic desktop application pioneering the concept of an "Agentic Development Environment" (ADE). Developed by Arne and Raban, it enables developers to run multiple autonomous coding agents simultaneously. A key feature is the isolation of each agent within its own Git worktree, allowing for flexible deployment either locally or remotely via SSH. Emdash was conceived to address common pain points in modern software development, such as the complexity of managing numerous terminals and branches, and the inefficiencies associated with waiting for outputs from AI coding assistants like Codex. By putting the terminal at the core of its interface, Emdash aims to provide a streamlined, highly organized, and efficient platform for leveraging AI agents to automate and accelerate coding tasks, ultimately improving developer productivity.

05

OpenAI, the US government and Persona built an identity surveillance machine

A recent investigation brings to light allegations that OpenAI, in conjunction with the US government and the identity verification platform Persona, has contributed to the creation of an advanced 'identity surveillance machine.' The report critically examines the collaboration, suggesting that the integration of OpenAI's cutting-edge artificial intelligence with Persona's robust identity verification technology, under government involvement, could establish a comprehensive system for monitoring and tracking individuals. This development sparks considerable debate on the ethical ramifications of deploying powerful AI systems in sensitive areas like national security and public identity management. Concerns are expected to center on potential privacy infringements, the expansion of government surveillance capabilities, and the broader implications for civil liberties. The article emphasizes the urgent need for transparency, accountability, and stringent regulatory frameworks to govern AI development and its application, particularly in partnerships involving state entities and private companies, to mitigate risks associated with surveillance and data exploitation.

06

So You Want to Cure Your Own Disease – Using AI to Take Agency over Your Health

This article explores the burgeoning concept of leveraging artificial intelligence to empower individuals in managing their own health and potentially addressing diseases, advocating for greater personal agency over healthcare decisions. It delves into how AI technologies can facilitate access to vast medical knowledge, offer personalized insights based on individual data, and support proactive health management strategies. The narrative emphasizes a shift from traditional healthcare reliance towards self-directed health initiatives, enabled by advanced AI tools for diagnostics, treatment exploration, and wellness optimization. While promoting the transformative potential of AI in fostering self-care and customized health pathways, the discussion implicitly acknowledges the critical need for responsible application and addresses the inherent complexities and ethical considerations associated with individuals utilizing AI for self-diagnosis and treatment, highlighting both the opportunities and challenges in this evolving domain of personal health empowerment.

huggingface

6 stories
01

Agents of Chaos

We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under benign and adversarial conditions. Focusing on failures emerging from the integration of language models with autonomy, tool use, and multi-party communication, we document eleven representative case studies. Observed behaviors include unauthorized compliance with non-owners, disclosure of sensitive information, execution of destructive system-level actions, denial-of-service conditions, uncontrolled resource consumption, identity spoofing vulnerabilities, cross-agent propagation of unsafe practices, and partial system takeover. In several cases, agents reported task completion while the underlying system state contradicted those reports. We also report on some of the failed attempts. Our findings establish the existence of security-, privacy-, and governance-relevant vulnerabilities in realistic deployment settings. These behaviors raise unresolved questions regarding accountability, delegated authority, and responsibility for downstream harms, and warrant urgent attention from legal scholars, policymakers, and researchers across disciplines. This report serves as an initial empirical contribution to that broader conversation.

02

Anatomy of Agentic Memory: Taxonomy and Empirical Analysis of Evaluation and System Limitations

Agentic memory systems enable large language model (LLM) agents to maintain state across long interactions, supporting long-horizon reasoning and personalization beyond fixed context windows. Despite rapid architectural development, the empirical foundations of these systems remain fragile: existing benchmarks are often underscaled, evaluation metrics are misaligned with semantic utility, performance varies significantly across backbone models, and system-level costs are frequently overlooked. This survey presents a structured analysis of agentic memory from both architectural and system perspectives. We first introduce a concise taxonomy of MAG systems based on four memory structures. Then, we analyze key pain points limiting current systems, including benchmark saturation effects, metric validity and judge sensitivity, backbone-dependent accuracy, and the latency and throughput overhead introduced by memory maintenance. By connecting the memory structure to empirical limitations, this survey clarifies why current agentic memory systems often underperform their theoretical promise and outlines directions for more reliable evaluation and scalable system design.

03

SkillOrchestra: Learning to Route Agents via Skill Transfer

Compound AI systems promise capabilities beyond those of individual models, yet their success depends critically on effective orchestration. Existing routing approaches face two limitations: (1) input-level routers make coarse query-level decisions that ignore evolving task requirements; (2) RL-trained orchestrators are expensive to adapt and often suffer from routing collapse, repeatedly invoking one strong but costly option in multi-turn scenarios. We introduce SkillOrchestra, a framework for skill-aware orchestration. Instead of directly learning a routing policy end-to-end, SkillOrchestra learns fine-grained skills from execution experience and models agent-specific competence and cost under those skills. At deployment, the orchestrator infers the skill demands of the current interaction and selects agents that best satisfy them under an explicit performance-cost trade-off. Extensive experiments across ten benchmarks demonstrate that SkillOrchestra outperforms SoTA RL-based orchestrators by up to 22.5% with 700x and 300x learning cost reduction compared to Router-R1 and ToolOrchestra, respectively. These results show that explicit skill modeling enables scalable, interpretable, and sample-efficient orchestration, offering a principled alternative to data-intensive RL-based approaches. The code is available at: https://github.com/jiayuww/SkillOrchestra.

04

VLANeXt: Recipes for Building Strong VLA Models

Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT-2 and OpenVLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. VLANeXt outperforms prior state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks and demonstrates strong generalization in real-world experiments. We will release a unified, easy-to-use codebase that serves as a common platform for the community to reproduce our findings, explore the design space, and build new VLA variants on top of a shared foundation.

05

Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact vision-language-diffusion model that brings unified multimodal intelligence to a mobile device. Its core module, the Mobile Conditioning Projector (MCP), fuses vision-language features with a diffusion generator using depthwise-separable convolutions and layerwise alignment. This design enables efficient cross-modal conditioning with minimal computational cost. Trained on only a few million samples and post-trained in a novel quadruplet format (generation prompt, image, question, answer), Mobile-O jointly enhances both visual understanding and generation capabilities. Despite its efficiency, Mobile-O attains competitive or superior performance compared to other unified models, achieving 74% on GenEval and outperforming Show-O and JanusFlow by 5% and 11%, while running 6x and 11x faster, respectively. For visual understanding, Mobile-O surpasses them by 15.3% and 5.1% averaged across seven benchmarks. Running in only ~3s per 512x512 image on an iPhone, Mobile-O establishes the first practical framework for real-time unified multimodal understanding and generation on edge devices. We hope Mobile-O will ease future research in real-time unified multimodal intelligence running entirely on-device with no cloud dependency. Our code, models, datasets, and mobile application are publicly available at https://amshaker.github.io/Mobile-O/

06

DSDR: Dual-Scale Diversity Regularization for Exploration in LLM Reasoning

Reinforcement learning with verifiers (RLVR) is a central paradigm for improving large language model (LLM) reasoning, yet existing methods often suffer from limited exploration. Policies tend to collapse onto a few reasoning patterns and prematurely stop deep exploration, while conventional entropy regularization introduces only local stochasticity and fails to induce meaningful path-level diversity, leading to weak and unstable learning signals in group-based policy optimization. We propose DSDR, a Dual-Scale Diversity Regularization reinforcement learning framework that decomposes diversity in LLM reasoning into global and coupling components. Globally, DSDR promotes diversity among correct reasoning trajectories to explore distinct solution modes. Locally, it applies a length-invariant, token-level entropy regularization restricted to correct trajectories, preventing entropy collapse within each mode while preserving correctness. The two scales are coupled through a global-to-local allocation mechanism that emphasizes local regularization for more distinctive correct trajectories. We provide theoretical support showing that DSDR preserves optimal correctness under bounded regularization, sustains informative learning signals in group-based optimization, and yields a principled global-to-local coupling rule. Experiments on multiple reasoning benchmarks demonstrate consistent improvements in accuracy and pass@k, highlighting the importance of dual-scale diversity for deep exploration in RLVR. Code is available at https://github.com/SUSTechBruce/DSDR.