NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-01-29ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Project Genie: Experimenting with infinite, interactive worlds

Google DeepMind's Project Genie introduces an innovative research initiative focused on creating AI models capable of generating "infinite, interactive worlds." This project aims to explore the frontiers of generative AI by developing systems that can build virtual environments with dynamic elements and interactive capabilities, moving beyond static content generation. The core idea is to train AI to learn from a wide array of visual data and human actions, enabling it to synthesize novel and consistent virtual spaces where users or other AI agents can engage. This research holds significant implications for various applications, including advanced game development, realistic simulations for training AI, and novel forms of creative expression. Project Genie represents a step towards truly immersive and autonomously evolving digital realities, pushing the boundaries of AI-powered world creation and interaction. The experimentation explores how AI can master the complexities of environmental design and interactive narrative within a dynamically expanding digital canvas.

02

Claude Code Daily Benchmarks for Degradation Tracking

MarginLab has implemented a robust system for 'Claude Code Daily Benchmarks for Degradation Tracking,' an essential initiative aimed at continuously monitoring the performance of Claude's code generation capabilities. This system is designed to execute a comprehensive suite of coding benchmarks on a daily basis, meticulously tracking key metrics to identify any potential degradation in the model's output quality or efficiency. The primary objective is to ensure the consistent high performance of Claude, a prominent large language model, particularly in complex programming tasks where precision and reliability are critical. By establishing daily performance baselines and actively comparing subsequent results, the platform can promptly detect and alert developers to any regressions caused by model updates, infrastructure changes, or other factors. This proactive degradation tracking is indispensable for maintaining the integrity and utility of advanced AI models in production environments, facilitating rapid intervention and corrective actions to safeguard the user experience and the overall quality of AI-generated code.

03

Launch HN: AgentMail (YC S25) – An API that gives agents their own email inboxes

AgentMail (YC S25), founded by Haakam, Michael, and Adi, has launched an API specifically designed to provide AI agents with their own dedicated email inboxes. The founders clarify that their innovation focuses on providing email capabilities *for* AI, rather than applying AI *to* email. Email is presented as an ideal interface for long-running autonomous agents due to its multithreaded and asynchronous nature, comprehensive support for rich text and file attachments, and its status as a universal protocol with built-in identity and authentication. A key advantage is that significant workflow context already exists within email ecosystems. The impetus for AgentMail arose from the limitations of existing email APIs, such as Gmail's, which restrict programmatic inbox creation and impose prohibitive rate limits, thereby impeding the development of truly independent agents. AgentMail aims to empower agents to autonomously receive and complete tasks, and to communicate proactively via email when human intervention is necessary, all without requiring users to delegate their personal identity.

04

US cybersecurity chief leaked sensitive government files to ChatGPT: Report

A recent report has brought to light an alleged incident where a US cybersecurity chief is accused of leaking sensitive government files by inputting them into ChatGPT. This revelation immediately triggers profound concerns regarding national security, data confidentiality, and the ethical frameworks governing the use of advanced artificial intelligence technologies within critical governmental infrastructure. The incident underscores the inherent risks associated with integrating public-facing large language models (LLMs) into environments handling classified or highly sensitive information, potentially exposing vulnerabilities that could be exploited by malicious actors or lead to unintended data compromise. Experts are now emphasizing the urgent need for comprehensive policy revisions, enhanced employee education on secure AI interaction protocols, and the potential development of specialized, secure AI solutions tailored for government use to mitigate such risks. This event serves as a critical case study, highlighting the imperative for robust cybersecurity measures and stringent guidelines to prevent inadvertent information disclosures when leveraging powerful AI systems in sensitive operational contexts.

05

SpaceX in Merger Talks with xAI

Reports indicate that Elon Musk's aerospace and satellite internet company, SpaceX, is engaged in advanced merger discussions with xAI, his artificial intelligence venture. This potential consolidation, which sources suggest is being pursued in anticipation of a planned initial public offering (IPO), signifies a strategic move to integrate or align the ambitious technological endeavors of both entities. xAI, known for its focus on developing large language models and advanced AI capabilities, could potentially benefit from SpaceX's vast infrastructure, data resources, or computational power, while SpaceX might seek to leverage xAI's AI advancements for various applications within its space exploration, satellite internet, or other technological initiatives. The proposed merger underscores a growing trend of cross-sector integration among technology giants, particularly those helmed by visionary leaders like Musk, aiming to create synergistic value and accelerate innovation across diverse domains. Observers are closely watching how this potential merger might reshape the competitive landscape in both the AI and aerospace sectors, particularly given Musk's significant influence in both fields and the ambitious goals of his companies.

06

Waymo robotaxi hits a child near an elementary school in Santa Monica

A Waymo robotaxi was involved in an incident where it struck a child near an elementary school in Santa Monica, immediately drawing significant attention to the safety and operational reliability of autonomous vehicles. This event is poised to intensify public and regulatory scrutiny on Waymo, a prominent player in the self-driving sector, and could have substantial implications for the broader autonomous vehicle industry's ongoing expansion and acceptance. The incident highlights the inherent complexities and unique challenges associated with deploying AI-driven transportation systems in dense urban environments, particularly those with vulnerable road users like children. It underscores the critical need for continuous advancements in AI perception and decision-making capabilities, robust safety protocols, and transparent incident reporting mechanisms. The accident reignites discussions surrounding the readiness of self-driving technology for widespread commercial deployment and the rigorous testing required to ensure absolute safety, as autonomous vehicle companies navigate the intricate balance between technological innovation and public trust.

huggingface

6 stories
01

Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning

Reinforcement learning has empowered large language models to act as intelligent agents, yet training them for long-horizon tasks remains challenging due to the scarcity of high-quality trajectories, especially under limited resources. Existing methods typically scale up rollout sizes and indiscriminately allocate computational resources among intermediate steps. Such attempts inherently waste substantial computation budget on trivial steps while failing to guarantee sample quality. To address this, we propose Spark (Strategic Policy-Aware exploRation via Key-state dynamic branching), a novel framework that selectively branches at critical decision states for resource-efficient exploration. Our key insight is to activate adaptive branching exploration at critical decision points to probe promising trajectories, thereby achieving precise resource allocation that prioritizes sampling quality over blind coverage. This design leverages the agent's intrinsic decision-making signals to reduce dependence on human priors, enabling the agent to autonomously expand exploration and achieve stronger generalization. Experiments across diverse tasks (e.g., embodied planning), demonstrate that Spark achieves superior success rates with significantly fewer training samples, exhibiting robust generalization even in unseen scenarios.

02

OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task Execution

Graphical User Interface (GUI) agents show great potential for enabling foundation models to complete real-world tasks, revolutionizing human-computer interaction and improving human productivity. In this report, we present OmegaUse, a general-purpose GUI agent model for autonomous task execution on both mobile and desktop platforms, supporting computer-use and phone-use scenarios. Building an effective GUI agent model relies on two factors: (1) high-quality data and (2) effective training methods. To address these, we introduce a carefully engineered data-construction pipeline and a decoupled training paradigm. For data construction, we leverage rigorously curated open-source datasets and introduce a novel automated synthesis framework that integrates bottom-up autonomous exploration with top-down taxonomy-guided generation to create high-fidelity synthetic data. For training, to better leverage these data, we adopt a two-stage strategy: Supervised Fine-Tuning (SFT) to establish fundamental interaction syntax, followed by Group Relative Policy Optimization (GRPO) to improve spatial grounding and sequential planning. To balance computational efficiency with agentic reasoning capacity, OmegaUse is built on a Mixture-of-Experts (MoE) backbone. To evaluate cross-terminal capabilities in an offline setting, we introduce OS-Nav, a benchmark suite spanning multiple operating systems: ChiM-Nav, targeting Chinese Android mobile environments, and Ubu-Nav, focusing on routine desktop interactions on Ubuntu. Extensive experiments show that OmegaUse is highly competitive across established GUI benchmarks, achieving a state-of-the-art (SOTA) score of 96.3% on ScreenSpot-V2 and a leading 79.1% step success rate on AndroidControl. OmegaUse also performs strongly on OS-Nav, reaching 74.24% step success on ChiM-Nav and 55.9% average success on Ubu-Nav.

03

Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent performance on general vision tasks. Contrary to the trend of relying on massive domain-specific pretraining and opaque pipelines, our work demonstrates that principled training design and transparent methodology can yield strong scientific intelligence with substantially reduced data requirements. (i) First, we provide a fully transparent, end-to-end reproducible training pipeline, covering data collection, cleaning, preprocessing, supervised fine-tuning, reinforcement learning, and evaluation, along with detailed optimization recipes. This facilitates systematic extension by the community. (ii) Second, Innovator-VL exhibits remarkable data efficiency, achieving competitive performance on various scientific tasks using fewer than five million curated samples without large-scale pretraining. These results highlight that effective reasoning can be achieved through principled data selection rather than indiscriminate scaling. (iii) Third, Innovator-VL demonstrates strong generalization, achieving competitive performance on general vision, multimodal reasoning, and scientific benchmarks. This indicates that scientific alignment can be integrated into a unified model without compromising general-purpose capabilities. Our practices suggest that efficient, reproducible, and high-performing scientific multimodal models can be built even without large-scale data, providing a practical foundation for future research.

04

VERGE: Formal Refinement and Guidance Engine for Verifiable LLM Reasoning

Despite the syntactic fluency of Large Language Models (LLMs), ensuring their logical correctness in high-stakes domains remains a fundamental challenge. We present a neurosymbolic framework that combines LLMs with SMT solvers to produce verification-guided answers through iterative refinement. Our approach decomposes LLM outputs into atomic claims, autoformalizes them into first-order logic, and verifies their logical consistency using automated theorem proving. We introduce three key innovations: (1) multi-model consensus via formal semantic equivalence checking to ensure logic-level alignment between candidates, eliminating the syntactic bias of surface-form metrics, (2) semantic routing that directs different claim types to appropriate verification strategies: symbolic solvers for logical claims and LLM ensembles for commonsense reasoning, and (3) precise logical error localization via Minimal Correction Subsets (MCS), which pinpoint the exact subset of claims to revise, transforming binary failure signals into actionable feedback. Our framework classifies claims by their logical status and aggregates multiple verification signals into a unified score with variance-based penalty. The system iteratively refines answers using structured feedback until acceptance criteria are met or convergence is achieved. This hybrid approach delivers formal guarantees where possible and consensus verification elsewhere, advancing trustworthy AI. With the GPT-OSS-120B model, VERGE demonstrates an average performance uplift of 18.7% at convergence across a set of reasoning benchmarks compared to single-pass approaches.

05

Advancing Open-source World Models

We present LingBot-World, an open-sourced world simulator stemming from video generation. Positioned as a top-tier world model, LingBot-World offers the following features. (1) It maintains high fidelity and robust dynamics in a broad spectrum of environments, including realism, scientific contexts, cartoon styles, and beyond. (2) It enables a minute-level horizon while preserving contextual consistency over time, which is also known as "long-term memory". (3) It supports real-time interactivity, achieving a latency of under 1 second when producing 16 frames per second. We provide public access to the code and model in an effort to narrow the divide between open-source and closed-source technologies. We believe our release will empower the community with practical applications across areas like content creation, gaming, and robot learning.

06

Reinforcement Learning via Self-Distillation

Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with verifiable rewards (RLVR) learn only from a scalar outcome reward per attempt, creating a severe credit-assignment bottleneck. Many verifiable environments actually provide rich textual feedback, such as runtime errors or judge evaluations, that explain why an attempt failed. We formalize this setting as reinforcement learning with rich feedback and introduce Self-Distillation Policy Optimization (SDPO), which converts tokenized feedback into a dense learning signal without any external teacher or explicit reward model. SDPO treats the current model conditioned on feedback as a self-teacher and distills its feedback-informed next-token predictions back into the policy. In this way, SDPO leverages the model's ability to retrospectively identify its own mistakes in-context. Across scientific reasoning, tool use, and competitive programming on LiveCodeBench v6, SDPO improves sample efficiency and final accuracy over strong RLVR baselines. Notably, SDPO also outperforms baselines in standard RLVR environments that only return scalar feedback by using successful rollouts as implicit feedback for failed attempts. Finally, applying SDPO to individual questions at test time accelerates discovery on difficult binary-reward tasks, achieving the same discovery probability as best-of-k sampling or multi-turn conversations with 3x fewer attempts.