NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-10中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Yann LeCun raises $1B to build AI that understands the physical world

Yann LeCun, a pioneering figure in the field of artificial intelligence and a leading researcher at Meta AI, has reportedly raised a substantial $1 billion to fund ambitious research aimed at developing AI systems with a profound understanding of the physical world. This significant capital injection underscores a strategic pivot towards addressing some of the most challenging frontiers in AI, moving beyond the current focus on data-driven patterns in digital spaces. The initiative seeks to build AI capable of comprehending real-world physics, interacting with objects, and navigating complex environments with common-sense reasoning. Such advancements are critical for fostering autonomous agents and robotics that can robustly operate and learn in dynamic physical settings, unlike present systems often limited to narrow tasks. This funding represents a major commitment to advancing foundational AI research, focusing on embodied intelligence, causality, and a more holistic grasp of our physical surroundings, ultimately aiming for more general and adaptable AI capabilities.

02

Launch HN: RunAnywhere (YC W26) Faster AI Inference on Apple Silicon

RunAnywhere, founded by Sanchit and Shubham (YC W26), has launched MetalRT, a high-performance inference engine specifically designed for Apple Silicon. This engine significantly accelerates AI tasks such as Large Language Models (LLMs), speech-to-text, and text-to-speech, outperforming existing solutions like llama.cpp, Apple's MLX, Ollama, and sherpa-onnx across various modalities. MetalRT achieves its superior speed through custom Metal shaders and by eliminating traditional framework overhead. Alongside MetalRT, the company has open-sourced RCLI, an end-to-end voice AI pipeline that operates entirely on-device, offering a fully local solution without reliance on cloud services or API keys. RCLI provides a fast, interactive experience, enabling users to set up and run voice AI models directly on their Apple devices with simple command-line tools. Benchmarking data confirms its efficiency, highlighting substantial performance gains for AI inference on Apple's M-series processors.

03

Redox OS has adopted a Certificate of Origin policy and a strict no-LLM policy

Redox OS, an experimental microkernel-based operating system, has officially implemented a Certificate of Origin (CoC) policy alongside a strict prohibition against contributions generated by Large Language Models (LLMs). The Certificate of Origin policy mandates that all contributors attest to the originality and their legal right to submit code under the project's specified license, thereby bolstering intellectual property protection and ensuring clear provenance for the codebase. Concurrently, the newly adopted no-LLM policy explicitly forbids the integration of any code produced by artificial intelligence models, such as ChatGPT or similar generative tools, into the Redox OS repository. This decision highlights a proactive stance by the project maintainers to address burgeoning concerns within the open-source ecosystem regarding the ethical, legal, and quality control challenges posed by AI-generated content. These challenges include potential licensing conflicts, difficulties in verifying code authorship, and the introduction of unvetted or insecure code patterns. By establishing these stringent guidelines, Redox OS reinforces its commitment to high standards of code integrity, transparent development, and robust intellectual property management, setting a distinct precedent in the evolving landscape of software development.

04

Maybe the G in AGI stands for Gemini

This discussion speculates on the potential and implications of Google's Gemini model within the broader pursuit of Artificial General Intelligence (AGI). The provocative title, "Maybe the G in AGI stands for Gemini," suggests an exploration of whether the advanced capabilities demonstrated by Google's multimodal AI model position it as a significant step towards achieving AGI, or if it redefines the very essence of what 'general intelligence' entails in an artificial context. The discourse likely delves into Gemini's capacity for complex reasoning, multimodal understanding, and problem-solving, examining if these attributes bridge the gap towards systems that can understand, learn, and apply intelligence across a wide range of tasks at a human-like level. It prompts a critical evaluation of current AI progress relative to the ambitious goals of AGI.

05

Amazon wins court order to block Perplexity's AI shopping agent

Amazon has successfully secured a court order to prevent Perplexity AI from operating its new AI-powered shopping agent. This legal action marks a significant development in the competitive landscape of e-commerce and artificial intelligence, as established market leaders begin to leverage legal frameworks against innovative AI applications that potentially disrupt their business models. Perplexity's AI shopping agent, designed to streamline product discovery and purchasing across various platforms, likely posed a direct competitive threat to Amazon's retail dominance by offering an aggregated and intelligent alternative to its native shopping experience. The court's decision underscores growing concerns over data access, intellectual property rights, and fair competition in the rapidly evolving AI sector. This case sets a precedent for how traditional retail giants might protect their market share from new AI agents that aggregate information or facilitate transactions outside their direct control. The ruling could influence future development and deployment strategies for AI agents operating in commercial spaces, prompting a reevaluation of operational methods to avoid legal challenges.

06

LoGeR  3D reconstruction from extremely long videos (DeepMind, UC Berkeley)

DeepMind and UC Berkeley have introduced LoGeR, a novel framework designed for high-quality 3D reconstruction from exceptionally long video sequences. Traditional 3D reconstruction methods often struggle with the scale and temporal inconsistencies inherent in extended video data, leading to computational bottlenecks and accumulated errors. LoGeR addresses these challenges by employing advanced techniques that manage long-term temporal dependencies and maintain geometric consistency over extended periods. This research aims to significantly improve the accuracy and efficiency of generating detailed 3D models from diverse video sources, paving the way for advancements in areas such as virtual reality, robotics, and autonomous systems where robust environmental mapping from continuous visual input is crucial. The project highlights a significant step forward in robustly processing large-scale visual data for sophisticated spatial understanding.

huggingface

6 stories
01

LiveWorld: Simulating Out-of-Sight Dynamics in Generative Video World Models

Recent generative video world models aim to simulate visual environment evolution, allowing an observer to interactively explore the scene via camera control. However, they implicitly assume that the world only evolves within the observer's field of view. Once an object leaves the observer's view, its state is "frozen" in memory, and revisiting the same region later often fails to reflect events that should have occurred in the meantime. In this work, we identify and formalize this overlooked limitation as the "out-of-sight dynamics" problem, which impedes video world models from representing a continuously evolving world. To address this issue, we propose LiveWorld, a novel framework that extends video world models to support persistent world evolution. Instead of treating the world as static observational memory, LiveWorld models a persistent global state composed of a static 3D background and dynamic entities that continue evolving even when unobserved. To maintain these unseen dynamics, LiveWorld introduces a monitor-based mechanism that autonomously simulates the temporal progression of active entities and synchronizes their evolved states upon revisiting, ensuring spatially coherent rendering. For evaluation, we further introduce LiveBench, a dedicated benchmark for the task of maintaining out-of-sight dynamics. Extensive experiments show that LiveWorld enables persistent event evolution and long-term scene consistency, bridging the gap between existing 2D observation-based memory and true 4D dynamic world simulation. The baseline and benchmark will be publicly available at https://zichengduan.github.io/LiveWorld/index.html.

02

Agentic Critical Training

Training large language models (LLMs) as autonomous agents often begins with imitation learning, but it only teaches agents what to do without understanding why: agents never contrast successful actions against suboptimal alternatives and thus lack awareness of action quality. Recent approaches attempt to address this by introducing self-reflection supervision derived from contrasts between expert and alternative actions. However, the training paradigm fundamentally remains imitation learning: the model imitates pre-constructed reflection text rather than learning to reason autonomously. We propose Agentic Critical Training (ACT), a reinforcement learning paradigm that trains agents to identify the better action among alternatives. By rewarding whether the model's judgment is correct, ACT drives the model to autonomously develop reasoning about action quality, producing genuine self-reflection rather than imitating it. Across three challenging agent benchmarks, ACT consistently improves agent performance when combined with different post-training methods. It achieves an average improvement of 5.07 points over imitation learning and 4.62 points over reinforcement learning. Compared to approaches that inject reflection capability through knowledge distillation, ACT also demonstrates clear advantages, yielding an average improvement of 2.42 points. Moreover, ACT enables strong out-of-distribution generalization on agentic benchmarks and improves performance on general reasoning benchmarks without any reasoning-specific training data, highlighting the value of our method. These results suggest that ACT is a promising path toward developing more reflective and capable LLM agents.

03

OneMillion-Bench: How Far are Language Agents from Human Experts?

As language models (LMs) evolve from chat assistants to long-horizon agents capable of multi-step reasoning and tool use, existing benchmarks remain largely confined to structured or exam-style tasks that fall short of real-world professional demands. To this end, we introduce OneMillion-Bench, a benchmark of 400 expert-curated tasks spanning Law, Finance, Industry, Healthcare, and Natural Science, built to evaluate agents across economically consequential scenarios. Unlike prior work, the benchmark requires retrieving authoritative sources, resolving conflicting evidence, applying domain-specific rules, and making constraint decisions, where correctness depends as much on the reasoning process as the final answer. We adopt a rubric-based evaluation protocol scoring factual accuracy, logical coherence, practical feasibility, and professional compliance, focused on expert-level problems to ensure meaningful differentiation across agents. Together, OneMillion-Bench provides a unified testbed for assessing agentic reliability, professional depth, and practical readiness in domain-intensive scenarios.

04

Sparse-BitNet: 1.58-bit LLMs are Naturally Friendly to Semi-Structured Sparsity

Semi-structured N:M sparsity and low-bit quantization (e.g., 1.58-bit BitNet) are two promising approaches for improving the efficiency of large language models (LLMs), yet they have largely been studied in isolation. In this work, we investigate their interaction and show that 1.58-bit BitNet is naturally more compatible with N:M sparsity than full-precision models. To study this effect, we propose Sparse-BitNet, a unified framework that jointly applies 1.58-bit quantization and dynamic N:M sparsification while ensuring stable training for the first time. Across multiple model scales and training regimes (sparse pretraining and dense-to-sparse schedules), 1.58-bit BitNet consistently exhibits smaller performance degradation than full-precision baselines at the same sparsity levels and can tolerate higher structured sparsity before accuracy collapse. Moreover, using our custom sparse tensor core, Sparse-BitNet achieves substantial speedups in both training and inference, reaching up to 1.30X. These results highlight that combining extremely low-bit quantization with semi-structured N:M sparsity is a promising direction for efficient LLMs. Code available at https://github.com/AAzdi/Sparse-BitNet

05

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) reasoning. However, existing CoT-based T2I methods largely rely on abstract natural-language planning, which lacks the precision required for complex spatial layouts, structured visual elements, and dense textual content. In this work, we propose CoCo (Code-as-CoT), a code-driven reasoning framework that represents the reasoning process as executable code, enabling explicit and verifiable intermediate planning for image generation. Given a text prompt, CoCo first generates executable code that specifies the structural layout of the scene, which is then executed in a sandboxed environment to render a deterministic draft image. The model subsequently refines this draft through fine-grained image editing to produce the final high-fidelity result. To support this training paradigm, we construct CoCo-10K, a curated dataset containing structured draft-final image pairs designed to teach both structured draft construction and corrective visual refinement. Empirical evaluations on StructT2IBench, OneIG-Bench, and LongText-Bench show that CoCo achieves improvements of +68.83%, +54.8%, and +41.23% over direct generation, while also outperforming other generation methods empowered by CoT. These results demonstrate that executable code is an effective and reliable reasoning paradigm for precise, controllable, and structured text-to-image generation. The code is available at: https://github.com/micky-li-hd/CoCo

06

Breaking Training Bottlenecks: Effective and Stable Reinforcement Learning for Coding Models

Modern code generation models exhibit longer outputs, accelerated capability growth, and changed training dynamics, rendering traditional training methodologies, algorithms, and datasets ineffective for improving their performance. To address these training bottlenecks, we propose MicroCoder-GRPO, an improved Group Relative Policy Optimization approach with three innovations: conditional truncation masking to improve long output potential while maintaining training stability, diversity-determined temperature selection to maintain and encourage output diversity, and removal of KL loss with high clipping ratios to facilitate solution diversity. MicroCoder-GRPO achieves up to 17.6% relative improvement over strong baselines on LiveCodeBench v6, with more pronounced gains under extended context evaluation. Additionally, we release MicroCoder-Dataset, a more challenging training corpus that achieves 3x larger performance gains than mainstream datasets on LiveCodeBench v6 within 300 training steps, and MicroCoder-Evaluator, a robust framework with approximately 25% improved evaluation accuracy and around 40% faster execution. Through comprehensive analysis across more than thirty controlled experiments, we reveal 34 training insights across seven main aspects, demonstrating that properly trained models can achieve competitive performance with larger counterparts.