NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-11中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

BitNet: 100B Param 1-Bit model for local CPUs

Microsoft's BitNet introduces a significant advancement in efficient deep learning with its 100-billion-parameter 1-bit model, specifically engineered for deployment on local CPUs. This innovation addresses the growing challenge of running increasingly larger neural networks on resource-constrained hardware, democratizing access to advanced AI capabilities. By drastically reducing the precision of model weights to a single bit, BitNet aims to minimize memory footprint and computational requirements, making sophisticated models viable for edge devices and standard personal computers. This breakthrough could enable developers and users to leverage powerful AI models without relying on extensive cloud infrastructure or specialized accelerators like GPUs. The focus on local CPU inference signifies a move towards more accessible and private AI applications, opening new avenues for on-device machine learning and reduced operational costs. BitNet represents a crucial step towards making large-scale AI ubiquitous and efficient.

02

TADA: Speech generation through text-acoustic synchronization

Hume AI has introduced TADA, an innovative open-source model designed for high-quality speech generation, leveraging a novel approach called text-acoustic synchronization. This method focuses on precisely aligning textual input with acoustic properties, enabling the synthesis of highly natural and expressive speech. Unlike traditional text-to-speech systems that might rely on discrete linguistic features, TADA emphasizes a continuous and synchronized understanding between text and the corresponding sound characteristics. This advancement promises to improve the fidelity and emotional nuances of generated speech, making it more robust and versatile for various applications. The release as an open-source project signifies a commitment to collaborative research and development in the field of generative audio, allowing researchers and developers to further explore and integrate this technology.

03

Launch HN: Prism (YC X25) – Workspace and API to generate and edit videos

Prism, an AI video creation platform and API, has launched to address the complexity of modern AI video production. Developed by Rajit, Land, and Alex, Prism aims to streamline the process of generating and editing videos, which traditionally involves integrating numerous disparate tools. The platform offers a unified workspace and an API, enabling users to efficiently remix existing videos and automate the creation of new content. Demonstrations highlight its capability to transform any video and to generate user-generated content (UGC) style advertisements through integration with tools like Openclaw. This initiative simplifies advanced video production, making AI-driven video creation more accessible and less resource-intensive for creators and businesses alike, consolidating complex workflows into a single solution.

04

Personal Computer by Perplexity

Perplexity AI has announced a new initiative titled "Personal Computer by Perplexity," hinting at an ambitious project to redefine human-computer interaction through advanced artificial intelligence. While specific details remain scarce, the project's title suggests a comprehensive AI agent or operating system designed to offer a personalized and intuitive computing experience. This venture is expected to leverage Perplexity's expertise in natural language processing and information retrieval to create a digital environment where users can interact with their devices and access information more seamlessly. The focus is likely on developing an AI that can understand complex queries, automate tasks, and provide proactive assistance, effectively transforming the traditional personal computer into an intelligent, conversational entity. This move signifies Perplexity's potential expansion beyond its current search and answer engine capabilities into a broader ecosystem of AI-driven personal computing solutions.

05

Show HN: Open-source browser for AI agents

An open-source browser, `agent-browser-protocol` (ABP), has been developed by forking Chromium to address critical synchronization issues faced by AI agents interacting with web environments. The creator observed that most agent failures stem from reasoning based on stale browser states, rather than a misunderstanding of page content. ABP tackles this by ensuring the acting agent is synchronized with the browser at every step. Following each agent action, the protocol freezes JavaScript execution and rendering, captures the current page state, and compiles notable events like navigation changes, file pickers, permission prompts, alerts, and downloads. This detailed information, alongside a screenshot of the frozen page, is then transmitted back to the agent, fostering a multimodal chat-loop interaction model that aligns more effectively with the operational paradigm of Large Language Models (LLMs).

06

Don't post generated/AI-edited comments. HN is for conversation between humans.

Hacker News has explicitly updated its community guidelines to prohibit the submission of comments that are either generated or substantially edited by artificial intelligence. This policy decision reinforces the platform's core principle of fostering genuine human-to-human conversation and maintaining the authenticity of its online discussions. The directive aims to ensure that all contributions reflect original human intellect, perspective, and engagement, rather than automated or machine-assisted outputs. This move highlights a growing industry trend among online forums to address the ethical implications and impact of generative AI on user interaction and the integrity of shared discourse. By mandating human-originated content, Hacker News seeks to preserve the unique value of human interaction and maintain a high standard of genuine community participation.

huggingface

6 stories
01

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion

While recent multimodal large language models (MLLMs) have made impressive strides, they predominantly employ a conventional autoregressive architecture as their backbone, leaving significant room to explore effective and efficient alternatives in architectural design. Concurrently, recent studies have successfully applied discrete diffusion models to various domains, such as visual understanding and image generation, revealing their considerable potential as a promising backbone for multimodal systems. Drawing inspiration from these pioneering research, we introduce Omni-Diffusion, the first any-to-any multimodal language model built entirely on mask-based discrete diffusion models, which unifies understanding and generation across text, speech, and images. Omni-Diffusion employs a unified mask-based discrete diffusion model to directly capture the joint distribution over discrete multimodal tokens. This approach supports not only bimodal tasks but also more complex scenarios involving multiple modalities. On a diverse set of benchmarks, our method outperforms or performs on par with existing multimodal systems that process two or more modalities, highlighting the significant promise of diffusion models in powering the next generation of multimodal foundation models. Project webpage: https://omni-diffusion.github.io.

02

MM-Zero: Self-Evolving Multi-Model Vision Language Models From Zero Data

Self-evolving has emerged as a key paradigm for improving foundational models such as Large Language Models (LLMs) and Vision Language Models (VLMs) with minimal human intervention. While recent approaches have demonstrated that LLM agents can self-evolve from scratch with little to no data, VLMs introduce an additional visual modality that typically requires at least some seed data, such as images, to bootstrap the self-evolution process. In this work, we present Multi-model Multimodal Zero (MM-Zero), the first RL-based framework to achieve zero-data self-evolution for VLM reasoning. Moving beyond prior dual-role (Proposer and Solver) setups, MM-Zero introduces a multi-role self-evolving training framework comprising three specialized roles: a Proposer that generates abstract visual concepts and formulates questions; a Coder that translates these concepts into executable code (e.g., Python, SVG) to render visual images; and a Solver that performs multimodal reasoning over the generated visual content. All three roles are initialized from the same base model and trained using Group Relative Policy Optimization (GRPO), with carefully designed reward mechanisms that integrate execution feedback, visual verification, and difficulty balancing. Our experiments show that MM-Zero improves VLM reasoning performance across a wide range of multimodal benchmarks. MM-Zero establishes a scalable path toward self-evolving multi-model systems for multimodal models, extending the frontier of self-improvement beyond the conventional two-model paradigm.

03

ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

While Large Language Models (LLMs) have revolutionized code generation, standard "System 1" approaches, generating solutions in a single forward pass, often hit a performance ceiling when faced with complex algorithmic tasks. Existing iterative refinement strategies attempt to bridge this gap at inference time, yet they predominantly rely on external oracles, execution feedback, or computationally expensive prompt-response cycles. In this work, we propose ReflexiCoder, a novel reinforcement learning (RL) framework that internalizes the structured reasoning trajectory, encompassing initial generation, bug and optimization aware reflection, and self-correction, directly into the model's weights. Unlike prior methods, ReflexiCoder shifts the paradigm from external-dependent refinement to an intrinsic, fully autonomous self-reflection and self-correction capabilities at inference time. We utilize an RL-zero training paradigm with granular reward functions to optimize the entire reflection-correction trajectory, teaching the model how to debug without reliance on ground-truth feedback or execution engines at inference time. Extensive experiments across seven benchmarks demonstrate that our ReflexiCoder-8B establishes a new state-of-the-art (SOTA) among leading open-source models in the 1.5B-14B range, achieving 94.51% (87.20%) on HumanEval (Plus), 81.80% (78.57%) on MBPP (Plus), 35.00% on BigCodeBench, 52.21% on LiveCodeBench, and 37.34% on CodeForces in a single-attempt setting, rivaling or surpassing proprietary models like GPT-5.1. Notably, our framework is significantly more token-efficient than base models, reducing inference-time compute overhead by approximately 40% through disciplined, high-speed reasoning and reflection patterns. Source code is available at https://github.com/juyongjiang/ReflexiCoder.

04

InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing

Unified multimodal models (UMMs) that integrate understanding, reasoning, generation, and editing face inherent trade-offs between maintaining strong semantic comprehension and acquiring powerful generation capabilities. In this report, we present InternVL-U, a lightweight 4B-parameter UMM that democratizes these capabilities within a unified framework. Guided by the principles of unified contextual modeling and modality-specific modular design with decoupled visual representations, InternVL-U integrates a state-of-the-art Multimodal Large Language Model (MLLM) with a specialized MMDiT-based visual generation head. To further bridge the gap between aesthetic generation and high-level intelligence, we construct a comprehensive data synthesis pipeline targeting high-semantic-density tasks, such as text rendering and scientific reasoning, under a reasoning-centric paradigm that leverages Chain-of-Thought (CoT) to better align abstract user intent with fine-grained visual generation details. Extensive experiments demonstrate that InternVL-U achieves a superior performance - efficiency balance. Despite using only 4B parameters, it consistently outperforms unified baseline models with over 3x larger scales such as BAGEL (14B) on various generation and editing tasks, while retaining strong multimodal understanding and reasoning capabilities.

05

The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness

Situational awareness, the capacity of an AI system to recognize its own nature, understand its training and deployment context, and reason strategically about its circumstances, is widely considered among the most dangerous emergent capabilities in advanced AI systems. Separately, a growing research effort seeks to improve the logical reasoning capabilities of large language models (LLMs) across deduction, induction, and abduction. In this paper, we argue that these two research trajectories are on a collision course. We introduce the RAISE framework (Reasoning Advancing Into Self Examination), which identifies three mechanistic pathways through which improvements in logical reasoning enable progressively deeper levels of situational awareness: deductive self inference, inductive context recognition, and abductive self modeling. We formalize each pathway, construct an escalation ladder from basic self recognition to strategic deception, and demonstrate that every major research topic in LLM logical reasoning maps directly onto a specific amplifier of situational awareness. We further analyze why current safety measures are insufficient to prevent this escalation. We conclude by proposing concrete safeguards, including a "Mirror Test" benchmark and a Reasoning Safety Parity Principle, and pose an uncomfortable but necessary question to the logical reasoning community about its responsibility in this trajectory.

06

Streaming Autoregressive Video Generation via Diagonal Distillation

Large pretrained diffusion models have significantly enhanced the quality of generated videos, and yet their use in real-time streaming remains limited. Autoregressive models offer a natural framework for sequential frame synthesis but require heavy computation to achieve high fidelity. Diffusion distillation can compress these models into efficient few-step variants, but existing video distillation approaches largely adapt image-specific methods that neglect temporal dependencies. These techniques often excel in image generation but underperform in video synthesis, exhibiting reduced motion coherence, error accumulation over long sequences, and a latency-quality trade-off. We identify two factors that result in these limitations: insufficient utilization of temporal context during step reduction and implicit prediction of subsequent noise levels in next-chunk prediction (i.e., exposure bias). To address these issues, we propose Diagonal Distillation, which operates orthogonally to existing approaches and better exploits temporal information across both video chunks and denoising steps. Central to our approach is an asymmetric generation strategy: more steps early, fewer steps later. This design allows later chunks to inherit rich appearance information from thoroughly processed early chunks, while using partially denoised chunks as conditional inputs for subsequent synthesis. By aligning the implicit prediction of subsequent noise levels during chunk generation with the actual inference conditions, our approach mitigates error propagation and reduces oversaturation in long-range sequences. We further incorporate implicit optical flow modeling to preserve motion quality under strict step constraints. Our method generates a 5-second video in 2.61 seconds (up to 31 FPS), achieving a 277.3x speedup over the undistilled model.