NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-09中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Show HN: The Mog Programming Language

Mog is a newly introduced statically typed, compiled, and embedded programming language explicitly designed for Large Language Models (LLMs) to write. Its entire specification is compact, fitting within 3,200 tokens, making it easily manageable within an AI's context window. An AI agent can write a Mog program, compile it, and then dynamically load it for use as a plugin, script, or hook. A key feature is its capability-based permissions system, allowing the host to precisely control which functions a Mog program can access, thus ensuring secure execution. Mog compiles directly to native code, offering low-latency execution without interpreter overhead, JIT compilation, or process startup costs. The compiler, written in safe Rust, is auditable for security, providing a robust toolchain for agents to extend their capabilities with self-generated code. Mog is MIT licensed and open for contributions.

02

Launch HN: Terminal Use (YC W26) – Vercel for filesystem-based agents

Terminal Use, a YC W26 startup founded by Filip, Stavros, and Vivek, has launched a platform described as "Vercel for filesystem-based agents." The service aims to simplify the deployment and management of AI agents that operate within sandboxed environments and rely on file system access for their functions. This encompasses a range of applications, including agents designed for coding, research, document processing, and internal tools that interact with files. The founders emphasize that the current landscape for hosting such agents is fragmented, requiring developers to integrate disparate components for agent packaging, secure sandboxing, real-time message streaming, state persistence across interactions, and efficient file transfer to and from agent workspaces. Terminal Use seeks to consolidate these functionalities into a single, cohesive platform, drawing parallels to Replicate's Cog but with a specific focus on robust filesystem capabilities. A demonstration video illustrates the platform's practical applications.

03

Code Review for Claude Code

Anthropic's blog post, 'Code Review for Claude Code,' provides insights into the essential practices and methodologies employed for rigorously reviewing code associated with its advanced AI assistant, Claude. The article is anticipated to outline the unique challenges and specific considerations inherent in conducting code reviews for large language models and their supporting infrastructure. This includes ensuring adherence to stringent safety protocols, upholding ethical guidelines, and optimizing for performance and efficiency in AI deployments. The discussion likely explores how human expert reviewers interact with both AI-generated code and the foundational code that powers Claude's functionalities, focusing on identifying potential vulnerabilities, mitigating biases, and resolving technical inefficiencies. The piece is expected to detail the best practices adopted by Anthropic's engineering and research teams to maintain exceptionally high code quality and reliability, underscoring the paramount importance of thorough and systematic scrutiny in modern AI development. Furthermore, the blog post could highlight specialized tools and refined workflows that facilitate highly effective code review processes within a complex, AI-centric development environment, ultimately contributing to the robustness, security, and trustworthiness of sophisticated AI systems such as Claude.

04

I gave my robot physical memory – it stopped repeating mistakes

A novel approach in robotics introduces the concept of 'physical memory,' designed to enable robots to learn from past experiences and autonomously avoid repeating errors. This system integrates a mechanism allowing robots to record and retrieve data specifically related to their physical interactions with the environment. By establishing a persistent memory of these encounters, the robot can significantly improve its ability to self-correct and refine actions, moving beyond solely pre-programmed instructions. This development is critical for enhancing robot reliability, efficiency, and safety, particularly in dynamic and complex operational settings. The innovation promises more intelligent and adaptive robotic systems capable of continuous experiential learning, holding significant implications for advanced automation, autonomous navigation, and other applications requiring robust, error-resistant machine learning.

05

Building a Procedural Hex Map with Wave Function Collapse

This article details the implementation of a procedural hexagonal map generation system leveraging the Wave Function Collapse (WFC) algorithm. WFC is a sophisticated constraint-based generative technique that significantly improves upon simple random generation by ensuring spatial coherence and structural integrity. The core principle involves defining a set of elementary patterns, often derived from a sample input, and establishing strict rules for how these patterns can adjoin one another. Through an iterative process, the "wave function" of each cell on the map\nas well as representing its possible states\nis collapsed by selecting a pattern that satisfies all surrounding constraints, thus propagating choices and eliminating incompatible options. When applied to hexagonal grids, this methodology facilitates the creation of visually diverse and logically structured maps, enabling seamless transitions between different terrain types such as forests, mountains, and rivers. This approach is highly relevant for developers in game design, architecture, and other simulation fields seeking to automatically generate expansive, unique worlds with controlled complexity and artistic consistency, fundamentally demonstrating how emergent global patterns can be derived from localized rules.

06

Is legal the same as legitimate: AI reimplementation and the erosion of copyleft

The article title "Is legal the same as legitimate: AI reimplementation and the erosion of copyleft" prompts a critical examination of the evolving landscape where Artificial Intelligence development intersects with established intellectual property and software licensing frameworks. It specifically questions the distinction between legally permissible actions and what is considered ethically legitimate, particularly concerning the reimplementation of functionalities or concepts by AI systems. The central thesis likely explores how AI's capacity to learn from and subsequently replicate or generate code and designs, often inspired by existing copylefted works, could undermine the foundational principles of copyleft. This erosion poses significant challenges to open-source software communities and raises complex questions about attribution, derivative works, and the enforcement of licensing agreements in an era where AI-driven creation is increasingly prevalent. The discussion aims to highlight the urgent need for new legal interpretations or ethical guidelines to address these unprecedented technological shifts and protect intellectual property rights without stifling innovation.

huggingface

6 stories
01

Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders

Vision Language Model (VLM) development has largely relied on scaling model size, which hinders deployment on compute-constrained mobile and edge devices such as smartphones and robots. In this work, we explore the performance limits of compact (e.g., 2B and 8B) VLMs. We challenge the prevailing practice that state-of-the-art VLMs must rely on vision encoders initialized via massive contrastive pretraining (e.g., CLIP/SigLIP). We identify an objective mismatch: contrastive learning, optimized for discrimination, enforces coarse and category-level invariances that suppress fine-grained visual cues needed for dense captioning and complex VLM reasoning. To address this issue, we present Penguin-VL, whose vision encoder is initialized from a text-only LLM. Our experiments reveal that Penguin-Encoder serves as a superior alternative to traditional contrastive pretraining, unlocking a higher degree of visual fidelity and data efficiency for multimodal understanding. Across various image and video benchmarks, Penguin-VL achieves performance comparable to leading VLMs (e.g., Qwen3-VL) in mathematical reasoning and surpasses them in tasks such as document understanding, visual knowledge, and multi-perspective video understanding. Notably, these gains are achieved with a lightweight architecture, demonstrating that improved visual representation rather than model scaling is the primary driver of performance. Our ablations show that Penguin-Encoder consistently outperforms contrastive-pretrained encoders, preserving fine-grained spatial and temporal cues that are critical for dense perception and complex reasoning. This makes it a strong drop-in alternative for compute-efficient VLMs and enables high performance in resource-constrained settings. Code: https://github.com/tencent-ailab/Penguin-VL

02

BandPO: Bridging Trust Regions and Ratio Clipping via Probability-Aware Bounds for LLM Reinforcement Learning

Proximal constraints are fundamental to the stability of the Large Language Model reinforcement learning. While the canonical clipping mechanism in PPO serves as an efficient surrogate for trust regions, we identify a critical bottleneck: fixed bounds strictly constrain the upward update margin of low-probability actions, disproportionately suppressing high-advantage tail strategies and inducing rapid entropy collapse. To address this, we introduce Band-constrained Policy Optimization (BandPO). BandPO replaces canonical clipping with Band, a unified theoretical operator that projects trust regions defined by f-divergences into dynamic, probability-aware clipping intervals. Theoretical analysis confirms that Band effectively resolves this exploration bottleneck. We formulate this mapping as a convex optimization problem, guaranteeing a globally optimal numerical solution while deriving closed-form solutions for specific divergences. Extensive experiments across diverse models and datasets demonstrate that BandPO consistently outperforms canonical clipping and Clip-Higher, while robustly mitigating entropy collapse.

03

FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase. While various sparse attention mechanisms have been explored, they typically suffer from either significant search latency or insufficient sparsity. In this paper, we propose FlashPrefill, a framework enabling ultra-fast prefilling via instantaneous pattern discovery and thresholding. FlashPrefill leverages a fast block-searching technique to simultaneously locate dynamic vertical, slash, and block-sparse attention patterns. Crucially, it introduces a dynamic thresholding mechanism that bypasses the prohibitive overhead of sorting or accumulating attention scores while effectively eliminating the long-tail distribution to enhance sparsity. Extensive evaluations demonstrate that FlashPrefill achieves a substantial leap in efficiency, delivering an unprecedented 27.78x speedup on 256K sequences. Notably, unlike existing methods that incur efficiency degradation on shorter contexts, FlashPrefill maintains a 1.71x speedup even at a 4K context length, demonstrating its robustness and practical utility across varying sequence scales.

04

HiMAP-Travel: Hierarchical Multi-Agent Planning for Long-Horizon Constrained Travel

Sequential LLM agents fail on long-horizon planning with hard constraints like budgets and diversity requirements. As planning progresses and context grows, these agents drift from global constraints. We propose HiMAP-Travel, a hierarchical multi-agent framework that splits planning into strategic coordination and parallel day-level execution. A Coordinator allocates resources across days, while Day Executors plan independently in parallel. Three key mechanisms enable this: a transactional monitor enforcing budget and uniqueness constraints across parallel agents, a bargaining protocol allowing agents to reject infeasible sub-goals and trigger re-planning, and a single policy trained with GRPO that powers all agents through role conditioning. On TravelPlanner, HiMAP-Travel with Qwen3-8B achieves 52.78% validation and 52.65% test Final Pass Rate (FPR). In a controlled comparison with identical model, training, and tools, it outperforms the sequential DeepTravel baseline by +8.67~pp. It also surpasses ATLAS by +17.65~pp and MTP by +10.0~pp. On FlexTravelBench multi-turn scenarios, it achieves 44.34% (2-turn) and 37.42% (3-turn) FPR while reducing latency 2.5x through parallelization.

05

EffectMaker: Unifying Reasoning and Generation for Customized Visual Effect Creation

Visual effects (VFX) are essential for enhancing the expressiveness and creativity of video content, yet producing high-quality effects typically requires expert knowledge and costly production pipelines. Existing AIGC systems face significant challenges in VFX generation due to the scarcity of effect-specific data and the inherent difficulty of modeling supernatural or stylized effects. Moreover, these approaches often require per-effect fine-tuning, which severely limits their scalability and generalization to novel VFX. In this work, we present EffectMaker, a unified reasoning-generation framework that enables reference-based VFX customization. EffectMaker employs a multimodal large language model to interpret high-level effect semantics and reason about how they should adapt to a target subject, while a diffusion transformer leverages in-context learning to capture fine-grained visual cues from reference videos. These two components form a semantic-visual dual-path guidance mechanism that enables accurate, controllable, and effect-consistent synthesis without per-effect fine-tuning. Furthermore, we construct EffectData, the largest high-quality synthetic dataset containing 130k videos across 3k VFX categories, to improve generalization and scalability. Experiments show that EffectMaker achieves superior visual quality and effect consistency over state-of-the-art baselines, offering a scalable and flexible paradigm for customized VFX generation. Project page: https://effectmaker.github.io

06

Physical Simulator In-the-Loop Video Generation

Recent advances in diffusion-based video generation have achieved remarkable visual realism but still struggle to obey basic physical laws such as gravity, inertia, and collision. Generated objects often move inconsistently across frames, exhibit implausible dynamics, or violate physical constraints, limiting the realism and reliability of AI-generated videos. We address this gap by introducing Physical Simulator In-the-loop Video Generation (PSIVG), a novel framework that integrates a physical simulator into the video diffusion process. Starting from a template video generated by a pre-trained diffusion model, PSIVG reconstructs the 4D scene and foreground object meshes, initializes them within a physical simulator, and generates physically consistent trajectories. These simulated trajectories are then used to guide the video generator toward spatio-temporally physically coherent motion. To further improve texture consistency during object movement, we propose a Test-Time Texture Consistency Optimization (TTCO) technique that adapts text and feature embeddings based on pixel correspondences from the simulator. Comprehensive experiments demonstrate that PSIVG produces videos that better adhere to real-world physics while preserving visual quality and diversity. Project Page: https://vcai.mpi-inf.mpg.de/projects/PSIVG/