NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-05中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

GPT-5.4

OpenAI has officially announced the introduction of GPT-5.4, representing a notable advancement in its series of large language models. This new iteration is accompanied by a dedicated "Thinking System Card," signaling a strong emphasis on developing more robust reasoning capabilities, enhanced system reliability, and comprehensive ethical considerations within its operational architecture. While initial details suggest a focus on refining cognitive processes, this release is poised to significantly impact the landscape of artificial intelligence. It anticipates delivering superior performance across complex tasks, including advanced problem-solving and nuanced understanding. The introduction of GPT-5.4 further solidifies OpenAI's commitment to pushing the boundaries of AI research and development, influencing future applications in natural language processing and various generative AI domains. This launch marks another critical step in the rapid evolution of intelligent systems, setting new standards for sophisticated AI models.

02

Show HN: Jido 2.0, Elixir Agent Framework

Jido, an Elixir Agent Framework, has announced its 2.0 release, delivering a production-hardened platform for building, managing, and running AI Agents on the BEAM. This significant update introduces a comprehensive suite of agentic features, including advanced Tool Calling and Agent Skills, alongside robust multi-agent support across distributed BEAM processes with integrated Supervision. Jido 2.0 offers diverse reasoning strategies such as ReAct, Chain of Thought, and Tree of Thought, enabling complex AI workflows. The framework ensures durability through a sophisticated Storage and Persistence layer, provides Agentic Memory, and facilitates external service integration via MCP and Sensors. Furthermore, it boasts deep observability and debugging capabilities, including full-stack OTel. This release leverages the BEAM's architecture, which is increasingly recognized as highly suitable for agentic workloads, positioning Jido 2.0 as a crucial development for the Elixir ecosystem in the burgeoning field of AI agents.

03

A GitHub Issue Title Compromised 4k Developer Machines

A critical vulnerability, dubbed 'clinejection,' has been identified, demonstrating how a carefully crafted GitHub issue title can lead to the compromise of developer machines. The attack vector leverages AI-powered developer tools, which, when processing the malicious issue title, are tricked into executing arbitrary commands on the host system. This sophisticated injection technique, explored in detail by Grith AI, highlights a significant security flaw where AI tools designed to enhance productivity inadvertently become conduits for supply chain attacks. The incident reportedly led to the compromise of 4,000 developer machines, underscoring the severe risks associated with integrating AI into sensitive development workflows without robust security measures. This research warns of emerging threats where the interfaces of AI tools, particularly those interacting with external, untrusted content like GitHub issue titles, can be exploited for remote code execution and system compromise, urging developers and organizations to re-evaluate their AI tooling security posture.

04

Show HN: PageAgent, A GUI agent that lives inside your web app

PageAgent introduces an innovative open-source (MIT) library for embedding AI agents directly into web application frontends, moving away from traditional external client or server-side AI operations. The project promotes an "inside-out" paradigm, enabling a client-side agent to interact natively with the live DOM tree and inherit the user's active session, which is highly beneficial for Single Page Applications (SPAs). This approach aims to unlock a significant design space for deploying general AI agents natively within the web apps users already utilize, transforming the web into an integral part of the AI ecosystem rather than merely a passive target. By integrating the agent directly, it offers seamless interaction and efficient context retention. An optional browser extension is also developed to serve as a "bridge," facilitating cross-page tasks and expanding the agent's operational scope within the broader browsing experience. This unique architecture promises more dynamic and integrated AI functionalities for web users.

05

Pentagon Formally Labels Anthropic Supply-Chain Risk

The Pentagon has officially designated Anthropic, a prominent artificial intelligence developer recognized for its large language models, as a formal supply-chain risk. This unprecedented move signals heightened concerns within the U.S. Department of Defense regarding the security, reliability, or broader national security implications associated with integrating Anthropic's advanced AI technologies into defense systems or critical infrastructure. The formal labeling indicates an escalation in governmental scrutiny of leading AI companies, particularly those that may contribute to or be integrated within the defense industrial base. This decision underscores the complex challenges inherent in managing the dual-use nature of cutting-edge AI, prompting a re-evaluation of procurement processes and vendor relationships to mitigate potential vulnerabilities. This development reflects ongoing efforts to establish comprehensive AI governance frameworks and to assess the strategic implications of technological dependencies on external AI providers. It anticipates a more cautious and regulated approach to the adoption of advanced AI within the national security apparatus, potentially influencing future collaborations and contracts between AI firms and government agencies.

06

Relicensing with AI-Assisted Rewrite

The initiative of 'Relicensing with AI-Assisted Rewrite' focuses on leveraging artificial intelligence technologies to modernize and accelerate the complex process of changing software licenses. This innovative approach involves utilizing AI's advanced capabilities in natural language processing and code generation to automatically identify, adapt, and rewrite sections of existing code, documentation, or legal agreements to conform to new licensing terms. The primary objective is to significantly improve efficiency and reduce the manual effort typically associated with license transitions, thereby minimizing human error and accelerating project timelines. This method presents a compelling solution for organizations and open-source projects navigating legal compliance changes. However, it also necessitates careful consideration of the AI's output accuracy, its adherence to specific legal frameworks, and the potential for introducing unintended semantic shifts. This development highlights a crucial convergence of legal informatics, software engineering, and sophisticated AI applications, aiming to revolutionize how software licenses are managed and updated in the digital age.

huggingface

6 stories
01

Helios: Real Real-Time Long Video Generation Model

We introduce Helios, the first 14B video generation model that runs at 19.5 FPS on a single NVIDIA H100 GPU and supports minute-scale generation while matching the quality of a strong baseline. We make breakthroughs along three key dimensions: (1) robustness to long-video drifting without commonly used anti-drifting heuristics such as self-forcing, error-banks, or keyframe sampling; (2) real-time generation without standard acceleration techniques such as KV-cache, sparse/linear attention, or quantization; and (3) training without parallelism or sharding frameworks, enabling image-diffusion-scale batch sizes while fitting up to four 14B models within 80 GB of GPU memory. Specifically, Helios is a 14B autoregressive diffusion model with a unified input representation that natively supports T2V, I2V, and V2V tasks. To mitigate drifting in long-video generation, we characterize typical failure modes and propose simple yet effective training strategies that explicitly simulate drifting during training, while eliminating repetitive motion at its source. For efficiency, we heavily compress the historical and noisy context and reduce the number of sampling steps, yielding computational costs comparable to -- or lower than -- those of 1.3B video generative models. Moreover, we introduce infrastructure-level optimizations that accelerate both inference and training while reducing memory consumption. Extensive experiments demonstrate that Helios consistently outperforms prior methods on both short- and long-video generation. We plan to release the code, base model, and distilled model to support further development by the community.

02

Heterogeneous Agent Collaborative Reinforcement Learning

We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new learning paradigm that addresses the inefficiencies of isolated on-policy optimization. HACRL enables collaborative optimization with independent execution: heterogeneous agents share verified rollouts during training to mutually improve, while operating independently at inference time. Unlike LLM-based multi-agent reinforcement learning (MARL), HACRL does not require coordinated deployment, and unlike on-/off-policy distillation, it enables bidirectional mutual learning among heterogeneous agents rather than one-directional teacher-to-student transfer. Building on this paradigm, we propose HACPO, a collaborative RL algorithm that enables principled rollout sharing to maximize sample utilization and cross-agent knowledge transfer. To mitigate capability discrepancies and policy distribution shifts, HACPO introduces four tailored mechanisms with theoretical guarantees on unbiased advantage estimation and optimization correctness. Extensive experiments across diverse heterogeneous model combinations and reasoning benchmarks show that HACPO consistently improves all participating agents, outperforming GSPO by an average of 3.3% while using only half the rollout cost.

03

T2S-Bench & Structure-of-Thought: Benchmarking and Prompting Comprehensive Text-to-Structure Reasoning

Think about how human handles complex reading tasks: marking key points, inferring their relationships, and structuring information to guide understanding and responses. Likewise, can a large language model benefit from text structure to enhance text-processing performance? To explore it, in this work, we first introduce Structure of Thought (SoT), a prompting technique that explicitly guides models to construct intermediate text structures, consistently boosting performance across eight tasks and three model families. Building upon this insight, we present T2S-Bench, the first benchmark designed to evaluate and improve text-to-structure capabilities of models. T2S-Bench includes 1.8K samples across 6 scientific domains and 32 structural types, rigorously constructed to ensure accuracy, fairness, and quality. Evaluation on 45 mainstream models reveals substantial improvement potential: the average accuracy on the multi-hop reasoning task is only 52.1%, and even the most advanced model achieves 58.1% node accuracy in end-to-end extraction. Furthermore, on Qwen2.5-7B-Instruct, SoT alone yields an average +5.7% improvement across eight diverse text-processing tasks, and fine-tuning on T2S-Bench further increases this gain to +8.6%. These results highlight the value of explicit text structuring and the complementary contributions of SoT and T2S-Bench. Dataset and eval code have been released at https://t2s-bench.github.io/T2S-Bench-Page/.

04

Phi-4-reasoning-vision-15B Technical Report

We present Phi-4-reasoning-vision-15B, a compact open-weight multimodal reasoning model, and share the motivations, design choices, experiments, and learnings that informed its development. Our goal is to contribute practical insight to the research community on building smaller, efficient multimodal reasoning models and to share the result of these learnings as an open-weight model that is good at common vision and language tasks and excels at scientific and mathematical reasoning and understanding user interfaces. Our contributions include demonstrating that careful architecture choices and rigorous data curation enable smaller, open-weight multimodal models to achieve competitive performance with significantly less training and inference-time compute and tokens. The most substantial improvements come from systematic filtering, error correction, and synthetic augmentation -- reinforcing that data quality remains the primary lever for model performance. Systematic ablations show that high-resolution, dynamic-resolution encoders yield consistent improvements, as accurate perception is a prerequisite for high-quality reasoning. Finally, a hybrid mix of reasoning and non-reasoning data with explicit mode tokens allows a single model to deliver fast direct answers for simpler tasks and chain-of-thought reasoning for complex problems.

05

Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory

Large language model (LLM) agents are fundamentally bottlenecked by finite context windows on long-horizon tasks. To overcome this, we introduce Memex, an indexed experience memory mechanism that compresses context without discarding evidence. Memex maintains a compact working context using structured summaries and stable indices, storing full-fidelity interactions in an external experience database. Agents can dereference indices to retrieve exact past evidence. We optimize this with MemexRL, a reinforcement learning framework using reward shaping for indexed memory usage under context budgets, enabling agents to learn summarizing, archiving, indexing, and retrieval. This approach offers a less lossy form of long-horizon memory compared to summary-only methods. Theoretical analysis shows Memex's potential to preserve decision quality with bounded dereferencing and effective in-context computation. Empirically, Memex agents trained with MemexRL improve task success on challenging long-horizon tasks while using significantly smaller working contexts.

06

MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Models

Safety evaluation and red-teaming of large language models remain predominantly text-centric, and existing frameworks lack the infrastructure to systematically test whether alignment generalizes to audio, image, and video inputs. We present MUSE (Multimodal Unified Safety Evaluation), an open-source, run-centric platform that integrates automatic cross-modal payload generation, three multi-turn attack algorithms (Crescendo, PAIR, Violent Durian), provider-agnostic model routing, and an LLM judge with a five-level safety taxonomy into a single browser-based system. A dual-metric framework distinguishes hard Attack Success Rate (Compliance only) from soft ASR (including Partial Compliance), capturing partial information leakage that binary metrics miss. To probe whether alignment generalizes across modality boundaries, we introduce Inter-Turn Modality Switching (ITMS), which augments multi-turn attacks with per-turn modality rotation. Experiments across six multimodal LLMs from four providers show that multi-turn strategies can achieve up to 90-100% ASR against models with near-perfect single-turn refusal. ITMS does not uniformly raise final ASR on already-saturated baselines, but accelerates convergence by destabilizing early-turn defenses, and ablation reveals that the direction of modality effects is model-family-specific rather than universal, underscoring the need for provider-aware cross-modal safety testing.