NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-04-07中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Project Glasswing: Securing critical software for the AI era

Project Glasswing, spearheaded by Anthropic, represents a pivotal initiative dedicated to confronting the intricate and expanding security challenges inherent in integrating Artificial Intelligence into critical software systems. The project's core mission is to develop and implement robust frameworks, advanced methodologies, and specialized tools designed to fortify the integrity, confidentiality, and availability of software essential for the AI era. This comprehensive approach targets a spectrum of threats, including adversarial attacks on machine learning models, data poisoning, and vulnerabilities throughout AI development pipelines, supply chains, and deployment environments. By specifically addressing critical software, Glasswing aims to significantly enhance the overall trustworthiness, resilience, and reliability of AI applications, ensuring their secure operation across various sensitive industries and societal functions. The initiative underscores the paramount importance of proactive and adaptive cybersecurity strategies to support the safe, ethical, and responsible advancement of AI technologies, thereby mitigating systemic risks associated with their increasing autonomy and pervasive influence. Its ultimate objective is to construct a secure and dependable foundation for the next generation of AI-powered critical infrastructure.

02

System Card: Claude Mythos Preview [pdf]

Anthropic has released a System Card preview for its forthcoming 'Claude Mythos' model, providing an initial look into its advanced capabilities and foundational design. This document, characteristic of a comprehensive System Card, is anticipated to detail the model's architectural innovations, performance benchmarks, and a thorough assessment of its ethical implications. It likely emphasizes the robust safety protocols and responsible AI practices Anthropic is integrating to address potential biases and ensure secure deployment. The 'Mythos' preview serves as an important informational resource for the AI community, including researchers and developers, offering insights into the technical specifications, potential applications, and the strategic vision behind this next-generation large language model. This early disclosure aims to foster transparency and informed discussion, paving the way for future advancements and responsible integration of cutting-edge AI technologies into various sectors, highlighting Anthropic's commitment to cautious and beneficial AI development.

03

GLM-5.1: Towards Long-Horizon Tasks

Z.AI has announced the release of GLM-5.1, a significant advancement in large language model technology specifically engineered to tackle long-horizon tasks. This new iteration focuses on enhancing the model's ability to engage in complex, multi-step reasoning, and sustained problem-solving, moving beyond traditional short-form conversational interactions. GLM-5.1 aims to improve planning capabilities, memory retention over extended sequences, and consistency in executing intricate workflows, paving the way for more sophisticated AI agents. The development represents a strategic shift towards applications requiring not just knowledge recall but also strategic foresight and persistent execution across various domains. By improving the model's capacity for deep contextual understanding and sequential action, GLM-5.1 is poised to empower AI systems to handle real-world challenges that demand sustained intelligence, such as advanced robotics control, complex software engineering tasks, intricate scientific simulations, and autonomous decision-making in dynamic environments. This release marks a crucial step towards developing more autonomous and truly intelligent AI systems capable of operating effectively over extended timeframes and complex operational parameters.

04

Assessing Claude Mythos Preview's cybersecurity capabilities

This report details an assessment of Claude Mythos Preview's cybersecurity capabilities, examining its robustness and potential vulnerabilities within complex digital environments. The evaluation likely encompasses various attack vectors, including prompt injection, data leakage, adversarial attacks, and system integrity compromises. Given that Claude is a large language model, the assessment would focus on its ability to generate secure code, identify security flaws in natural language descriptions, and resist manipulation attempts designed to extract sensitive information or misuse its generative functions. The findings aim to provide insights into the model's readiness for deployment in security-sensitive applications, highlighting areas of strength and identifying where further enhancements are required to mitigate risks and ensure responsible AI development. This proactive analysis is crucial for establishing trust and ensuring the secure integration of advanced AI models into critical infrastructure and enterprise systems.

05

Show HN: Gemma 4 Multimodal Fine-Tuner for Apple Silicon

A developer has unveiled a new project on Hacker News: a Gemma 4 multimodal fine-tuner engineered for Apple Silicon, specifically targeting M2 Ultra Macs. This ambitious initiative commenced approximately six months ago, originally focusing on the local fine-tuning of the Whisper model. Faced with a massive dataset comprising 15,000 hours of audio data stored in Google Cloud Storage, a bespoke streaming system was devised to facilitate data transfer directly from GCS to the local machine during the training process, bypassing local storage constraints. The project evolved to integrate Gemma 3n and has recently been updated to support the newly released Gemma 4, with its fine-tuning capabilities now distinct from the Whisper component. The developer highlighted a significant operational challenge: the frequent occurrence of Out-Of-Memory (OOM) errors on a 64GB RAM Mac Studio when fine-tuning longer sequences, underscoring the substantial memory demands associated with advanced local model training on consumer-grade hardware.

06

Google open-sources experimental agent orchestration testbed Scion

Google has officially open-sourced Scion, an experimental testbed specifically engineered for the orchestration and sophisticated management of AI agents. This strategic release aims to equip developers and researchers with a robust, comprehensive platform to design, test, and deploy complex multi-agent systems with greater efficiency and control. Scion is positioned to streamline the entire development lifecycle of AI agents, providing essential functionalities for seamless experimentation, resilient deployment, and detailed performance monitoring in diverse operational environments. By making Scion openly accessible, Google significantly contributes to fostering innovation and collaboration across the global artificial intelligence community. This initiative is particularly pivotal for advancing the frontier of agent-based AI solutions, where intricate coordination, communication, and interaction between autonomous entities are paramount. It is expected to substantially simplify the architectural complexity involved in creating sophisticated AI applications, thereby accelerating progress in the development and deployment of next-generation AI systems and empowering a wider range of AI-driven projects.

huggingface

6 stories
01

TriAttention: Efficient Long Reasoning with Trigonometric KV Compression

Extended reasoning in large language models (LLMs) creates severe KV cache memory bottlenecks. Leading KV cache compression methods estimate KV importance using attention scores from recent post-RoPE queries. However, queries rotate with position during RoPE, making representative queries very few, leading to poor top-key selection and unstable reasoning. To avoid this issue, we turn to the pre-RoPE space, where we observe that Q and K vectors are highly concentrated around fixed non-zero centers and remain stable across positions -- Q/K concentration. We show that this concentration causes queries to preferentially attend to keys at specific distances (e.g., nearest keys), with the centers determining which distances are preferred via a trigonometric series. Based on this, we propose TriAttention to estimate key importance by leveraging these centers. Via the trigonometric series, we use the distance preference characterized by these centers to score keys according to their positions, and also leverage Q/K norms as an additional signal for importance estimation. On AIME25 with 32K-token generation, TriAttention matches Full Attention reasoning accuracy while achieving 2.5x higher throughput or 10.7x KV memory reduction, whereas leading baselines achieve only about half the accuracy at the same efficiency. TriAttention enables OpenClaw deployment on a single consumer GPU, where long context would otherwise cause out-of-memory with Full Attention.

02

AURA: Always-On Understanding and Real-Time Assistance via Video Streams

Video Large Language Models (VideoLLMs) have achieved strong performance on many video understanding tasks, but most existing systems remain offline and are not well-suited for live video streams that require continuous observation and timely response. Recent streaming VideoLLMs have made progress, yet current approaches often rely on decoupled trigger-response pipelines or are limited to captioning-style narration, reducing their effectiveness for open-ended question answering and long-horizon interaction. We propose AURA (Always-On Understanding and Real-Time Assistance), an end-to-end streaming visual interaction framework that enables a unified VideoLLM to continuously process video streams and support both real-time question answering and proactive responses. AURA integrates context management, data construction, training objectives, and deployment optimization for stable long-horizon streaming interaction. It achieves state-of-the-art performance on streaming benchmarks and supports a real-time demo system with ASR and TTS running at 2 FPS on two 80G accelerators. We release the AURA model together with a real-time inference framework to facilitate future research.

03

Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw

OpenClaw, the most widely deployed personal AI agent in early 2026, operates with full local system access and integrates with sensitive services such as Gmail, Stripe, and the filesystem. While these broad privileges enable high levels of automation and powerful personalization, they also expose a substantial attack surface that existing sandboxed evaluations fail to capture. To address this gap, we present the first real-world safety evaluation of OpenClaw and introduce the CIK taxonomy, which unifies an agent's persistent state into three dimensions, i.e., Capability, Identity, and Knowledge, for safety analysis. Our evaluations cover 12 attack scenarios on a live OpenClaw instance across four backbone models (Claude Sonnet 4.5, Opus 4.6, Gemini 3.1 Pro, and GPT-5.4). The results show that poisoning any single CIK dimension increases the average attack success rate from 24.6% to 64-74%, with even the most robust model exhibiting more than a threefold increase over its baseline vulnerability. We further assess three CIK-aligned defense strategies alongside a file-protection mechanism; however, the strongest defense still yields a 63.8% success rate under Capability-targeted attacks, while file protection blocks 97% of malicious injections but also prevents legitimate updates. Taken together, these findings show that the vulnerabilities are inherent to the agent architecture, necessitating more systematic safeguards to secure personal AI agents. Our project page is https://ucsc-vlaa.github.io/CIK-Bench.

04

SkillX: Automatically Constructing Skill Knowledge Bases for Agents

Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited experience, resulting in redundant exploration and poor generalization. To address this problem, we propose SkillX, a fully automated framework for constructing a plug-and-play skill knowledge base that can be reused across agents and environments. SkillX operates through a fully automated pipeline built on three synergistic innovations: (i) Multi-Level Skills Design, which distills raw trajectories into three-tiered hierarchy of strategic plans, functional skills, and atomic skills; (ii) Iterative Skills Refinement, which automatically revises skills based on execution feedback to continuously improve library quality; and (iii) Exploratory Skills Expansion, which proactively generates and validates novel skills to expand coverage beyond seed training data. Using a strong backbone agent (GLM-4.6), we automatically build a reusable skill library and evaluate its transferability on challenging long-horizon, user-interactive benchmarks, including AppWorld, BFCL-v3, and τ^2-Bench. Experiments show that SkillKB consistently improves task success and execution efficiency when plugged into weaker base agents, highlighting the importance of structured, hierarchical experience representations for generalizable agent learning.

05

Can LLMs Learn to Reason Robustly under Noisy Supervision?

Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels due to expert scarcity remains critically underexplored. In this work, we take the first step toward a systematic analysis of noisy label mechanisms in RLVR. In contrast to supervised classification, most RLVR algorithms incorporate a rollout-based condition: a label's influence on training is contingent on whether the current policy can generate rollouts that realize it, a property that naturally extends to noisy labels. Based on this observation, we distinguish two types of noise: inactive noisy labels, which reduce data efficiency, and active noisy labels, which are reinforced and risk skewing the model toward incorrect distributions. From experiments on training with noisy samples, we identify an Early Correctness Coherence phenomenon: although noisy samples begin to lag behind in later stages, accuracy on both clean and noisy samples increases similarly in early training. Motivated by this dynamic, we propose Online Label Refinement (OLR), which progressively corrects potentially noisy labels with majority-voted answers when two conditions hold: a positive slope in the majority answer's rollout pass rate and stable historical consistency across updates, enabling gradual self-correction as the policy improves. We evaluate OLR on six in-distribution mathematical reasoning benchmarks (AIME24/25, AMC, MATH-500, Minerva, and Olympiad) and three out-of-distribution tasks (ARC-c, GPQA-diamond, and MMLU-pro). Across noise ratios from 0.1 to 0.9, OLR consistently improves robustness under both inactive and active noisy-label settings, achieving average gains of 3.6% to 3.9% on in-distribution benchmarks and 3.3% to 4.6% on out-of-distribution evaluations.

06

ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent editing of subjects and scenes remains a critical challenge. Recent attempts to incorporate richer environment control through rigid 3D geometric compositions often encounter a stark trade-off between precise control and generative flexibility. Furthermore, the heavy 3D pre-processing still limits practical scalability. In this paper, we propose ONE-SHOT, a parameter-efficient framework for compositional human-environment video generation. Our key insight is to factorize the generative process into disentangled signals. Specifically, we introduce a canonical-space injection mechanism that decouples human dynamics from environmental cues via cross-attention. We also propose Dynamic-Grounded-RoPE, a novel positional embedding strategy that establishes spatial correspondences between disparate spatial domains without any heuristic 3D alignments. To support long-horizon synthesis, we introduce a Hybrid Context Integration mechanism to maintain subject and scene consistency across minute-level generations. Experiments demonstrate that our method significantly outperforms state-of-the-art methods, offering superior structural control and creative diversity for video synthesis.