NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-04-08中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Muse Spark: Scaling towards personal superintelligence

Meta AI has officially introduced "Muse Spark," a groundbreaking initiative conceptualized to scale towards the realization of "personal superintelligence." This strategic move indicates Meta's profound commitment to advancing artificial intelligence beyond current paradigms, focusing on developing highly sophisticated and deeply personalized AI systems. While granular technical specifications and the underlying architectural framework of Muse Spark remain to be fully disclosed, the core objective points towards creating AI agents or models capable of delivering unprecedented levels of autonomy, adaptability, and intellectual capability tailored to individual user needs. This undertaking is expected to integrate state-of-the-art developments in large language models and sophisticated AI agent design, aiming to augment human cognition and interaction within digital ecosystems. Muse Spark represents a pivotal step in Meta's long-term vision for AI, striving to redefine personal computing through intelligent, anticipatory, and highly capable digital companions.

02

MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU

A groundbreaking paper titled 'MegaTrain' introduces a novel approach enabling the full precision training of Large Language Models (LLMs) with over 100 billion parameters on a single Graphics Processing Unit (GPU). This development marks a substantial advancement in AI model training, traditionally requiring extensive distributed computing setups and often relying on reduced precision techniques to manage memory demands. MegaTrain's methodology effectively addresses the formidable memory and computational challenges associated with training ultra-large models, allowing researchers and developers to conduct high-fidelity training without the need for large-scale GPU clusters. This innovation promises to democratize access to advanced LLM research and development, potentially accelerating the iteration cycles for next-generation AI models. The ability to train such massive models in full precision on a single GPU could also lead to more stable training, improved model performance, and reduced infrastructure costs, making advanced AI capabilities more accessible.

03

Claude Managed Agents

Anthropic has officially unveiled "Claude Managed Agents," a significant new offering poised to revolutionize how businesses deploy and operate AI agents powered by its advanced Claude large language models. This initiative is specifically designed to alleviate the substantial complexities associated with developing, scaling, and securely maintaining sophisticated AI agents for diverse enterprise applications. By providing a comprehensive managed service, Anthropic aims to significantly lower the technical and operational barriers for organizations aspiring to integrate cutting-edge AI capabilities into their workflows. The service promises enhanced reliability, optimized performance, and robust security measures, allowing businesses to harness the full potential of agentic AI systems without the extensive overhead of managing intricate infrastructure or requiring deep in-house specialized expertise. This strategic move underscores Anthropic's commitment to making sophisticated AI agent technology more accessible, scalable, and enterprise-ready, thereby accelerating the widespread adoption of autonomous AI solutions across various industry sectors and fostering innovation in automated problem-solving.

04

We Fingerprinted 178 AI Models' Writing Styles and Similarity Clusters

A research initiative has meticulously fingerprinted the writing styles of 178 distinct AI models by analyzing a dataset of 3,095 standardized responses across 43 varied prompts. This process involved extracting a 32-dimension stylometric fingerprint from each response, encompassing features like lexical richness, sentence structure, punctuation habits, formatting patterns, and discourse markers. Key findings from this extensive analysis include the identification of nine "clone clusters," where models demonstrated over 90% cosine similarity in their normalized feature vectors. The study also highlighted notable stylistic resemblances, such as Gemini 2.5 Flash Lite exhibiting 78% similarity to Claude 3 Opus, at a significantly reduced cost. Meta was identified as possessing the most distinctive provider "house style." Interestingly, prompts like "satirical fake news" led to the greatest writing convergence among models, while "count letters" prompted the most divergence in writing styles. This study provides valuable insights into the stylistic landscape of contemporary AI models through a comprehensive, multi-metric similarity assessment.

05

Show HN: Skrun – Deploy any agent skill as an API

Skrun, a new open-source project showcased on Hacker News, offers a streamlined solution for operationalizing artificial intelligence agent skills by enabling their deployment as standard API endpoints. This platform is designed to significantly simplify the integration of complex AI agent functionalities into existing software applications and services. By abstracting the underlying agent frameworks and exposing specific skills through a well-defined API, Skrun empowers developers to seamlessly incorporate sophisticated AI capabilities, such as automated tasks, decision-making processes, or specialized data analysis, without requiring extensive knowledge of agent-specific architectures. The initiative aims to enhance the interoperability and reusability of AI agents, fostering a more modular and interconnected ecosystem for AI development. This approach not only democratizes access to advanced AI agent technology but also accelerates the development lifecycle, allowing for faster prototyping and scalable deployment of agent-driven solutions across various industries.

06

Show HN: TUI-use: Let AI agents control interactive terminal programs

TUI-use, a project highlighted on Hacker News, introduces a novel approach to empower AI agents by enabling them to effectively control and interact with interactive terminal programs, commonly known as Text-based User Interfaces (TUIs). This development is crucial for advancing AI agent autonomy, particularly in environments where traditional graphical interfaces are not available or preferred. The project addresses the challenge of allowing AI to perceive and understand the dynamic state of a TUI, subsequently executing precise actions to navigate, input data, or manipulate program flows. Such capabilities could revolutionize automation in software development, system administration, and other technical domains that heavily rely on command-line tools. By bridging the gap between sophisticated AI reasoning and the practicalities of terminal interactions, TUI-use promises significant improvements in automating complex, multi-step tasks that previously demanded human intervention, thereby enhancing operational efficiency and broadening the scope of AI applications in technical workflows.

huggingface

6 stories
01

Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents

Large language models are increasingly deployed as autonomous agents executing multi-step workflows in real-world software environments. However, existing agent benchmarks suffer from three critical limitations: (1) trajectory-opaque grading that checks only final outputs, (2) underspecified safety and robustness evaluation, and (3) narrow modality coverage and interaction paradigms. We introduce Claw-Eval, an end-to-end evaluation suite addressing all three gaps. It comprises 300 human-verified tasks spanning 9 categories across three groups (general service orchestration, multimodal perception and generation, and multi-turn professional dialogue). Every agent action is recorded through three independent evidence channels (execution traces, audit logs, and environment snapshots), enabling trajectory-aware grading over 2,159 fine-grained rubric items. The scoring protocol evaluates Completion, Safety, and Robustness, reporting Average Score, Pass@k, and Pass^k across three trials to distinguish genuine capability from lucky outcomes. Experiments on 14 frontier models reveal that: (1) trajectory-opaque evaluation is systematically unreliable, missing 44% of safety violations and 13% of robustness failures that our hybrid pipeline catches; (2) controlled error injection primarily degrades consistency rather than peak capability, with Pass^3 dropping up to 24% while Pass@3 remains stable; (3) multimodal performance varies sharply, with most models performing poorer on video than on document or image, and no single model dominating across all modalities. Beyond benchmarking, Claw-Eval highlights actionable directions for agent development, shedding light on what it takes to build agents that are not only capable but reliably deployable.

02

In-Place Test-Time Training

The static "train then deploy" paradigm fundamentally limits Large Language Models (LLMs) from dynamically adapting their weights in response to continuous streams of new information inherent in real-world tasks. Test-Time Training (TTT) offers a compelling alternative by updating a subset of model parameters (fast weights) at inference time, yet its potential in the current LLM ecosystem is hindered by critical barriers including architectural incompatibility, computational inefficiency and misaligned fast weight objectives for language modeling. In this work, we introduce In-Place Test-Time Training (In-Place TTT), a framework that seamlessly endows LLMs with Test-Time Training ability. In-Place TTT treats the final projection matrix of the ubiquitous MLP blocks as its adaptable fast weights, enabling a "drop-in" enhancement for LLMs without costly retraining from scratch. Furthermore, we replace TTT's generic reconstruction objective with a tailored, theoretically-grounded objective explicitly aligned with the Next-Token-Prediction task governing autoregressive language modeling. This principled objective, combined with an efficient chunk-wise update mechanism, results in a highly scalable algorithm compatible with context parallelism. Extensive experiments validate our framework's effectiveness: as an in-place enhancement, it enables a 4B-parameter model to achieve superior performance on tasks with contexts up to 128k, and when pretrained from scratch, it consistently outperforms competitive TTT-related approaches. Ablation study results further provide deeper insights on our design choices. Collectively, our results establish In-Place TTT as a promising step towards a paradigm of continual learning in LLMs.

03

Action Images: End-to-End Policy Learning via Multiview Video Generation

World action models (WAMs) have emerged as a promising direction for robot policy learning, as they can leverage powerful video backbones to model the future states. However, existing approaches often rely on separate action modules, or use action representations that are not pixel-grounded, making it difficult to fully exploit the pretrained knowledge of video models and limiting transfer across viewpoints and environments. In this work, we present Action Images, a unified world action model that formulates policy learning as multiview video generation. Instead of encoding control as low-dimensional tokens, we translate 7-DoF robot actions into interpretable action images: multi-view action videos that are grounded in 2D pixels and explicitly track robot-arm motion. This pixel-grounded action representation allows the video backbone itself to act as a zero-shot policy, without a separate policy head or action module. Beyond control, the same unified model supports video-action joint generation, action-conditioned video generation, and action labeling under a shared representation. On RLBench and real-world evaluations, our model achieves the strongest zero-shot success rates and improves video-action joint generation quality over prior video-space world models, suggesting that interpretable action images are a promising route to policy learning.

04

MMEmb-R1: Reasoning-Enhanced Multimodal Embedding with Pair-Aware Selection and Adaptive Control

MLLMs have been successfully applied to multimodal embedding tasks, yet their generative reasoning capabilities remain underutilized. Directly incorporating chain-of-thought reasoning into embedding learning introduces two fundamental challenges. First, structural misalignment between instance-level reasoning and pairwise contrastive supervision may lead to shortcut behavior, where the model merely learns the superficial format of reasoning. Second, reasoning is not universally beneficial for embedding tasks. Enforcing reasoning for all inputs may introduce unnecessary computation and latency, and can even obscure salient semantic signals for simple cases. To address these issues, we propose MMEmb-R1, an adaptive reasoning-based multimodal embedding framework. We formulate reasoning as a latent variable and introduce pair-aware reasoning selection that employs counterfactual intervention to identify reasoning paths beneficial for query-target alignment. Furthermore, we adopt reinforcement learning to selectively invoke reasoning only when necessary. Experiments on the MMEB-V2 benchmark demonstrate that our model achieves a score of 71.2 with only 4B parameters, establishing a new state-of-the-art while significantly reducing reasoning overhead and inference latency.

05

Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning

Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals into editable TikZ code. While TikZ is the de facto standard for scientific schematics due to its programmatic flexibility, its requirement for rigorous spatial precision presents a significant challenge for Multimodal Large Language Models. Progress is currently stifled by two primary gaps: (1) Data Quality Gap: existing image-TikZ corpora often lack strict executability and reliable visual alignment; (2) Evaluation Gap: a lack of benchmarks for both structural and visual fidelity. To address these, we present a closed-loop framework featuring: SciTikZ-230K, a large-scale, high-quality dataset from our Execution-Centric Data Engine covering 11 diverse scientific disciplines; SciTikZ-Bench, a multifaceted benchmark spanning from basic geometric constructs to intricate hierarchical schematics to evaluate both visual fidelity and structural logic. To further broaden the scope of visual-code optimization methodology, we introduce a novel Dual Self-Consistency Reinforcement Learning optimization paradigm, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency. Empowered by these, our trained model SciTikZer-8B achieves state-of-the-art performance, consistently outperforming proprietary giants like Gemini-2.5-Pro and massive models like Qwen3-VL-235B-A22B-Instruct.

06

QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization

Large Language Models (LLMs) achieve strong program repair performance but often suffer from over-editing, where excessive modifications overwrite correct code and hinder bug localization. We systematically quantify its impact and introduce precise repair task, which maximizes reuse of correct code while fixing only buggy parts. Building on this insight, we propose PRepair, a framework that mitigates over-editing and improves repair accuracy. PRepair has two components: Self-Breaking, which generates diverse buggy programs via controlled bug injection and min-max sampling, and Self-Repairing, which trains models with Edit-Aware Group Relative Policy Optimization (EA-GRPO) using an edit-aware reward to encourage minimal yet correct edits. Experiments show that PRepair improves repair precision by up to 31.4% under fix_1@1, a metric that jointly considers repair correctness and extent, and significantly increases decoding throughput when combined with speculative editing, demonstrating its potential for precise and practical code repair.