NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-31中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

The Claude Code Source Leak: fake tools, frustration regexes, undercover mode

The recent disclosure, termed "The Claude Code Source Leak," presents a detailed technical analysis of code allegedly originating from the Claude AI model. This examination uncovers several notable elements, including the presence of "fake tools" which appear to be internal development or testing utilities, possibly designed for simulating external interactions without actual deployment. Furthermore, the analysis reportedly highlights a set of "frustration regexes," indicating sophisticated pattern-matching logic likely employed to identify and manage problematic or undesirable outputs and user interactions, revealing internal strategies for robustness and safety. The concept of an "undercover mode" also surfaces, potentially describing a discreet operational state for specific testing, debugging, or controlled performance evaluations. This incident offers an uncommon window into the intricate internal development practices, error-handling mechanisms, and strategic design choices that underpin the creation and ongoing refinement of a leading large language model, shedding light on the engineering challenges inherent in advanced AI development.

02

Slop is not necessarily the future

The article "Slop is not necessarily the future" presents a critical perspective on the emerging concept of "AI slopware," refuting the idea that the future of artificial intelligence is destined to be dominated by low-quality, generic, and uninspired outputs. It challenges the growing concern that the ease of AI content generation will inherently lead to a deluge of mediocre information, code, or creative work that lacks human touch, originality, or significant value. The piece likely advocates for a more optimistic and proactive outlook, suggesting that the industry and users will increasingly prioritize quality, specificity, and contextual relevance over sheer volume. This shift implies a trajectory where AI tools evolve to become more sophisticated assistants for high-value tasks, demanding better models, more refined prompts, and greater human oversight to produce truly meaningful results. Ultimately, the narrative asserts that the drive for excellence and utility will prevail, steering AI development away from the production of mere "slop" towards innovative and impactful applications.

03

Show HN: Cerno CAPTCHA that targets LLM reasoning, not human biology

Cerno introduces a novel CAPTCHA solution specifically engineered to challenge Large Language Models (LLMs) by testing their reasoning capabilities, rather than relying on human biological traits like visual perception or motor skills. This innovative approach seeks to address the increasing sophistication of AI models, which are often capable of bypassing traditional CAPTCHAs designed to distinguish humans from automated bots. Cerno aims to provide a more robust defense against AI-driven automation, ensuring that only human users can access protected online resources. The system's focus on cognitive reasoning tasks marks a significant shift in bot detection methodologies, adapting to the evolving landscape of artificial intelligence and its potential for malicious use. This development is particularly relevant for applications requiring enhanced security against advanced AI agents.

04

Microsoft: Copilot is for entertainment purposes only

Microsoft has introduced a significant clause within the terms of use for its AI assistant, Copilot, stipulating that the tool is intended "for entertainment purposes only." This clarification from the technology giant serves primarily to manage user expectations and delineate legal liabilities concerning the outputs generated by the advanced artificial intelligence system. By categorizing Copilot in this manner, Microsoft is likely seeking to mitigate risks associated with potential inaccuracies, factual errors, or 'hallucinations' that can occur with large language models. The disclaimer signals that while Copilot leverages sophisticated generative AI capabilities to assist users with various tasks, its results should not be implicitly trusted for critical, factual, or professional applications without independent verification. This strategic move highlights the ongoing challenges faced by AI developers in balancing the impressive capabilities of their models with the inherent limitations and unpredictability of current AI technology. It underscores a broader industry trend towards clearer disclaimers as AI tools become more ubiquitous, reflecting the evolving legal and ethical considerations surrounding the deployment and usage of generative AI in public domains.

05

Show HN: PhAIL Real-robot benchmark for AI models

PhAIL is introduced as a novel real-robot benchmark designed to evaluate the practical performance of Vision-Language-Action (VLA) AI models in commercial settings. Developed to address the scarcity of reliable performance data for VLA models, PhAIL rigorously tests four prominent models—OpenPI/pi0.5, GR00T, ACT, and SmolVLA—on a standardized bin-to-bin order picking task. The benchmark utilizes a Franka FR3 robot, identical objects, and hundreds of blind runs to ensure objectivity, with operators unaware of the model being tested. Initial results reveal a significant performance gap: the best VLA model achieved 64 Units Per Hour (UPH), considerably lower than a human teleoperating the same robot (330 UPH) or a human performing the task manually (1,300+ UPH). All experimental data, including synced video, telemetry, fine-tuning datasets, and training scripts, is publicly available, and the leaderboard is open for new submissions, fostering transparency and collaborative advancement in robotics AI.

06

Universal Claude.md cut Claude output tokens

This Hacker News story introduces "Universal Claude.md," a notable initiative aimed at significantly optimizing the output token usage for Claude AI models. The project, hosted on GitHub, provides methods and guidelines specifically designed to reduce the number of tokens generated by Claude, leading to enhanced operational efficiency and potential cost savings on API calls. By implementing token-cutting strategies, "Universal Claude.md" seeks to enable more concise and direct responses from the model, ensuring information quality is maintained while minimizing superfluous output. This optimization is particularly relevant for developers and researchers leveraging large language models, as managing token consumption is a critical factor in scaling AI applications effectively. The approach offers practical techniques for more token-efficient prompt engineering and response generation, addressing a key challenge in the practical deployment and sustained operation of advanced AI systems like Claude. It represents a valuable resource for improving the performance and economic viability of LLM-powered solutions.

huggingface

6 stories
01

Towards a Medical AI Scientist

Autonomous systems that generate scientific hypotheses, conduct experiments, and draft manuscripts have recently emerged as a promising paradigm for accelerating discovery. However, existing AI Scientists remain largely domain-agnostic, limiting their applicability to clinical medicine, where research is required to be grounded in medical evidence with specialized data modalities. In this work, we introduce Medical AI Scientist, the first autonomous research framework tailored to clinical autonomous research. It enables clinically grounded ideation by transforming extensively surveyed literature into actionable evidence through clinician-engineer co-reasoning mechanism, which improves the traceability of generated research ideas. It further facilitates evidence-grounded manuscript drafting guided by structured medical compositional conventions and ethical policies. The framework operates under 3 research modes, namely paper-based reproduction, literature-inspired innovation, and task-driven exploration, each corresponding to a distinct level of automated scientific inquiry with progressively increasing autonomy. Comprehensive evaluations by both large language models and human experts demonstrate that the ideas generated by the Medical AI Scientist are of substantially higher quality than those produced by commercial LLMs across 171 cases, 19 clinical tasks, and 6 data modalities. Meanwhile, our system achieves strong alignment between the proposed method and its implementation, while also demonstrating significantly higher success rates in executable experiments. Double-blind evaluations by human experts and the Stanford Agentic Reviewer suggest that the generated manuscripts approach MICCAI-level quality, while consistently surpassing those from ISBI and BIBM. The proposed Medical AI Scientist highlights the potential of leveraging AI for autonomous scientific discovery in healthcare.

02

Emergent Social Intelligence Risks in Generative Multi-Agent Systems

Multi-agent systems composed of large generative models are rapidly moving from laboratory prototypes to real-world deployments, where they jointly plan, negotiate, and allocate shared resources to solve complex tasks. While such systems promise unprecedented scalability and autonomy, their collective interaction also gives rise to failure modes that cannot be reduced to individual agents. Understanding these emergent risks is therefore critical. Here, we present a pioneer study of such emergent multi-agent risk in workflows that involve competition over shared resources (e.g., computing resources or market share), sequential handoff collaboration (where downstream agents see only predecessor outputs), collective decision aggregation, and others. Across these settings, we observe that such group behaviors arise frequently across repeated trials and a wide range of interaction conditions, rather than as rare or pathological cases. In particular, phenomena such as collusion-like coordination and conformity emerge with non-trivial frequency under realistic resource constraints, communication protocols, and role assignments, mirroring well-known pathologies in human societies despite no explicit instruction. Moreover, these risks cannot be prevented by existing agent-level safeguards alone. These findings expose the dark side of intelligent multi-agent systems: a social intelligence risk where agent collectives, despite no instruction to do so, spontaneously reproduce familiar failure patterns from human societies.

03

Density-aware Soft Context Compression with Semi-Dynamic Compression Ratio

Soft context compression reduces the computational workload of processing long contexts in LLMs by encoding long context into a smaller number of latent tokens. However, existing frameworks apply uniform compression ratios, failing to account for the extreme variance in natural language information density. While adopting a density-aware dynamic compression ratio seems intuitive, empirical investigations reveal that models struggle intrinsically with operations parameterized by input dependent, continuous structural hyperparameters. To resolve this pitfall, we introduce Semi-Dynamic Context Compression framework. Our approach features a Discrete Ratio Selector, which predicts a compression target based on intrinsic information density and quantizes it to a predefined set of discrete compression ratios. It is efficiently jointly trained with the compressor on synthetic data, with the summary lengths as a proxy to create labels for compression ratio prediction. Extensive evaluations confirm that our density-aware framework, utilizing mean pooling as the backbone, consistently outperforms static baselines, establishing a robust Pareto frontier for context compression techniques. Our code, data and model weights are available at https://github.com/yuyijiong/semi-dynamic-context-compress

04

On-the-fly Repulsion in the Contextual Space for Rich Diversity in Diffusion Transformers

Modern Text-to-Image (T2I) diffusion models have achieved remarkable semantic alignment, yet they often suffer from a significant lack of variety, converging on a narrow set of visual solutions for any given prompt. This typicality bias presents a challenge for creative applications that require a wide range of generative outcomes. We identify a fundamental trade-off in current approaches to diversity: modifying model inputs requires costly optimization to incorporate feedback from the generative path. In contrast, acting on spatially-committed intermediate latents tends to disrupt the forming visual structure, leading to artifacts. In this work, we propose to apply repulsion in the Contextual Space as a novel framework for achieving rich diversity in Diffusion Transformers. By intervening in the multimodal attention channels, we apply on-the-fly repulsion during the transformer's forward pass, injecting the intervention between blocks where text conditioning is enriched with emergent image structure. This allows for redirecting the guidance trajectory after it is structurally informed but before the composition is fixed. Our results demonstrate that repulsion in the Contextual Space produces significantly richer diversity without sacrificing visual fidelity or semantic adherence. Furthermore, our method is uniquely efficient, imposing a small computational overhead while remaining effective even in modern "Turbo" and distilled models where traditional trajectory-based interventions typically fail.

05

DreamLite: A Lightweight On-Device Unified Model for Image Generation and Editing

Diffusion models have made significant progress in both text-to-image (T2I) generation and text-guided image editing. However, these models are typically built with billions of parameters, leading to high latency and increased deployment challenges. While on-device diffusion models improve efficiency, they largely focus on T2I generation and lack support for image editing. In this paper, we propose DreamLite, a compact unified on-device diffusion model (0.39B) that supports both T2I generation and text-guided image editing within a single network. DreamLite is built on a pruned mobile U-Net backbone and unifies conditioning through in-context spatial concatenation in the latent space. It concatenates images horizontally as input, using a (target | blank) configuration for generation tasks and (target | source) for editing tasks. To stabilize the training of this compact model, we introduce a task-progressive joint pretraining strategy that sequentially targets T2I, editing, and joint tasks. After high-quality SFT and reinforcement learning, DreamLite achieves GenEval (0.72) for image generation and ImgEdit (4.11) for image editing, outperforming existing on-device models and remaining competitive with several server-side models. By employing step distillation, we further reduce denoising processing to just 4 steps, enabling our DreamLite could generate or edit a 1024 x 1024 image in less than 1s on a Xiaomi 14 smartphone. To the best of our knowledge, DreamLite is the first unified on-device diffusion model that supports both image generation and image editing.

06

KAT-Coder-V2 Technical Report

We present KAT-Coder-V2, an agentic coding model developed by the KwaiKAT team at Kuaishou. KAT-Coder-V2 adopts a "Specialize-then-Unify" paradigm that decomposes agentic coding into five expert domains - SWE, WebCoding, Terminal, WebSearch, and General - each undergoing independent supervised fine-tuning and reinforcement learning, before being consolidated into a single model via on-policy distillation. We develop KwaiEnv, a modular infrastructure sustaining tens of thousands of concurrent sandbox instances, and scale RL training along task complexity, intent alignment, and scaffold generalization. We further propose MCLA for stabilizing MoE RL training and Tree Training for eliminating redundant computation over tree-structured trajectories with up to 6.2x speedup. KAT-Coder-V2 achieves 79.6% on SWE-bench Verified (vs. Claude Opus 4.6 at 80.8%), 88.7 on PinchBench (surpassing GLM-5 and MiniMax M2.7), ranks first across all three frontend aesthetics scenarios, and maintains strong generalist scores on Terminal-Bench Hard (46.8) and tau^2-Bench (93.9). Our model is publicly available at https://streamlake.com/product/kat-coder.