NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-27中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Show HN: Open-Source Animal Crossing–Style UI for Claude Code Agents

Outworked is an open-source desktop application that provides an Animal Crossing-style user interface for managing Claude Code agents. This innovative platform allows a team of AI agents to collaboratively work towards a defined goal, where an orchestrator breaks down complex tasks and assigns them for parallel execution. Agents possess the capability to communicate with each other, write code, and leverage the web. Recent significant updates introduce iMessage channel support, enabling agents to text people and vice versa, a built-in browser for navigating and interacting with websites, and task scheduling for automated operations. Furthermore, the system includes built-in tunneling for sharing local resources over the internet, more robust support for custom skills (MCP), and auto-approval for agent requests, enhancing its extensibility and ease of use. Designed to run locally and integrate seamlessly with existing Claude Code setups, Outworked offers a practical and user-friendly environment for leveraging advanced AI agent capabilities.

02

Anatomy of the .claude/ folder

The article "Anatomy of the .claude/ folder" offers a detailed technical examination of the directory structure and contents associated with the Claude AI assistant, focusing on its local operational footprint. This deep dive illuminates how the `.claude/` folder manages various aspects of user interaction and model functionality. Typically, such a directory would house critical configuration files, session history, cached responses, and potentially logs of API calls or user prompts, enabling persistent context and efficient operation for applications leveraging Claude. Understanding this folder's layout is crucial for developers and power users aiming to debug issues, manage data privacy, or customize their AI agent's behavior. The analysis likely covers the purpose of key subdirectories and files, shedding light on data storage practices, potential areas for optimization, and security considerations related to locally stored sensitive information. This detailed anatomical view serves as an essential guide for anyone working with or analyzing the backend operations of AI agents powered by Anthropic's Claude.

03

Some uncomfortable truths about AI coding agents

The article 'Some uncomfortable truths about AI coding agents' critically examines the practical challenges and current limitations associated with the deployment of artificial intelligence tools in software development. It addresses the discrepancy between the often-hyped potential of AI coding agents and their real-world efficacy, highlighting several 'uncomfortable truths' for developers and organizations. These include the significant human oversight and debugging efforts still required to validate and refine AI-generated code, the agents' inherent difficulties in handling complex system architectures or nuanced problem specifications, and the potential for introducing subtle errors or inefficiencies that can be difficult to trace. The piece likely concludes that while AI coding agents offer valuable assistance in specific, well-defined coding tasks, they are not yet capable of fully autonomous development and necessitate a realistic appreciation of their current imperfections and the continued indispensable role of human expertise in the software engineering lifecycle. This perspective challenges an overly optimistic view, advocating for a balanced understanding of AI's role in coding.

04

Schedule tasks on the web

The document "Schedule tasks on the web" from code.claude.com likely details methodologies and tools for automating and managing recurring or time-sensitive operations within web-based applications, particularly those integrated with or supporting Artificial Intelligence functionalities. This capability is crucial for orchestrating complex AI workflows, such as periodic data collection, model training, inference job execution, and reporting. Implementing robust web task scheduling ensures that AI systems operate efficiently and reliably, minimizing manual intervention and maximizing throughput. The documentation probably covers various approaches, including server-side cron jobs, cloud-based scheduling services, or platform-specific APIs offered by Claude for seamless integration. Effective task scheduling is essential for maintaining the operational health and scalability of AI-driven web services, allowing developers to define precise execution windows, handle retries, and monitor task statuses, thereby contributing to the overall stability and performance of AI applications.

05

Hold on to Your Hardware

The article "Hold on to Your Hardware" advocates for a re-evaluation of the strategic importance of maintaining and leveraging local computing infrastructure amidst the prevailing trend towards cloud-centric solutions. It posits that while cloud services offer scalability and flexibility, there are compelling arguments for retaining and investing in physical hardware assets. This perspective highlights the long-term value derived from on-premise data processing, which can lead to greater cost predictability, enhanced data security, and reduced latency for critical operations. The piece likely delves into the benefits of owning infrastructure for specialized computational tasks, such as high-performance computing, large-scale data analytics, and the training of complex AI models, where specific hardware configurations and direct control over resources are paramount. It encourages a balanced approach, suggesting that foundational hardware investment can strategically complement cloud services, thereby ensuring resilience, optimizing performance, and providing greater autonomy for organizations facing evolving technological and economic landscapes. This discussion is particularly relevant as industries grapple with data sovereignty, operational continuity, and the efficiency of resource allocation.

06

OpenAI's US ad pilot exceeds $100M in annualized revenue in six weeks

OpenAI has achieved significant financial success with its US ad pilot program, reportedly surpassing an annualized revenue run rate of $100 million within just six weeks of its launch. This rapid growth underscores the company's aggressive strategy to monetize its advanced artificial intelligence capabilities and diversify its revenue streams beyond subscription services and API access. The pilot's strong performance highlights the commercial viability of integrating AI-powered advertising solutions into mainstream markets, suggesting a burgeoning demand for sophisticated, data-driven ad platforms. This milestone positions OpenAI as a formidable player not only in AI research and development but also in the competitive digital advertising landscape, signaling a potential shift in how AI technologies are leveraged for large-scale commercial applications and demonstrating the strong market adoption of OpenAI's offerings. The company's ability to quickly scale its ad pilot into a substantial revenue generator is a testament to its innovation and market penetration.

huggingface

6 stories
01

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.

02

MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance degradation as the number of input references grows. We identify the root cause as a fundamental data bottleneck: existing datasets are dominated by single- or few-reference pairs and lack the structured, long-context supervision needed to learn dense inter-reference dependencies. To address this, we introduce MacroData, a large-scale dataset of 400K samples, each containing up to 10 reference images, systematically organized across four complementary dimensions -- Customization, Illustration, Spatial reasoning, and Temporal dynamics -- to provide comprehensive coverage of the multi-reference generation space. Recognizing the concurrent absence of standardized evaluation protocols, we further propose MacroBench, a benchmark of 4,000 samples that assesses generative coherence across graded task dimensions and input scales. Extensive experiments show that fine-tuning on MacroData yields substantial improvements in multi-reference generation, and ablation studies further reveal synergistic benefits of cross-task co-training and effective strategies for handling long-context complexity. The dataset and benchmark will be publicly released.

03

AVControl: Efficient Framework for Training Audio-Visual Controls

Controlling video and audio generation requires diverse modalities, from depth and pose to camera trajectories and audio transformations, yet existing approaches either train a single monolithic model for a fixed set of controls or introduce costly architectural changes for each new modality. We introduce AVControl, a lightweight, extendable framework built on LTX-2, a joint audio-visual foundation model, where each control modality is trained as a separate LoRA on a parallel canvas that provides the reference signal as additional tokens in the attention layers, requiring no architectural changes beyond the LoRA adapters themselves. We show that simply extending image-based in-context methods to video fails for structural control, and that our parallel canvas approach resolves this. On the VACE Benchmark, we outperform all evaluated baselines on depth- and pose-guided generation, inpainting, and outpainting, and show competitive results on camera control and audio-visual benchmarks. Our framework supports a diverse set of independently trained modalities: spatially-aligned controls such as depth, pose, and edges, camera trajectory with intrinsics, sparse motion control, video editing, and, to our knowledge, the first modular audio-visual controls for a joint generation model. Our method is both compute- and data-efficient: each modality requires only a small dataset and converges within a few hundred to a few thousand training steps, a fraction of the budget of monolithic alternatives. We publicly release our code and trained LoRA checkpoints.

04

FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol

This paper introduces FinMCP-Bench, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of financial model context protocols. FinMCP-Bench contains 613 samples spanning 10 main scenarios and 33 sub-scenarios, featuring both real and synthetic user queries to ensure diversity and authenticity. It incorporates 65 real financial MCPs and three types of samples, single tool, multi-tool, and multi-turn, allowing evaluation of models across different levels of task complexity. Using this benchmark, we systematically assess a range of mainstream LLMs and propose metrics that explicitly measure tool invocation accuracy and reasoning capabilities. FinMCP-Bench provides a standardized, practical, and challenging testbed for advancing research on financial LLM agents.

05

Vega: Learning to Drive with Natural Language Instructions

Vision-language-action models have reshaped autonomous driving to incorporate languages into the decision-making process. However, most existing pipelines only utilize the language modality for scene descriptions or reasoning and lack the flexibility to follow diverse user instructions for personalized driving. To address this, we first construct a large-scale driving dataset (InstructScene) containing around 100,000 scenes annotated with diverse driving instructions with the corresponding trajectories. We then propose a unified Vision-Language-World-Action model, Vega, for instruction-based generation and planning. We employ the autoregressive paradigm to process visual inputs (vision) and language instructions (language) and the diffusion paradigm to generate future predictions (world modeling) and trajectories (action). We perform joint attention to enable interactions between the modalities and use individual projection layers for different modalities for more capabilities. Extensive experiments demonstrate that our method not only achieves superior planning performance but also exhibits strong instruction-following abilities, paving the way for more intelligent and personalized driving systems.

06

Reaching Beyond the Mode: RL for Distributional Reasoning in Language Models

Given a question, a language model (LM) implicitly encodes a distribution over possible answers. In practice, post-training procedures for LMs often collapse this distribution onto a single dominant mode. While this is generally not a problem for benchmark-style evaluations that assume one correct answer, many real-world tasks inherently involve multiple valid answers or irreducible uncertainty. Examples include medical diagnosis, ambiguous question answering, and settings with incomplete information. In these cases, we would like LMs to generate multiple plausible hypotheses, ideally with confidence estimates for each one, and without computationally intensive repeated sampling to generate non-modal answers. This paper describes a multi-answer reinforcement learning approach for training LMs to perform distributional reasoning over multiple answers during inference. We modify the RL objective to enable models to explicitly generate multiple candidate answers in a single forward pass, internalizing aspects of inference-time search into the model's generative process. Across question-answering, medical diagnostic, and coding benchmarks, we observe improved diversity, coverage, and set-level calibration scores compared to single answer trained baselines. Models trained with our approach require fewer tokens to generate multiple answers than competing approaches. On coding tasks, they are also substantially more accurate. These results position multi-answer RL as a principled and compute-efficient alternative to inference-time scaling procedures such as best-of-k. Code and more information can be found at https://multi-answer-rl.github.io/.