NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2025-10-17中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Asking AI to build scrapers should be easy right?

The article titled 'Asking AI to build scrapers should be easy right?' explores the common misconception that leveraging artificial intelligence for web scraping is a straightforward task. While the promise of AI agents automating data extraction is appealing, the reality presents significant challenges. Websites are often dynamic, employing complex structures, JavaScript rendering, and anti-bot measures such as CAPTCHAs, which current general-purpose AI models, particularly Large Language Models (LLMs), struggle to navigate reliably. The core issue lies in the robustness and adaptability required for scrapers to handle frequent website changes, session management, and diverse data formats without constant human intervention. Effective AI-driven scraping necessitates sophisticated AI agents capable of understanding context, making decisions, and autonomously adapting to unforeseen variations in web interfaces. This discussion highlights the technical hurdles in developing truly autonomous and resilient web scraping solutions, indicating a gap between theoretical AI capabilities and practical, scalable deployment in dynamic web environments.

02

OpenAI Needs $400B In The Next 12 Months

OpenAI, a prominent leader in artificial intelligence research and development, is reportedly anticipating an unprecedented financial requirement of up to $400 billion over the next 12 months to fund its ambitious endeavors. This colossal figure highlights the escalating costs associated with pioneering advancements in the AI sector, particularly for companies operating at the forefront of large language model development. The projected capital would be critical for several key areas, including the massive procurement of advanced computing infrastructure, such as state-of-the-art GPUs, essential for training and deploying increasingly complex AI models. Additionally, these funds would support extensive research initiatives, attract top-tier AI talent, and facilitate the global scaling of its innovative AI products and services. The scale of this investment reflects the intense competition and significant infrastructure demands inherent in pushing the boundaries of artificial intelligence, as well as OpenAI's strategic intent to maintain its leadership position in the global AI race.

03

Claude Skills are awesome, maybe a bigger deal than MCP

The recent introduction of 'Claude Skills' by Anthropic marks a pivotal advancement in the field of artificial intelligence, with some observers suggesting its potential impact could exceed that of previously established benchmarks or paradigms, colloquially referred to as 'MCP'. These new capabilities are anticipated to significantly enhance Claude's ability to perform complex, multi-step tasks, utilize external tools more effectively, and interact with dynamic environments with greater autonomy. The concept of 'Skills' likely refers to a sophisticated framework allowing the AI to learn, adapt, and execute specialized functions, moving beyond conventional conversational or generative tasks. This development points towards a future where large language models are not merely predictive text engines but highly capable AI agents, capable of orchestrating sophisticated workflows and solving problems that require deep integration with real-world systems. The discussion around its significance highlights a growing trend towards more agentic AI, emphasizing practical utility and robust problem-solving, potentially redefining the landscape of AI applications and human-AI collaboration.

04

Andrej Karpathy – AGI is still a decade away

Prominent AI researcher Andrej Karpathy recently articulated his perspective on the timeline for achieving Artificial General Intelligence (AGI), suggesting that such a milestone remains approximately a decade in the future. This assessment implies that despite rapid advancements in specialized AI domains, fundamental breakthroughs are still required to reach a level of machine intelligence capable of performing any intellectual task that a human can. Karpathy's view likely stems from an analysis of current AI limitations, including challenges in common-sense reasoning, true understanding, and the ability to generalize across vastly different contexts without extensive retraining. His prediction offers a counterpoint to more optimistic timelines, emphasizing the complexity of mimicking human-level cognitive flexibility and adaptability. It underscores the ongoing need for significant research and development in areas beyond current narrow AI capabilities to bridge the gap towards truly general intelligence, influencing the strategic direction of future AI investment and research priorities across the industry.

05

Show HN: We packaged an MCP server inside Chromium

BrowserOS, a YC startup (S24) developing an open-source Chromium fork, has announced the integration of an MCP (Machine Control Protocol) server directly into its browser binary. Positioned as a privacy-first alternative to emerging AI browsers, this feature responds to significant user demand. Unlike Google's `chrome-devtools-mcp`, which requires external setup and utilizes fresh headless instances, BrowserOS's approach simplifies the developer experience. The direct packaging offers three main benefits: streamlined setup without `npx install` or specific Chrome DevTools Protocol (CDP) flags, enabling AI agents to interact with a user's currently logged-in browser sessions, and generally enhancing the efficiency of web application development and debugging through AI assistance within a familiar, privacy-focused environment. This innovation aims to provide a more integrated and user-friendly platform for leveraging AI agents.

06

AI has a cargo cult problem

The article, titled "AI has a cargo cult problem," identifies a critical issue within the artificial intelligence community: the tendency for practitioners to adopt and apply popular AI models and methodologies by merely mimicking their superficial characteristics, rather than deeply understanding their theoretical underpinnings, operational mechanisms, and inherent limitations. This phenomenon, analogous to a cargo cult, can lead to the widespread deployment of AI solutions based on observed empirical success in narrow contexts, often without sufficient scrutiny regarding their broader applicability, ethical implications, or robustness across varying data distributions. Such an approach risks the development of brittle, uninterpretable, and potentially unreliable AI systems, ultimately impeding genuine innovation and responsible progress in the field. The piece underscores the imperative for a more rigorous, principle-driven approach to AI development, advocating for a transition from simply replicating current trends to fostering a comprehensive understanding of AI's intricate complexities to build more resilient, accountable, and impactful technologies.

GitHub

2 stories
01

MiniMind

MiniMind is an open-source project focused on enabling the training of ultra-small language models (LLMs) from scratch with minimal resources – as little as $0.50 and 2 hours on a single GPU. It offers a comprehensive, simplified framework for building LLMs, including native PyTorch implementations of essential components like tokenizers, pre-training, supervised fine-tuning (SFT), LoRA, Direct Preference Optimization (DPO), and model distillation. The project also extends to multimodal capabilities with MiniMind-V and provides cleaned, high-quality datasets. Designed as both a full-stage LLM reproduction and a learning tutorial, MiniMind aims to democratize LLM development, allowing individuals to understand and build models from the ground up, fostering broader AI community engagement and innovation.

02

HuLa: A Real-time Communication System built with Tauri, Vite 7, Vue 3, and TypeScript

HuLa is a robust, cross-platform instant messaging system engineered with modern web technologies including Tauri, Vite 7, Vue 3, and TypeScript. It offers a comprehensive suite of communication features, spanning individual and group chats, message recall, @mentions, read statuses, emojis, and message reactions. The system prioritizes user experience with a modern UI design, dark/light themes, and customizable skins. Core functionalities also extend to social management, enabling friend additions/deletions, group creation, and online status tracking. HuLa's technical foundation leverages Tauri for a lightweight, high-performance desktop container, Vue 3 for a responsive user interface, Vite 7 for rapid development, and TypeScript for enhanced type safety, making it an efficient and secure solution. Furthermore, the system supports various platforms including Windows, macOS, Linux, iOS, and Android, and integrates AI features such as an AI chat assistant, designed to enhance user interaction and offer multi-platform AI support.

huggingface

6 stories
01

Attention Is All You Need for KV Cache in Diffusion LLMs

This work studies how to adaptively recompute key-value (KV) caches for diffusion large language models (DLMs) to maximize prediction accuracy while minimizing decoding latency. Prior methods' decoders recompute QKV for all tokens at every denoising step and layer, despite KV states changing little across most steps, especially in shallow layers, leading to substantial redundancy. We make three observations: (1) distant MASK tokens primarily act as a length-bias and can be cached block-wise beyond the active prediction window; (2) KV dynamics increase with depth, suggesting that selective refresh starting from deeper layers is sufficient; and (3) the most-attended token exhibits the smallest KV drift, providing a conservative lower bound on cache change for other tokens. Building on these, we propose Elastic-Cache, a training-free, architecture-agnostic strategy that jointly decides when to refresh (via an attention-aware drift test on the most-attended token) and where to refresh (via a depth-aware schedule that recomputes from a chosen layer onward while reusing shallow-layer caches and off-window MASK caches). Unlike fixed-period schemes, Elastic-Cache performs adaptive, layer-aware cache updates for diffusion LLMs, reducing redundant computation and accelerating decoding with negligible loss in generation quality. Experiments on LLaDA-Instruct, LLaDA-1.5, and LLaDA-V across mathematical reasoning and code generation tasks demonstrate consistent speedups: 8.7times on GSM8K (256 tokens), 45.1times on longer sequences, and 4.8times on HumanEval, while consistently maintaining higher accuracy than the baseline. Our method achieves significantly higher throughput (6.8times on GSM8K) than existing confidence-based approaches while preserving generation quality, enabling practical deployment of diffusion LLMs.

02

When Models Lie, We Learn: Multilingual Span-Level Hallucination Detection with PsiloQA

Hallucination detection remains a fundamental challenge for the safe and reliable deployment of large language models (LLMs), especially in applications requiring factual accuracy. Existing hallucination benchmarks often operate at the sequence level and are limited to English, lacking the fine-grained, multilingual supervision needed for a comprehensive evaluation. In this work, we introduce PsiloQA, a large-scale, multilingual dataset annotated with span-level hallucinations across 14 languages. PsiloQA is constructed through an automated three-stage pipeline: generating question-answer pairs from Wikipedia using GPT-4o, eliciting potentially hallucinated answers from diverse LLMs in a no-context setting, and automatically annotating hallucinated spans using GPT-4o by comparing against golden answers and retrieved context. We evaluate a wide range of hallucination detection methods -- including uncertainty quantification, LLM-based tagging, and fine-tuned encoder models -- and show that encoder-based models achieve the strongest performance across languages. Furthermore, PsiloQA demonstrates effective cross-lingual generalization and supports robust knowledge transfer to other benchmarks, all while being significantly more cost-efficient than human-annotated datasets. Our dataset and results advance the development of scalable, fine-grained hallucination detection in multilingual settings.

03

Agentic Entropy-Balanced Policy Optimization

Recently, Agentic Reinforcement Learning (Agentic RL) has made significant progress in incentivizing the multi-turn, long-horizon tool-use capabilities of web agents. While mainstream agentic RL algorithms autonomously explore high-uncertainty tool-call steps under the guidance of entropy, excessive reliance on entropy signals can impose further constraints, leading to the training collapse. In this paper, we delve into the challenges caused by entropy and propose the Agentic Entropy-Balanced Policy Optimization (AEPO), an agentic RL algorithm designed to balance entropy in both the rollout and policy update phases. AEPO comprises two core components: (1) a dynamic entropy-balanced rollout mechanism that adaptively allocate global and branch sampling budget through entropy pre-monitoring, while imposing a branch penalty on consecutive high-entropy tool-call steps to prevent over-branching issues; and (2) Entropy-Balanced Policy Optimization that inserts a stop-gradient operation into the high-entropy clipping term to preserve and properly rescale gradients on high-entropy tokens, while incorporating entropy-aware advantage estimation to prioritize learning on high-uncertainty tokens. Results across 14 challenging datasets show that AEPO consistently outperforms 7 mainstream RL algorithms. With just 1K RL samples, Qwen3-14B with AEPO achieves impressive results: 47.6% on GAIA, 11.2% on Humanity's Last Exam, and 43.0% on WebWalker for Pass@1; 65.0% on GAIA, 26.0% on Humanity's Last Exam, and 70.0% on WebWalker for Pass@5. Further analysis reveals that AEPO improves rollout sampling diversity while maintaining stable policy entropy, facilitating scalable web agent training.

04

VLA^2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation

Current vision-language-action (VLA) models, pre-trained on large-scale robotic data, exhibit strong multi-task capabilities and generalize well to variations in visual and language instructions for manipulation. However, their success rate drops significantly when faced with object concepts outside the training data, such as unseen object descriptions and textures in the dataset. To address this, we propose a novel agentic framework, VLA^2, which leverages OpenVLA as the execution backbone and effectively leverages external modules such as web retrieval and object detection to provide visual and textual knowledge about target objects to the VLA. This approach mitigates generalization failure when handling out-of-distribution objects. Based on the LIBERO simulation environment, we introduced novel objects and object descriptions to construct a new evaluation benchmark with three difficulty levels to test the effectiveness of our method. Our framework successfully outperformed the current state-of-the-art models on our designed hard-level generalization benchmark. Compared to the standalone OpenVLA baseline, VLA^2 achieves a 44.2% improvement in the success rate in the hard-level benchmark and an average improvement of 20.2% in all customized environments without any performance degradation on in-domain tasks. Project website: https://vla-2.github.io.

05

ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints

Video generation models have achieved remarkable progress, particularly excelling in realistic scenarios; however, their performance degrades notably in imaginative scenarios. These prompts often involve rarely co-occurring concepts with long-distance semantic relationships, falling outside training distributions. Existing methods typically apply test-time scaling for improving video quality, but their fixed search spaces and static reward designs limit adaptability to imaginative scenarios. To fill this gap, we propose ImagerySearch, a prompt-guided adaptive test-time search strategy that dynamically adjusts both the inference search space and reward function according to semantic relationships in the prompt. This enables more coherent and visually plausible videos in challenging imaginative settings. To evaluate progress in this direction, we introduce LDT-Bench, the first dedicated benchmark for long-distance semantic prompts, consisting of 2,839 diverse concept pairs and an automated protocol for assessing creative generation capabilities. Extensive experiments show that ImagerySearch consistently outperforms strong video generation baselines and existing test-time scaling approaches on LDT-Bench, and achieves competitive improvements on VBench, demonstrating its effectiveness across diverse prompt types. We will release LDT-Bench and code to facilitate future research on imaginative video generation.

06

WithAnyone: Towards Controllable and ID Consistent Image Generation

Identity-consistent generation has become an important focus in text-to-image research, with recent models achieving notable success in producing images aligned with a reference identity. Yet, the scarcity of large-scale paired datasets containing multiple images of the same individual forces most approaches to adopt reconstruction-based training. This reliance often leads to a failure mode we term copy-paste, where the model directly replicates the reference face rather than preserving identity across natural variations in pose, expression, or lighting. Such over-similarity undermines controllability and limits the expressive power of generation. To address these limitations, we (1) construct a large-scale paired dataset MultiID-2M, tailored for multi-person scenarios, providing diverse references for each identity; (2) introduce a benchmark that quantifies both copy-paste artifacts and the trade-off between identity fidelity and variation; and (3) propose a novel training paradigm with a contrastive identity loss that leverages paired data to balance fidelity with diversity. These contributions culminate in WithAnyone, a diffusion-based model that effectively mitigates copy-paste while preserving high identity similarity. Extensive qualitative and quantitative experiments demonstrate that WithAnyone significantly reduces copy-paste artifacts, improves controllability over pose and expression, and maintains strong perceptual quality. User studies further validate that our method achieves high identity fidelity while enabling expressive controllable generation.