NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-10-20ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Code on the Web

Anthropic has introduced "Claude Code on the Web," a new initiative designed to significantly enhance the coding and web interaction capabilities of its Claude large language model. This advancement focuses on enabling Claude to more effectively understand, generate, and potentially execute code within web-based environments. The development positions Claude as a more sophisticated tool for developers and users requiring automated web tasks, complex code solutions, or dynamic content generation for online platforms. By integrating robust coding functions with web-specific interaction, Anthropic aims to bridge the gap between high-level natural language instructions and practical web operations. This move underscores the company's commitment to evolving AI agents that can seamlessly navigate and contribute to complex digital ecosystems, thereby expanding the practical applications of large language models in real-world scenarios. The announcement signifies a strategic effort to enhance Claude's versatility for programming and web-related challenges.

02

BERT is just a single text diffusion step

A recent analysis proposes a novel conceptual framework, asserting that the widely-used BERT (Bidirectional Encoder Representations from Transformers) model can be understood as performing a single text diffusion step. This perspective re-contextualizes BERT's masked language modeling objective, where the model predicts masked tokens based on surrounding context, as analogous to a denoising operation characteristic of diffusion models. In this view, the input text, with its masked tokens, is considered a 'noisy' version of the original, and BERT's task is to 'denoise' it by reconstructing the missing information. This reframing offers significant theoretical implications, potentially bridging the gap between discriminative pre-training objectives like those in BERT and the generative mechanisms found in diffusion-based models. It suggests that the power of masked language models might stem from their inherent ability to perform a robust single-step generative denoising process. Such an interpretation could lead to new architectural designs, training methodologies, or deeper insights into the emergent generative capabilities of Transformer-based models, fostering advancements in natural language understanding and generation by unifying these previously distinct paradigms.

03

Production RAG: what I learned from processing 5M+ documents

An analysis of insights gained from developing and deploying a Production Retrieval Augmented Generation (RAG) system, detailing lessons learned from processing over five million documents. The article highlights critical considerations for scaling RAG solutions, emphasizing challenges related to data ingestion, embedding generation, and efficient retrieval across large corpuses. It covers strategies for optimizing vector database performance, ensuring data quality, and managing infrastructure to support high-volume operations. Key takeaways include best practices for maintaining a robust RAG pipeline, improving retrieval accuracy, and integrating seamlessly with large language models in real-world applications. The learnings are crucial for engineers and data scientists building scalable and reliable AI systems that leverage external knowledge bases.

04

Show HN: Playwright Skill for Claude Code – Less context than playwright-MCP

A new Playwright Skill has been developed for Claude Code, specifically designed to address the high token consumption issue experienced with `playwright-mcp` within Claude's 200K token limit. This innovative solution operates by enabling Claude to directly write and execute Playwright code for browser automation, significantly reducing the amount of context exchanged. Unlike `playwright-mcp`, which sends extensive accessibility tree snapshots for every action, this skill provides only screenshots and console output in return. This approach drastically minimizes overhead, leveraging just 314 lines of instructions compared to managing a persistent `playwright-mcp` server. The system efficiently loads full Playwright API documentation only when required, ensuring optimized resource usage while maintaining comparable browser automation capabilities. It is available as a Claude Code plugin or for manual installation, offering a more efficient alternative for developers working with large language models and web automation tasks, as highlighted by a reported token limit issue with `playwright-mcp`.

05

Modeling Others' Minds as Code

This paper introduces a novel approach to understanding and predicting the behavior of intelligent agents, whether human or artificial, by conceptualizing their cognitive processes as executable code. The core idea involves developing computational frameworks that can model the "mind" of an agent by representing its internal states, reasoning mechanisms, and decision-making logic as programmatic constructs. This paradigm shift offers significant potential for advancements in AI agents' ability to develop sophisticated "theory of mind" capabilities, enabling them to better anticipate actions, infer intentions, and strategically interact with other entities in complex environments. By treating mental models as code, researchers aim to create more interpretable, debuggable, and adaptable AI systems capable of robust social and collaborative intelligence. This research has implications for areas such as human-computer interaction, multi-agent systems, and the development of more empathetic and effective AI assistants. The methodology promises to enhance an AI's capacity for complex social reasoning and interaction.

06

The FTC Is Disappearing Blog Posts About AI Published During Lina Khan's Tenure

The Federal Trade Commission (FTC) is reportedly removing blog posts related to artificial intelligence (AI) that were published during the tenure of its current chair, Lina Khan. This action, highlighted by a Wired report, raises questions regarding the agency's transparency and consistency in its public stance on AI regulation and policy. The disappearing content, which previously offered insights into the FTC's perspective on the rapidly evolving AI landscape, could signal a recalibration of the agency's approach to emerging technologies. This move has potential implications for how businesses and researchers interpret future regulatory enforcement and guidance in the AI space. Stakeholders are observing whether this reflects a strategic shift in the FTC's focus or an internal review process, particularly concerning critical areas such as data privacy, algorithmic bias, and market competition, which fall under the FTC's broad mandate. The lack of clear communication surrounding these removals contributes to an environment of uncertainty regarding the future direction of AI governance and consumer protection under the current administration, prompting calls for greater clarity from the federal regulator.

GitHub

5 stories
01

Claude Cookbooks

The Claude Cookbooks offer a comprehensive collection of code and guides for developers building applications with Anthropic's Claude AI. Designed for easy integration, the resource provides copy-able Python code snippets, though concepts are adaptable to any programming language supporting the Claude API. Key capabilities explored include text classification, retrieval-augmented generation (RAG), and summarization. The cookbooks also delve into advanced techniques such as integrating Claude with external tools for functions like customer service or SQL queries, and facilitating third-party integrations with vector databases (e.g., Pinecone), Wikipedia, and web search services. Furthermore, it covers multimodal functionalities like vision for image analysis and chart interpretation, as well as image generation via integration with Stable Diffusion. Developers can also find guides on advanced topics like sub-agents, PDF processing, automated prompt evaluations, JSON mode, moderation filters, and prompt caching, making it a vital resource for enhancing Claude-powered solutions.

02

System Prompts and Models of AI Tools

This repository, "System Prompts and Models of AI Tools," provides an extensive collection of over 30,000 lines of insights into the structure and functionality of various AI system prompts and models. It serves as a valuable resource for developers focused on building reliable AI agents and prompts, offering an understanding of underlying mechanisms. The project is actively updated, with new instructions often released early via Discord. It also highlights an open-source AI engineering platform, Latitude, and features a security notice for AI startups, advocating for robust data protection against exposed prompts and model configurations through services like ZeroLeaks, an AI security audit platform. The collection encompasses prompts and tools from numerous AI agents and platforms, promoting knowledge sharing within the AI development community.

03

LeRobot: State-of-the-art AI for real-world robotics

LeRobot is a Hugging Face library designed to democratize real-world robotics by providing state-of-the-art models, datasets, and tools within the PyTorch ecosystem. It aims to lower the barrier to entry for robotics development, enabling wider contributions and sharing of pre-trained models and datasets. The library focuses on imitation learning and reinforcement learning approaches, which have demonstrated effective transfer to real-world applications. LeRobot currently offers pre-trained models, human-collected demonstration datasets, and simulation environments, allowing users to start without needing physical robot hardware. Future plans include expanding support for affordable and capable real-world robots. The project also provides tutorials for building specific robots like the dexterous HopeJR and the cost-effective SO-101, alongside tools for dataset visualization and reproducing state-of-the-art policy training.

04

sing-box

Sing-box stands as a universal proxy platform, engineered to facilitate comprehensive network traffic management. This robust solution empowers users with the capability to route internet traffic via various proxy protocols, thereby bolstering privacy, security, and accessibility. While its core function is to establish a flexible and adaptable proxy infrastructure, its sponsorship highlights a specific application within advanced development environments, particularly for "coding with multiple AI agents." This suggests its potential utility in orchestrating network interactions for AI-driven processes. Operating under the GNU General Public License v3 or later, Sing-box underscores its commitment to open-source principles. Detailed documentation is readily available at sing-box.sagernet.org, offering extensive guidance for deployment and configuration, positioning it as a versatile tool for both individual users and developers requiring sophisticated network proxy solutions.

05

ebook2audiobook

"ebook2audiobook" is a versatile CPU/GPU-accelerated converter designed to transform legally acquired eBooks into audiobooks, complete with chapters and metadata. It integrates advanced Text-to-Speech (TTS) engines, including XTTSv2, Bark, Vits, Fairseq, YourTTS, and Tacotron, supporting an impressive range of over 1110 languages. A key feature is its optional voice cloning capability, allowing users to personalize the audiobook narration with their own voice. The tool is engineered for efficiency, capable of running on systems with as little as 4GB RAM. It supports a wide array of eBook formats such as `.epub`, `.pdf`, and `.mobi`, and outputs audiobooks in various formats like `m4b`, `mp3`, and `wav`. Users can operate it via a Gradio web interface, headless command-line mode, or Docker, with robust support for both local and remote deployment environments like Hugging Face Spaces and Google Colab. The project emphasizes responsible use and provides extensive documentation for installation, usage, fine-tuning TTS models, and troubleshooting common issues.

huggingface

7 stories
01

OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding

Advancing machine intelligence requires developing the ability to perceive across multiple modalities, much as humans sense the world. We introduce OmniVinci, an initiative to build a strong, open-source, omni-modal LLM. We carefully study the design choices across model architecture and data curation. For model architecture, we present three key innovations: (i) OmniAlignNet for strengthening alignment between vision and audio embeddings in a shared omni-modal latent space; (ii) Temporal Embedding Grouping for capturing relative temporal alignment between vision and audio signals; and (iii) Constrained Rotary Time Embedding for encoding absolute temporal information in omni-modal embeddings. We introduce a curation and synthesis pipeline that generates 24M single-modal and omni-modal conversations. We find that modalities reinforce one another in both perception and reasoning. Our model, OmniVinci, outperforms Qwen2.5-Omni with +19.05 on DailyOmni (cross-modal understanding), +1.7 on MMAR (audio), and +3.9 on Video-MME (vision), while using just 0.2T training tokens - a 6 times reduction compared to Qwen2.5-Omni's 1.2T. We finally demonstrate omni-modal advantages in downstream applications spanning robotics, medical AI, and smart factory.

02

Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning, these findings were limited to finetuning and activation steering, leaving out in-context learning (ICL). We therefore ask: does EM emerge in ICL? We find that it does: across three datasets, three frontier models produce broadly misaligned responses at rates between 2% and 17% given 64 narrow in-context examples, and up to 58% with 256 examples. We also examine mechanisms of EM by eliciting step-by-step reasoning (while leaving in-context examples unchanged). Manual analysis of the resulting chain-of-thought shows that 67.5% of misaligned traces explicitly rationalize harmful outputs by adopting a reckless or dangerous ''persona'', echoing prior results on finetuning-induced EM.

03

Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset

Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this fundamental challenge. At its heart, Ditto features a novel data generation pipeline that fuses the creative diversity of a leading image editor with an in-context video generator, overcoming the limited scope of existing models. To make this process viable, our framework resolves the prohibitive cost-quality trade-off by employing an efficient, distilled model architecture augmented by a temporal enhancer, which simultaneously reduces computational overhead and improves temporal coherence. Finally, to achieve full scalability, this entire pipeline is driven by an intelligent agent that crafts diverse instructions and rigorously filters the output, ensuring quality control at scale. Using this framework, we invested over 12,000 GPU-days to build Ditto-1M, a new dataset of one million high-fidelity video editing examples. We trained our model, Editto, on Ditto-1M with a curriculum learning strategy. The results demonstrate superior instruction-following ability and establish a new state-of-the-art in instruction-based video editing.

04

A^2FM: An Adaptive Agent Foundation Model for Tool-Aware Hybrid Reasoning

Large language models split into two families: reasoning-centric LLMs, which strengthen internal chain-of-thought reasoning but cannot invoke external tools, and agentic LLMs, which learn to interact with environments and leverage tools but often lag in deep reasoning. This divide arises from fundamentally different training objectives, leading to mismatched strengths and inefficiency on simple queries, where both families tend to overthink or over-call tools. In this work, we present Adaptive Agent Foundation Model (A^2FM), a unified framework that follows a route-then-align principle: the model first learns task-aware routing and then aligns mode-specific trajectories under a shared backbone. To address the inefficiency gap, we introduce a third mode-instant-that handles simple queries directly, preventing unnecessary reasoning or tool calls while complementing the agentic and reasoning modes. To jointly enhance accuracy and efficiency, we propose Adaptive Policy Optimization (APO), which enforces adaptive sampling across modes and applies a cost-regularized reward. On the 32B scale, A^2FM achieves 13.4% on BrowseComp, 70.4% on AIME25, and 16.7% on HLE, setting new SOTA among comparable models and performing competitively with frontier LLMs across agentic, reasoning, and general benchmarks. Notably, the adaptive execution achieves a cost of pass of only $0.00487 per correct answer-cutting cost by 45.2% relative to reasoning and 33.5% relative to agentic, thus delivering substantially higher cost efficiency while maintaining comparable accuracy.

05

BLIP3o-NEXT: Next Frontier of Native Image Generation

We present BLIP3o-NEXT, a fully open-source foundation model in the BLIP3 series that advances the next frontier of native image generation. BLIP3o-NEXT unifies text-to-image generation and image editing within a single architecture, demonstrating strong image generation and image editing capabilities. In developing the state-of-the-art native image generation model, we identify four key insights: (1) Most architectural choices yield comparable performance; an architecture can be deemed effective provided it scales efficiently and supports fast inference; (2) The successful application of reinforcement learning can further push the frontier of native image generation; (3) Image editing still remains a challenging task, yet instruction following and the consistency between generated and reference images can be significantly enhanced through post-training and data engine; (4) Data quality and scale continue to be decisive factors that determine the upper bound of model performance. Building upon these insights, BLIP3o-NEXT leverages an Autoregressive + Diffusion architecture in which an autoregressive model first generates discrete image tokens conditioned on multimodal inputs, whose hidden states are then used as conditioning signals for a diffusion model to generate high-fidelity images. This architecture integrates the reasoning strength and instruction following of autoregressive models with the fine-detail rendering ability of diffusion models, achieving a new level of coherence and realism. Extensive evaluations of various text-to-image and image-editing benchmarks show that BLIP3o-NEXT achieves superior performance over existing models.

06

VISTA: A Test-Time Self-Improving Video Generation Agent

Despite rapid advances in text-to-video synthesis, generated video quality remains critically dependent on precise user prompts. Existing test-time optimization methods, successful in other domains, struggle with the multi-faceted nature of video. In this work, we introduce VISTA (Video Iterative Self-improvemenT Agent), a novel multi-agent system that autonomously improves video generation through refining prompts in an iterative loop. VISTA first decomposes a user idea into a structured temporal plan. After generation, the best video is identified through a robust pairwise tournament. This winning video is then critiqued by a trio of specialized agents focusing on visual, audio, and contextual fidelity. Finally, a reasoning agent synthesizes this feedback to introspectively rewrite and enhance the prompt for the next generation cycle. Experiments on single- and multi-scene video generation scenarios show that while prior methods yield inconsistent gains, VISTA consistently improves video quality and alignment with user intent, achieving up to 60% pairwise win rate against state-to-the-art baselines. Human evaluators concur, preferring VISTA outputs in 66.4% of comparisons.

07

DLER: Doing Length pEnalty Right - Incentivizing More Intelligence per Token via Reinforcement Learning

Reasoning language models such as OpenAI-o1, DeepSeek-R1, and Qwen achieve strong performance via extended chains of thought but often generate unnecessarily long outputs. Maximizing intelligence per token--accuracy relative to response length--remains an open problem. We revisit reinforcement learning (RL) with the simplest length penalty--truncation--and show that accuracy degradation arises not from the lack of sophisticated penalties but from inadequate RL optimization. We identify three key challenges: (i) large bias in advantage estimation, (ii) entropy collapse, and (iii) sparse reward signal. We address them with Doing Length pEnalty Right (DLER), a training recipe combining batch-wise reward normalization, higher clipping, dynamic sampling, and a simple truncation length penalty. DLER achieves state-of-the-art accuracy--efficiency trade-offs, cutting output length by over 70 percent while surpassing all previous baseline accuracy. It also improves test-time scaling: compared to DeepSeek-R1-7B, DLER-7B generates multiple concise responses in parallel with 28 percent higher accuracy and lower latency. We further introduce Difficulty-Aware DLER, which adaptively tightens truncation on easier questions for additional efficiency gains. We also propose an update-selective merging method that preserves baseline accuracy while retaining the concise reasoning ability of the DLER model, which is useful for scenarios where RL training data is scarce.