NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-09-26DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

SimpleFold: Folding proteins is simpler than you think

Apple's machine learning division has unveiled 'SimpleFold,' a novel initiative addressing the complex challenge of protein folding. This project, detailed in an arXiv publication (2509.18480) and supported by an open-source GitHub repository (github.com/apple/ml-simplefold), aims to simplify the intricate process of predicting three-dimensional protein structures. The title, 'Folding proteins is simpler than you think,' suggests that SimpleFold introduces innovative algorithms or a more streamlined methodology, departing from traditional, often computationally intensive, approaches. This development holds significant implications for various scientific fields, including drug discovery, biotechnology, and fundamental biological research, as accurate protein structure prediction is critical for understanding cellular functions and disease mechanisms. The release of SimpleFold demonstrates Apple's growing commitment to contributing to foundational scientific problems through advanced machine learning. By potentially making protein structure prediction more accessible and efficient, SimpleFold could accelerate research and development efforts globally, fostering new insights into biological processes and therapeutic innovations.

02

How to stop AI's "lethal trifecta"

This article from The Economist addresses the critical challenge of mitigating the most severe risks posed by advanced artificial intelligence, conceptualizing these threats as a "lethal trifecta." It likely discusses how to counter the three major dangers: the potential for autonomous weapon systems to escalate conflicts, the risk of AI-driven misinformation and deepfakes destabilizing societies, and the challenge of maintaining human control over increasingly powerful and opaque AI decision-making systems. The piece advocates for a multifaceted approach involving robust international governance frameworks, the implementation of stringent safety and ethical guidelines for AI development, and enhanced public-private sector collaboration. It emphasizes the urgent need for proactive regulatory measures, fostering research into explainable AI, and ensuring transparency to align AI systems with human values, thereby safeguarding global stability and preventing catastrophic outcomes associated with uncontrolled or misused artificial intelligence.

03

DeepFabric \u2013 Generate high-quality synthetic datasets at scale

DeepFabric introduces a novel approach for generating high-quality synthetic datasets at scale, addressing critical challenges in machine learning development such as data scarcity, privacy concerns, and bias. This technology provides a robust solution for developers and researchers to create diverse and representative synthetic data, which can significantly enhance the training, testing, and validation processes of AI models. By enabling the generation of data that mimics real-world distributions without exposing sensitive information, DeepFabric facilitates the development of more accurate and ethical AI systems. Its ability to operate at scale ensures that even large-scale projects can benefit from tailored synthetic data, thereby accelerating innovation and deployment in various domains. The system emphasizes data utility and fidelity, ensuring that the generated datasets are not only private but also effective substitutes for real data.

04

Context is the bottleneck for coding agents now

The assertion that "Context is the bottleneck for coding agents now" underscores a significant challenge in the ongoing advancement of AI-driven software development tools. Modern coding agents, predominantly powered by Large Language Models (LLMs), critically depend on the breadth and depth of contextual information to effectively interpret user requirements, navigate intricate codebases, debug complex issues, and generate precise, functional code. The inherent limitation of an LLM's 'context window'u2014the maximum token count it can process simultaneouslyu2014becomes a formidable barrier as software projects increase in scale and complexity. When the necessary contextual data, including extensive code files, architectural documentation, and historical discussions, surpasses this window, agent performance diminishes, leading to an increase in errors, a reduction in the quality of generated code, and difficulty in managing interdependencies within large systems. This bottleneck significantly curtails the practical scalability and overall efficacy of coding agents. Consequently, researchers and developers are actively exploring and implementing sophisticated strategies such as intelligent context summarization, advanced retrieval-augmented generation (RAG) techniques, and multi-turn, iterative prompting to mitigate these constraints. However, these workarounds often introduce additional complexity and processing overhead. Addressing and overcoming this fundamental context limitation is paramount for unlocking the next generation of truly autonomous and highly capable programming assistants, pushing the boundaries of what AI can achieve in software engineering.

05

Why AI systems may never be secure, and what to do about it

This article explores the inherent challenges in securing Artificial Intelligence systems, suggesting that achieving absolute security may be an unattainable goal due to their probabilistic and adaptive nature. It delves into the fundamental vulnerabilities that make AI models susceptible to a range of sophisticated threats, including adversarial attacks, data poisoning, model inversion, and membership inference. Traditional cybersecurity paradigms, designed for deterministic software, are often ill-equipped to address these novel attack vectors and the dynamic threat landscape presented by AI. The discussion highlights the critical need for a paradigm shift in AI security strategies. The piece advocates for a comprehensive approach centered on building resilient AI systems through robust architectural design, continuous monitoring for anomalous behavior, and the integration of explainable AI techniques to enhance transparency and detect malicious manipulations. It underscores the necessity for developers and organizations to adopt a proactive stance, focusing on risk mitigation, ethical AI development, and establishing robust governance frameworks to manage the persistent and evolving security risks associated with advanced AI technologies, rather than pursuing the illusion of perfect, impregnable security. This re-evaluation of security expectations is crucial for the safe and responsible deployment of AI.

06

Better health conversations: Research on a "wayfinding" AI agent based on Gemini

Google Research is actively investigating a novel "wayfinding" AI agent, powered by the Gemini large language model, with the goal of significantly improving health-related conversations. This research focuses on developing an intelligent system capable of guiding individuals through complex medical information, assisting them in navigating health resources, comprehending diagnoses, and making informed decisions about their well-being. The "wayfinding" paradigm emphasizes personalized and context-aware assistance, aiming to move beyond simplistic question-answering to provide structured pathways for users seeking health knowledge. Initial insights from this research are expected to demonstrate how advanced AI, particularly large language models like Gemini, can be effectively leveraged to enhance patient engagement, mitigate information overload, and foster more productive dialogues between individuals and healthcare systems. The project ultimately seeks to empower users with improved access to accurate and relevant health information, thereby elevating health literacy and outcomes through sophisticated conversational interfaces.

GitHub

2 stories
01

🚀 RAG-Anything: All-in-One RAG Framework

RAG-Anything is a comprehensive, all-in-one multimodal document processing RAG (Retrieval-Augmented Generation) system, designed to overcome the limitations of traditional text-centric RAGs. Built on LightRAG, it provides a unified framework for seamlessly processing and querying diverse content modalities including text, images, tables, equations, and multimedia, eliminating the need for multiple specialized tools. The system features an end-to-end multimodal pipeline, universal support for various document formats (PDFs, Office documents, images), and specialized analysis engines for visual, structured, and mathematical content. Its multi-stage architecture includes high-fidelity document parsing via MinerU, intelligent content understanding, a multimodal analysis engine, a knowledge graph index for cross-modal relationships, and modality-aware hybrid retrieval. RAG-Anything also supports advanced VLM-enhanced queries and direct content list insertion. This makes it an invaluable tool for academic research, technical documentation, financial reports, and enterprise knowledge management, delivering comprehensive insights from rich, mixed-content documents through a single interface.

02

HumanLayer

HumanLayer is a project designed to address the critical challenge of safely integrating high-stakes functions into AI agent workflows. While Large Language Models (LLMs) excel at interacting with the outside world via tools, giving them direct access to sensitive operations carries significant risks due to potential inaccuracies or hallucinations. HumanLayer provides a robust set of tools, notably the `@require_approval` decorator and `human_as_tool` functionality, to deterministically guarantee human oversight for these critical function calls. This ensures a human-in-the-loop mechanism, even if the LLM makes an error. The framework is built to empower next-generation autonomous agents, enabling them to operate in "outer loop" scenarios where they initiate interactions and manage complex workflows, while still obtaining necessary human input for sensitive tasks. By bridging the gap between highly capable LLMs and the need for reliability in high-impact applications, HumanLayer facilitates the development of more trustworthy and impactful AI solutions for complex business processes.

huggingface

6 stories
01

Tree Search for LLM Agent Reinforcement Learning

Recent advances in reinforcement learning (RL) have significantly enhanced the agentic capabilities of large language models (LLMs). In long-term and multi-turn agent tasks, existing approaches driven solely by outcome rewards often suffer from the problem of sparse supervision. To address the challenge, we propose Tree-based Group Relative Policy Optimization (Tree-GRPO), a grouped agent RL method based on tree search, where each tree node represents the complete agent interaction step. By sharing common prefixes, the tree search sampling increases the number of rollouts achievable within a fixed budget of tokens or tool calls. Moreover, we find that the tree-structured trajectory naturally allows the construction of step-wise process supervised signals even using only the outcome reward. Based on this, Tree-GRPO estimates the grouped relative advantages both on intra-tree and inter-tree levels. Through theoretical analysis, we demonstrate that the objective of intra-tree level group relative policy optimization is equivalent to that of step-level direct preference learning. Experiments across 11 datasets and 3 types of QA tasks demonstrate the superiority of the proposed tree-based RL over the chain-based RL method.

02

MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources

Large multimodal reasoning models have achieved rapid progress, but their advancement is constrained by two major limitations: the absence of open, large-scale, high-quality long chain-of-thought (CoT) data, and the instability of reinforcement learning (RL) algorithms in post-training. Group Relative Policy Optimization (GRPO), the standard framework for RL fine-tuning, is prone to gradient vanishing when reward variance is low, which weakens optimization signals and impairs convergence. This work makes three contributions: (1) We propose Variance-Aware Sampling (VAS), a data selection strategy guided by Variance Promotion Score (VPS) that combines outcome variance and trajectory diversity to promote reward variance and stabilize policy optimization. (2) We release large-scale, carefully curated resources containing ~1.6M long CoT cold-start data and ~15k RL QA pairs, designed to ensure quality, difficulty, and diversity, along with a fully reproducible end-to-end training codebase. (3) We open-source a family of multimodal reasoning models in multiple scales, establishing standardized baselines for the community. Experiments across mathematical reasoning benchmarks demonstrate the effectiveness of both the curated data and the proposed VAS. Comprehensive ablation studies and analyses provide further insight into the contributions of each component. In addition, we theoretically establish that reward variance lower-bounds the expected policy gradient magnitude, with VAS serving as a practical mechanism to realize this guarantee. Our code, data, and checkpoints are available at https://github.com/LengSicong/MMR1.

03

Seedream 4.0: Toward Next-generation Multimodal Image Generation

We introduce Seedream 4.0, an efficient and high-performance multimodal image generation system that unifies text-to-image (T2I) synthesis, image editing, and multi-image composition within a single framework. We develop a highly efficient diffusion transformer with a powerful VAE which also can reduce the number of image tokens considerably. This allows for efficient training of our model, and enables it to fast generate native high-resolution images (e.g., 1K-4K). Seedream 4.0 is pretrained on billions of text-image pairs spanning diverse taxonomies and knowledge-centric concepts. Comprehensive data collection across hundreds of vertical scenarios, coupled with optimized strategies, ensures stable and large-scale training, with strong generalization. By incorporating a carefully fine-tuned VLM model, we perform multi-modal post-training for training both T2I and image editing tasks jointly. For inference acceleration, we integrate adversarial distillation, distribution matching, and quantization, as well as speculative decoding. It achieves an inference time of up to 1.8 seconds for generating a 2K image (without a LLM/VLM as PE model). Comprehensive evaluations reveal that Seedream 4.0 can achieve state-of-the-art results on both T2I and multimodal image editing. In particular, it demonstrates exceptional multimodal capabilities in complex tasks, including precise image editing and in-context reasoning, and also allows for multi-image reference, and can generate multiple output images. This extends traditional T2I systems into an more interactive and multidimensional creative tool, pushing the boundary of generative AI for both creativity and professional applications. Seedream 4.0 is now accessible on https://www.volcengine.com/experience/ark?launch=seedream.

04

ScaleDiff: Scaling Difficult Problems for Advanced Mathematical Reasoning

Large Reasoning Models (LRMs) have shown impressive capabilities in complex problem-solving, often benefiting from training on difficult mathematical problems that stimulate intricate reasoning. Recent efforts have explored automated synthesis of mathematical problems by prompting proprietary models or large-scale open-source models from seed data or inherent mathematical concepts. However, scaling up these methods remains challenging due to their high computational/API cost, complexity of prompting, and limited difficulty level of the generated problems. To overcome these limitations, we propose ScaleDiff, a simple yet effective pipeline designed to scale the creation of difficult problems. We efficiently identify difficult problems from existing datasets with only a single forward pass using an adaptive thinking model, which can perceive problem difficulty and automatically switch between "Thinking" and "NoThinking" modes. We then train a specialized difficult problem generator (DiffGen-8B) on this filtered difficult data, which can produce new difficult problems in large scale, eliminating the need for complex, per-instance prompting and its associated high API costs. Fine-tuning Qwen2.5-Math-7B-Instruct on the ScaleDiff-Math dataset yields a substantial performance increase of 11.3% compared to the original dataset and achieves a 65.9% average accuracy on AIME'24, AIME'25, HMMT-Feb'25, BRUMO'25, and MATH500, outperforming recent strong LRMs like OpenThinker3. Notably, this performance is achieved using the cost-efficient Qwen3-8B model as a teacher, demonstrating that our pipeline can effectively transfer advanced reasoning capabilities without relying on larger, more expensive teacher models. Furthermore, we observe a clear scaling phenomenon in model performance on difficult benchmarks as the quantity of difficult problems increases. Code: https://github.com/QizhiPei/ScaleDiff.

05

SceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent

Indoor scene synthesis has become increasingly important with the rise of Embodied AI, which requires 3D environments that are not only visually realistic but also physically plausible and functionally diverse. While recent approaches have advanced visual fidelity, they often remain constrained to fixed scene categories, lack sufficient object-level detail and physical consistency, and struggle to align with complex user instructions. In this work, we present SceneWeaver, a reflective agentic framework that unifies diverse scene synthesis paradigms through tool-based iterative refinement. At its core, SceneWeaver employs a language model-based planner to select from a suite of extensible scene generation tools, ranging from data-driven generative models to visual- and LLM-based methods, guided by self-evaluation of physical plausibility, visual realism, and semantic alignment with user input. This closed-loop reason-act-reflect design enables the agent to identify semantic inconsistencies, invoke targeted tools, and update the environment over successive iterations. Extensive experiments on both common and open-vocabulary room types demonstrate that SceneWeaver not only outperforms prior methods on physical, visual, and semantic metrics, but also generalizes effectively to complex scenes with diverse instructions, marking a step toward general-purpose 3D environment generation. Project website: https://scene-weaver.github.io/.

06

Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving

End-to-End (E2E) solutions have emerged as a mainstream approach for autonomous driving systems, with Vision-Language-Action (VLA) models representing a new paradigm that leverages pre-trained multimodal knowledge from Vision-Language Models (VLMs) to interpret and interact with complex real-world environments. However, these methods remain constrained by the limitations of imitation learning, which struggles to inherently encode physical rules during training. Existing approaches often rely on complex rule-based post-refinement, employ reinforcement learning that remains largely limited to simulation, or utilize diffusion guidance that requires computationally expensive gradient calculations. To address these challenges, we introduce ReflectDrive, a novel learning-based framework that integrates a reflection mechanism for safe trajectory generation via discrete diffusion. We first discretize the two-dimensional driving space to construct an action codebook, enabling the use of pre-trained Diffusion Language Models for planning tasks through fine-tuning. Central to our approach is a safety-aware reflection mechanism that performs iterative self-correction without gradient computation. Our method begins with goal-conditioned trajectory generation to model multi-modal driving behaviors. Based on this, we apply local search methods to identify unsafe tokens and determine feasible solutions, which then serve as safe anchors for inpainting-based regeneration. Evaluated on the NAVSIM benchmark, ReflectDrive demonstrates significant advantages in safety-critical trajectory generation, offering a scalable and reliable solution for autonomous driving systems.