NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-06DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Learning the Integral of a Diffusion Model

The research explores a novel approach to understanding and manipulating diffusion models by focusing on learning their integral. This method potentially offers a more direct and efficient way to model the generative process inherent in diffusion-based generative models, moving beyond iterative sampling. By integrating the diffusion process, researchers aim to develop 'flow maps' that can transform noise into data or vice-versa in a single, continuous step, rather than through a sequence of discrete steps. This could lead to significant advancements in computational efficiency for generating high-quality synthetic data, image synthesis, and other generative tasks. The underlying mathematical framework likely involves connections to continuous-time stochastic processes, optimal transport, and normalizing flows, providing a theoretical foundation for faster and more controlled generation.

02

Google Cloud fraud defense, the next evolution of reCAPTCHA

Google Cloud has announced the launch of Google Cloud Fraud Defense, marking the next significant evolution of its established reCAPTCHA technology. This new comprehensive service is engineered to deliver advanced, real-time fraud protection tailored for businesses operating within the Google Cloud ecosystem. Extending beyond reCAPTCHA's core function of discerning human users from automated bots, Google Cloud Fraud Defense incorporates sophisticated machine learning models and behavioral analytics. Its purpose is to proactively detect and mitigate a wider array of fraudulent activities, including account takeovers, payment fraud, synthetic account creation, and large-scale bot attacks. The solution is poised to harness Google's vast global threat intelligence and deep expertise in artificial intelligence, enabling it to continuously adapt to emerging fraud patterns. This strategic enhancement aims to fortify the security posture of enterprises, safeguard critical digital assets, and ensure trust in online interactions, reflecting Google Cloud's ongoing commitment to delivering intelligent and scalable cybersecurity defenses.

03

Agents can now create Cloudflare accounts, buy domains, and deploy

Cloudflare has announced a new capability enabling automated agents to interact directly with its platform, streamlining various development and deployment workflows. This advancement allows agents to programmatically create new Cloudflare accounts, acquire domain names, and deploy applications, potentially automating significant portions of web infrastructure management. The integration aims to enhance efficiency for developers and organizations leveraging AI agents or other automated systems to manage their online presence. This development signifies Cloudflare's push towards greater platform programmability and integration with emerging AI-driven tools, offering a more hands-off approach to managing web services from account creation to global deployment.

04

Vibe coding and agentic engineering are getting closer than I'd like

The article's concise title, "Vibe coding and agentic engineering are getting closer than I'd like," signals a growing apprehension regarding the evolving landscape of software development. "Vibe coding" generally describes an intuitive, less structured programming methodology, often characterized by rapid iteration and a focus on achieving a functional outcome without extensive upfront design or formal processes. This approach contrasts sharply with traditional, disciplined software engineering practices. Concurrently, "agentic engineering" represents the emerging trend of utilizing advanced AI agents to automate significant portions of the software development lifecycle, from ideation and code generation to testing and deployment. These AI agents, often powered by sophisticated Large Language Models, are designed to operate autonomously, making decisions and executing tasks with minimal human intervention. The author's concern arises from the potential convergence of these two paradigms. A fear exists that the inherent lack of rigorous structure in "vibe coding" could be amplified or exacerbated when integrated with autonomous AI agents. This convergence might lead to software systems that are less predictable, harder to debug, or more challenging to maintain due to a diminished human understanding of the underlying agent-driven processes. Such a scenario could compromise software quality, increase technical debt, and introduce unforeseen complexities, prompting a critical re-evaluation of current practices in AI-assisted software development.

05

Show HN: Tilde.run – Agent Sandbox with a Transactional, Versioned Filesystem

Tilde.run introduces an innovative agent sandbox platform featuring a transactional and versioned filesystem. This new offering aims to provide developers and researchers with a robust environment for deploying and testing software agents. The core distinction lies in its sophisticated filesystem, which ensures data integrity through transactional operations and enables comprehensive state management via versioning. This capability is crucial for debugging agent behavior, reproducing specific scenarios, and maintaining audit trails of agent interactions and modifications to their operational environment. By combining a secure sandbox with advanced data management, Tilde.run addresses common challenges in agent development, such as ensuring deterministic execution, managing complex dependencies, and rolling back to previous states. This platform is poised to facilitate the development of more reliable and accountable AI agents and automated systems across various applications, from research prototypes to production deployments. It provides a foundational infrastructure designed to enhance the development lifecycle of intelligent agents by offering control and transparency over their operational data.

06

Telus Uses AI to Alter Call-Agent Accents

Telus, a prominent telecommunications company, has reportedly implemented artificial intelligence technology to modify the accents of its call center agents in real-time. This innovative application of AI aims to enhance customer experience by potentially making communication clearer and more universally understandable, particularly for customers engaging with agents from diverse linguistic backgrounds. The technology likely leverages advanced speech synthesis and voice transformation algorithms, enabling a dynamic alteration of an agent's vocal characteristics to a perceived standard or a more neutral accent. While proponents suggest this could lead to improved service efficiency and customer satisfaction by reducing communication barriers, the deployment of such AI raises significant ethical and cultural considerations. Concerns may include the impact on agents' identity, the potential for accent discrimination, and broader questions about authenticity in human-computer interaction. The initiative underscores a growing trend in businesses adopting sophisticated AI tools to optimize customer service operations, pushing the boundaries of real-time audio manipulation and highlighting the complex interplay between technology, human interaction, and cultural perception.

huggingface

6 stories
01

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning

Reinforcement learning (RL) has become a central post-training tool for improving the reasoning abilities of large language models (LLMs). In these systems, the rollout, the trajectory sampled from a prompt to termination, including intermediate reasoning steps and optional tool or environment interactions, determines the data the optimizer learns from, yet rollout design is often underreported. This survey provides an optimizer-agnostic view of rollout strategies for RL-based post-training of reasoning LLMs. We formalize rollout pipelines with unified notation and introduce Generate-Filter-Control-Replay (GFCR), a lifecycle taxonomy that decomposes rollout pipelines into four modular stages: Generate proposes candidate trajectories and topologies; Filter constructs intermediate signals via verifiers, judges, critics; Control allocates compute and makes continuation/branching/stopping decisions under budgets; and Replay retains and reuses artifacts across rollouts without weight updates, including self-evolving curricula that autonomously generate new training tasks. We complement GFCR with a criterion taxonomy of reliability, coverage, and cost sensitivity that characterizes rollout trade-offs. Using this framework, we synthesize methods spanning RL with verifiable rewards, process supervision, judge-based gating, guided and tree/segment rollouts, adaptive compute allocation, early-exit and partial rollouts, throughput optimization, and replay/recomposition for self-improvement. We ground the framework with case studies in math, code/SQL, multimodal reasoning, tool-using agents, and agentic skill benchmarks that evaluate skill induction, reuse, and cross-task transfer. Finally, we provide a diagnostic index that maps common rollout pathologies to GFCR modules and mitigation levers, alongside open challenges for building reproducible, compute-efficient, and trustworthy rollout pipelines.

02

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration

This report describes ARIS (Auto-Research-in-sleep), an open-source research harness for autonomous research, including its architecture, assurance mechanisms, and early deployment experience. The performance of agent systems built on LLMs depends on both the model weights and the harness around them, which governs what information to store, retrieve, and present to the model. For long-horizon research workflows, the central failure mode is not a visible breakdown but a plausible unsupported success: a long-running agent can produce claims whose evidential support is incomplete, misreported, or silently inherited from the executor's framing. Therefore, we present ARIS as a research harness that coordinates machine-learning research workflows through cross-model adversarial collaboration as a default configuration: an executor model drives forward progress while a reviewer from a different model family is recommended to critique intermediate artifacts and request revisions. ARIS has three architectural layers. The execution layer provides more than 65 reusable Markdown-defined skills, model integrations via MCP, a persistent research wiki for iterative reuse of prior findings, and deterministic figure generation. The orchestration layer coordinates five end-to-end workflows with adjustable effort settings and configurable routing to reviewer models. The assurance layer includes a three-stage process for checking whether experimental claims are supported by evidence: integrity verification, result-to-claim mapping, and claim auditing that cross-checks manuscript statements against the claim ledger and raw evidence, as well as a five-pass scientific-editing pipeline, mathematical-proof checks, and visual inspection of the rendered PDF. A prototype self-improvement loop records research traces and proposes harness improvements that are adopted only after reviewer approval.

03

OpenSeeker-v2: Pushing the Limits of Search Agents with Informative and High-Difficulty Trajectories

Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet their development remains dominated by industrial giants. The typical industry recipe involves a highly resource-intensive pipeline spanning pre-training, continual pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). In this report, we show that when fueled with informative and high-difficulty trajectories, a simple SFT approach could be surprisingly powerful for training frontier search agents. By introducing three simple data synthesis modifications: scaling knowledge graph size for richer exploration, expanding the tool set size for broader functionality, and strict low-step filtering, we establish a stronger baseline. Trained on merely 10.6k data points, our OpenSeeker-v2 achieves state-of-the-art performance across 4 benchmarks (30B-sized agents with ReAct paradigm): 46.0% on BrowseComp, 58.1% on BrowseComp-ZH, 34.6% on Humanity's Last Exam, and 78.0% on xbench, surpassing even Tongyi DeepResearch trained with heavy CPT+SFT+RL pipeline, which achieves 43.4%, 46.7%, 32.9%, and 75.0%, respectively. Notably, OpenSeeker-v2 represents the first state-of-the-art search agent within its model scale and paradigm to be developed by a purely academic team using only SFT. We are excited to open-source the OpenSeeker-v2 model weights and share our simple yet effective findings to make frontier search agent research more accessible to the community.

04

Video Generation with Predictive Latents

Video Variational Autoencoder (VAE) enables latent video generative modeling by mapping the visual world into compact spatiotemporal latent spaces, improving training efficiency and stability. While existing video VAEs achieve commendable reconstruction quality, continued optimization of reconstruction does not necessarily translate into improved generative performance. How to enhance the diffusability of video latents remains a critical and unresolved challenge. In this work, inspired by principles of predictive world modeling, we investigate the potential of predictive learning to improve the video generative modeling. To this end, we introduce a simple and effective predictive reconstruction objective that unifies predictive learning with video reconstruction. Specifically, we randomly discard future frames and encode only partial past observations, while training the decoder to reconstruct the observed frames and predict future ones simultaneously. This design encourages the latent space to encode temporally predictive structures and build a more coherent understanding of video dynamics, thereby improving generation quality. Our model, termed Predictive Video VAE (PV-VAE), achieves superior performance on video generation, with 52% faster convergence and a 34.42 FVD improvement over the Wan2.2 VAE on UCF101. Furthermore, comprehensive analyses demonstrate that PV-VAE not only exhibits favorable scalability, with generative performance improving alongside VAE training, but also yields consistent gains in downstream video understanding, underscoring a latent space that effectively captures temporal coherence and motion priors.

05

Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL

The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce PRISM, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini 3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT-to-RLVR baseline on 4B and 8B, respectively. Our code, data, and model checkpoints are publicly available at https://github.com/XIAO4579/PRISM.

06

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

Iterative Retrieval-Augmented Generation (iRAG) has emerged as a powerful paradigm for answering complex multi-hop questions by progressively retrieving and reasoning over external documents. However, current systems predominantly operate on parsed text, which creates two critical bottlenecks: (1) Coarse-grained attribution, where users are burdened with manually locating evidence within lengthy documents based on vague text-level citations; and (2) Visual semantic loss, where the conversion of visually rich documents (e.g., slides, PDFs with charts) into text discards spatial logic and layout cues essential for reasoning. To bridge this gap, we present Chain of Evidence (CoE), a retriever-agnostic visual attribution framework that leverages Vision-Language Models to reason directly over screenshots of retrieved document candidates. CoE eliminates format-specific parsing and outputs precise bounding boxes, visualizing the complete reasoning chain within the retrieved candidate set. We evaluate CoE on two distinct benchmarks: Wiki-CoE, a large-scale dataset of structured web pages derived from 2WikiMultiHopQA, and SlideVQA, a challenging dataset of presentation slides featuring complex diagrams and free-form layouts. Experiments demonstrate that fine-tuned Qwen3-VL-8B-Instruct achieves robust performance, significantly outperforming text-based baselines in scenarios requiring visual layout understanding, while establishing a retriever-agnostic solution for pixel-level interpretable iRAG. Our code is available at https://github.com/PeiYangLiu/CoE.git.