NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-26DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Eagle 3.1: Collaboration Between the EAGLE Team, vLLM Team, and TorchSpec Team

This technical blog post highlights Eagle 3.1, a major collaborative initiative involving the EAGLE team, the vLLM team, and the TorchSpec team. The integration aims to advance the efficiency of speculative decoding for large language models. Speculative decoding has emerged as a crucial optimization technique to accelerate model inference without compromising generation quality. Through this tripartite engineering effort, Eagle 3.1 introduces optimized system integrations that leverage vLLM's high-throughput serving engine and TorchSpec's specialized speculative execution frameworks. This synergy significantly reduces latency and enhances token-per-second generation speeds. By bridging the gap between algorithmic design and hardware-aware runtime optimization, the collaboration provides a highly performant, open-source solution for modern LLM deployment, paving the way for more responsive AI applications and scalable enterprise-level serving architectures.

02

A sleep-like consolidation mechanism for LLMs

This paper introduces an innovative, sleep-like memory consolidation mechanism designed specifically for Large Language Models (LLMs). Recognizing that neural networks often suffer from catastrophic forgetting and inefficient long-term knowledge retention, the researchers propose a novel training phase that mimics biological sleep. During this dedicated offline phase, the LLM processes and reorganizes recently acquired information, strengthening valuable associations while pruning redundant parameters and noisy connections. The study demonstrates that integrating this biological-inspired sleep cycle significantly improves the model's factual recall, reasoning stability, and overall sample efficiency. By transitioning from continuous active learning to a balanced active-sleep cycle, the LLM exhibits enhanced long-term memory consolidation similar to human cognitive processes, offering a promising pathway toward more robust, adaptable, and energy-efficient artificial intelligence architectures.

03

DeepSWE: A contamination-free benchmark for long-horizon coding agents

DeepSWE is introduced as a novel, contamination-free benchmark designed specifically for testing and evaluating long-horizon coding agents. Modern language models and software engineering agents often suffer from data contamination when evaluated on public code repositories that might have been included in their training sets. DeepSWE addresses this problem by providing a clean, isolated environment with complex, multi-step software engineering tasks that are completely free from training set leakage. This benchmark enables researchers to measure the genuine reasoning, planning, and execution capabilities of advanced AI systems over extended sequences of coding operations, establishing a more reliable standard for autonomous software development evaluation.

04

Outsourcing plus local AI will soon become more economical vs. frontier labs

This article explores a shifting paradigm in the artificial intelligence industry, suggesting that the combination of outsourcing paired with localized, open-source AI models will soon offer a more cost-effective alternative to relying on proprietary frontier labs. While proprietary giant models currently dominate high-end capabilities, the cost of API calls and data privacy concerns make them less viable for sustained business scaling. Conversely, deploying localized, smaller, fine-tuned AI architectures alongside Human-in-the-Loop outsourcing models creates a highly efficient workflow. This combined strategy allows enterprises to achieve comparable accuracy and task completion rates while significantly reducing overall operational expenses, safeguarding sensitive data, and lessening dependency on major foundation model providers.

05

Launch HN: Minicor (YC P26) – Windows desktop automations at scale

Faiz and Saheed have launched Minicor, a Y Combinator-backed (P26) platform designed to enable artificial intelligence companies to quickly build and deploy scalable desktop Robotic Process Automation solutions. Minicor targets legacy systems that lack API access, such as clinic-based Windows medical record systems. Traditionally, scaling desktop RPA has been hindered by complex scripting, dynamic UI changes, orchestration challenges including virtual machine queuing and parallelization, and difficult debugging environments with high failure rates. Minicor addresses these persistent automation barriers by providing robust orchestration, enhanced observability, and reliable execution framework for desktop environments. By lowering failure rates and streamlining VM orchestration, Minicor enables modern AI agents and integration platforms to interact seamlessly with local desktop applications, bringing scalability and reliability to legacy system automation.

06

Bay Area mom out thousands after scammers use AI to mimic daughter's voice

This report details a sophisticated financial scam targeting a Bay Area mother who lost thousands of dollars to fraudsters using artificial intelligence to mimic her daughter's voice. The scammers executed a highly convincing, fake kidnapping scheme by leveraging advanced voice cloning technology, which requires only a short audio sample to generate highly realistic, real-time vocal imitations. This incident highlights a growing and alarming trend of AI-enabled social engineering attacks, where malicious actors exploit Generative AI to bypass traditional security awareness and emotionally manipulate victims. Security experts warn that the rapid democratization of high-fidelity synthetic voice generation tools has lowered the barrier to entry for cybercriminals. The case underscores an urgent need for enhanced public awareness, the adoption of family safety protocols like verbal passphrases, and the development of robust detection mechanisms to identify synthetic audio in real-time communication channels.

Twitter

6 stories
01

ZoubinGhahrama1_AlphaProof Nexus

Google DeepMind has introduced AlphaProof Nexus, an innovative agentic framework designed to advance research-level mathematics through artificial intelligence. By leveraging state-of-the-art agentic capabilities, the system aims to solve complex mathematical problems that were previously beyond the reach of automated reasoning tools. This development marks a significant milestone in integrating formal logic with advanced machine learning techniques, enabling AI to contribute meaningfully to higher-order mathematical research.

02

gdb_GPT-5.5 Coding Review

Greg Brockman, a prominent figure in the artificial intelligence industry, recently shared his positive assessment of GPT-5.5, highlighting its exceptional performance specifically in the domain of computer programming. The tweet suggests that this iteration of the model demonstrates superior coding capabilities, potentially marking a significant advancement in automated software development and code generation tasks, setting a new performance benchmark for large language models.

03

Google_Gemini AI Update

Google has officially announced significant updates to its Gemini AI ecosystem, focusing on enhancing the model's performance, speed, and multimodal capabilities. The latest improvements aim to streamline user interaction by integrating more advanced reasoning and processing power across its diverse product suite, demonstrating superior document comprehension, real-time data analysis, and advanced code generation.

04

Kling_ai_House of David Integration

Jon Erwin, founder of Wonder Project and creator of the Amazon Prime series House of David, has publicly highlighted the pivotal role of Kling AI in the production of his show's first two seasons. The collaboration reportedly achieved several industry firsts, demonstrating how advanced generative video technology can be integrated into high-budget television workflows to enhance visual storytelling.

05

natolambert_Gemma MoE Call

In a recent social media post, AI researcher Nathan Lambert publicly called for the release of Google's 100B parameter Gemma 4 Mixture-of-Experts (MoE) model. Driven by the official launch of Gemini Flash 3.5, Lambert suggests that this high-capacity open model represents an essential development for the open-source community to access large-scale, open-weight architectures.

06

Hailuo_AI_Video Generation Release

Hailuo AI has announced a breakthrough capability in automated content creation, showcasing the platform's ability to generate an entire season of narrative video content in just sixty seconds. By leveraging advanced generative models, the system significantly compresses production timelines for complex visual stories, showcasing a major paradigm shift in generative video synthesis.

huggingface

6 stories
01

DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning

Reinforcement Learning has become a standard paradigm for aligning Large Language Models with human intent and task requirements. While Group Relative Policy Optimization offers an efficient, value-model-free alternative to Proximal Policy Optimization, adapting it to real-world multi-reward settings remains challenging. Standard scalarization practices suffer from significant drawbacks. To address these limitations, we propose Dynamic Variance-adaptive Advantage Optimization (DVAO), which dynamically adjusts combination weights based on the empirical reward variance of each objective within a rollout group, effectively up-weighting objectives with a stronger learning signal while suppressing noisy ones. Extensive experiments on mathematical reasoning and tool-use benchmarks using Qwen3 and Qwen2.5 models demonstrate that DVAO significantly outperforms baseline methods.

02

Macaron-A2UI: A Model for Generative UI in Personal Agents

As personal agents evolve to handle complex, user-centric tasks, static plain-text chat is rapidly becoming a bottleneck. Generative UI emerges as the necessary new interface layer, dynamically synthesizing the right controls, options, and state from the interaction context in real time. We present Macaron-A2UI, a model for Generative UI in personal agents. Our goal is to move beyond text-only interaction by enabling agents to generate natural language together with lightweight, executable UI actions for information collection, preference refinement, confirmation, and multi-goal organization. We train 30B, 235B and 754B models with parameter-efficient LoRA-based supervised fine-tuning followed by reward-driven reinforcement learning. The best Macaron-A2UI model reaches 75.6 overall on A2UI-Bench without explicit schema hints.

03

QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasks

Deep research agents extend the role of search engines from retrieving keyword-matched pages to synthesizing knowledge. We release QUEST, a family of open models (ranging from 2B to 35B) that serve as general-purpose deep research agents designed to handle a wide range of long-horizon search tasks, with strong capabilities in fact seeking, citation grounding, and report synthesis. To build QUEST, we propose an effective training recipe combining mid-training, supervised fine-tuning, and reinforcement learning based on a curated data synthesis pipeline with unified rubric trees. QUEST approaches or even surpasses frontier closed-source agents across eight deep research benchmarks.

04

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its extension to computer-use agents (CUAs) has been bottlenecked by the scarcity of scalable training data with deterministic rewards. We present CUA-Gym, a scalable pipeline that co-generates task instructions, environment states, and reward functions. Using this pipeline, we construct CUA-Gym, a dataset of 32,112 verified RLVR training tuples grounded in 110 environments. Trained with GSPO on CUA-Gym, our agents achieve 62.1% and 72.6% on OSWorld-Verified, outperforming prior open-source CUAs at comparable scales.

05

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive. To bridge this gap, we introduce ProAct, a proactive agent architecture that leverages idle-time compute to anticipate and fulfill likely upcoming user needs. By analyzing evolving dialogue history together with persistent memory, ProAct predicts upcoming needs and iteratively acquires information, allowing the agent to resolve knowledge gaps and prepare evidence before the user initiates a query. We also introduce ProActEval, a comprehensive benchmark comprising 200 scenarios across 40 domains, where ProAct accelerates task completion by reducing required turns by 14.8%.

06

On-Policy Adversarial Flow Distillation for Autoregressive Video Generation

Autoregressive video generators are attractive for streaming, long-horizon, and interactive applications, but distilling strong black-box teachers into causal students remains difficult. We propose Adversarial Flow Distillation (AFD), an on-policy framework for heterogeneous black-box video distillation. AFD queries the teacher and rolls out the current student on the same prompts, trains a prompt-paired Bradley-Terry discriminator to estimate clean-sample teacher-student discrepancy, and converts the resulting on-policy advantage into forward-process flow-matching updates on the student's own noised states. AFD provides dense velocity-field supervision while requiring no teacher scores or reverse-chain reinforcement learning.