NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-29DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Notes from the Mistral AI Now Summit in Paris

This document outlines key insights and announcements from the Mistral AI Now Summit held in Paris. The summit highlighted Mistral AI's latest advancements in open-source and commercial language models, emphasizing their focus on high-performance, cost-effective AI solutions for developers and enterprises. Key discussion points included the deployment of specialized models, optimization techniques like quantization, and the growing ecosystem supporting Mistral's architecture. The event underscored the company's commitment to fostering an open ecosystem that challenges proprietary alternatives, while showcasing real-world use cases, API enhancements, and strategic partnerships aimed at accelerating the adoption of generative AI technologies globally.

02

Liquid AI reveals 8B-A1B MoE trained on 38T

Liquid AI has officially announced the launch of its latest model, an 8B-A1B Mixture of Experts architecture trained on a massive dataset of 38 trillion tokens. This release highlights the company's commitment to developing highly efficient Liquid Foundation Models that depart from traditional Transformer architectures. By utilizing advanced dynamic systems and state-space formulations, the new model achieves state-of-the-art performance benchmarks while operating with a significantly smaller footprint during inference. The training phase represents a major milestone, proving that non-Transformer designs can scale effectively when exposed to ultra-large-scale datasets. This release is expected to provide developers with a highly responsive, cost-effective alternative for deploying complex AI workflows, setting a new standard for computational efficiency in the generative AI space.

03

CAPTCHAs can still detect AI agents

This research analyzes the continued efficacy of Completely Automated Public Turing tests to tell Computers and Humans Apart (CAPTCHAs) in the era of advanced artificial intelligence. Despite rapid advancements in Large Language Models and computer vision algorithms that threaten traditional security gates, modern CAPTCHA implementations remain surprisingly effective at identifying and blocking automated AI agents. By utilizing sophisticated behavioral tracking, dynamic cognitive challenges, and anomaly detection, modern security systems can differentiate between human interactions and simulated agent workflows. The study highlights the evolution of CAPTCHA design from static image recognition to complex, interactive behavioral analysis, offering a critical defense layer against unauthorized web scraping, automated spamming, and bot-driven exploitation. Ultimately, the findings suggest that while basic visual puzzles are increasingly vulnerable to machine learning attacks, adaptive challenge-response mechanisms can still robustly detect and mitigate the actions of sophisticated AI agents across the web.

04

CVE-Bench: testing LLM agents on real-world vulnerability patches

CVE-Bench introduces a novel, rigorous benchmark dataset designed to evaluate the capabilities of Large Language Model agents in generating software vulnerability patches for real-world Common Vulnerabilities and Exposures. Unlike synthetic benchmarks, CVE-Bench leverages historical, actual security vulnerabilities across various programming languages to test how effectively AI agents can reason about codebases, identify security flaws, and produce functional, secure code modifications. The benchmark aims to bridge the gap between theoretical software engineering capabilities and practical cybersecurity deployment, providing a standardized environment to track progress in automated security patching. The results emphasize the current limitations of state-of-the-art LLMs in handling complex, multi-file code dependencies and highlight the necessity of advanced reasoning architectures for successful autonomous cyber defense operations.

05

Robinhood now lets your AI agents trade stocks

Robinhood has launched a new capability allowing user-configured AI agents to execute stock trades directly on its platform. This feature marks a significant shift toward autonomous financial transactions powered by AI. By providing secure API integrations and programmatic access for external AI systems, Robinhood enables developers and users to deploy autonomous software capable of analyzing market conditions and automatically executing buying or selling decisions. This integration bridges the gap between decentralized software agents and traditional financial markets, offering new infrastructure for automated algorithmic trading while presenting unique regulatory and safety considerations for real-time market operations.

06

Real-time LLM Inference on Standard GPUs: 3k tokens/s per request

This technical article outlines a breakthrough optimization technique for running Large Language Model (LLM) inference on standard GPU hardware, achieving an impressive processing speed of 3,000 tokens per second per individual request. By addressing key computational bottlenecks in standard memory bandwidth and execution pipelines, the development team has managed to scale up throughput significantly. This performance leap is realized without relying on specialized ultra-high-end hardware configurations, making real-time, high-speed LLM applications highly accessible for mainstream enterprise deployments. The architecture shifts the paradigm of real-time text generation by optimizing tensor parallelism, memory management, and attention mechanisms. Ultimately, this implementation proves that intelligent structural optimizations can dramatically democratize high-throughput artificial intelligence serving, opening new avenues for interactive agent behaviors and instant conversational responses.

Twitter

6 stories
01

New ChatGPT Model 5.5

OpenAI's leadership has announced the official launch of the new 5.5 instant model integrated into the ChatGPT platform. This release represents a significant advancement in model efficiency and real-time processing capabilities, aiming to provide users with faster and more responsive interactions. By introducing the 5.5 version, the organization continues to iterate on its large language model architecture to optimize performance metrics and latency. This update is designed to support more complex reasoning tasks while maintaining the conversational fluency expected from the ChatGPT ecosystem.

02

Google_Gemini AI Agent Launch

At the Google I/O developer conference, Google officially introduced Gemini Spark, a sophisticated personal AI agent designed to function as a 24/7 digital assistant. This new tool empowers users to navigate their digital lives more effectively by performing tasks and taking actions on their behalf under direct user supervision. A key technical highlight is its seamless cross-platform integration, allowing the agent to operate efficiently across mobile devices and desktop computers, even maintaining functionality while a laptop is closed. This product launch represents a significant milestone in Google's effort to integrate advanced AI capabilities into everyday productivity workflows.

03

runwayml_The Rogue Project

Runway has unveiled a behind-the-scenes look at the production of 'The Rogue,' a short film created entirely by a single developer in under one month using their AI video platform. This project serves as a cornerstone for the new initiative, Project Luxo, which aims to investigate and demonstrate how modern generative video technology has effectively bridged the uncanny valley. By showcasing high-quality, AI-synthesized imagery, Runway highlights the significant advancements in realism and production efficiency within the creative industry. The initiative underscores the company's commitment to pushing the boundaries of automated filmmaking.

04

Kling AI Showcase RAPHAEL

Kling AI has officially presented a behind-the-scenes look at the creation of RAPHAEL, an innovative AI-powered feature film. This showcase provides a comprehensive deep dive into the practical application of Kling AI tools throughout the entire filmmaking pipeline. By integrating generative video technology from initial creative ideation and storyboarding through to the production of high-fidelity cinematic frames, the project demonstrates significant improvements in production efficiency and creative output.

05

johnschulman2_LLM Renderers

John Schulman highlights the critical role of renderers within the Large Language Model software stack. Renderers act as a foundational layer by mapping tokens to messages, ensuring consistency across varying tokenizer implementations and formatting conventions. The author emphasizes that accurate message handling is essential for maintaining alignment between APIs, datasets, and Reinforcement Learning (RL) environments. Failure to manage these details effectively can introduce significant technical risks.

06

llama.cpp Official Launch

Georgi Gerganov, the creator of the llama.cpp project, has officially launched a dedicated website for the popular local artificial intelligence inference engine. This milestone represents a significant step forward in the project's mission to make high-performance local AI accessible to a wider audience. By providing a centralized hub, the project aims to streamline documentation, resources, and community engagement for developers and researchers working with Large Language Models on consumer-grade hardware.

huggingface

6 stories
01

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In this work, we study whether heterogeneous embodied decision-making problems can be unified within a single vision-language-action model. We present Qwen-VLA, a unified embodied foundation model that extends Qwen's vision-language modeling stack from perception, understanding, and reasoning to continuous action and trajectory generation through a DiT-based action decoder.

02

minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models

Recent video diffusion foundation models have achieved remarkable progress in high-quality video generation, yet turning them into real-time interactive video world models remains challenging. Interactive world models require controllable, causal, and low-latency rollout, which in practice demands a full pipeline spanning data construction, controllable fine-tuning, autoregressive training, few-step distillation, and streaming inference. In this work, we present minWM, a full-stack open-source framework for building real-time interactive video world models.

03

AdaState: Self-Evolving Anchors for Streaming Video Generation

Autoregressive video diffusion models generate streaming video by producing frames sequentially, conditioning each chunk on previously generated content. These models are structurally anchored to the first frame, which draws disproportionate attention and suppresses video dynamics. To address this, we replace the static anchor with an adaptive state, a hidden latent that the model denoises alongside content at every chunk but never renders, improving video dynamics and motion.

04

Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning

Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks. We propose Skill0.5, a novel agentic RL framework that explicitly differentiates skill treatments by combining general skill internalization with task-specific skill utilization, outperforming both memory-based and skill-based RL baselines.

05

AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenarios

Large language model (LLM)-based agents have shown strong capabilities in using external tools to solve complex tasks. However, existing evaluations often overlook the temporal dimension of tool use, especially tool response latency. We propose AsyncTool, a benchmark for assessing LLM-based agents in interactive multi-task tool-use environments with delayed tool feedback.

06

Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation

We propose Ptah, a multi-agent harness for interleaved report generation that orchestrates the lifecycle from user query to rendered web report. It utilizes specialized agents to construct visual-aware plans, collect claim-grounded evidence, maintain source-aligned images in a Visual Working Memory, and verify factual grounding and cross-modal consistency.