NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-08-15ENGLISH EDITION
This issue
—
All time
—

Twitter

6 stories
01

latentspacepod_Greg Brockman on OpenAI's Future, GPT-5, and AGI

In a Latent.Space podcast, OpenAI President Greg Brockman discusses the GPT-5 era, evaluating model intelligence, scaling compute, and the path to AGI. He emphasizes "energy turns into compute, turns into intelligence," sharing OpenAI's progress in reasoning evolution, online/offline learning, model routing, pricing optimization, and self-improving agents. The conversation reveals OpenAI's vision for future AI research and engineering, alongside reflections on the value of engineers in the age of AGI.

02

OpenAI_ChatGPT Weekly Updates: GPT-4o and GPT-5 Features & Availability Expansion

OpenAI announced its latest weekly ChatGPT updates, including GPT-4o being default for paid users, who can also enable more legacy models and GPT-5 Thinking mini. GPT-5 introduces Auto, Fast, and Thinking modes for varied response speed and depth. Plus and Team users now receive up to 3,000 GPT-5 Thinking messages per week. Furthermore, GPT-5 is now available for Enterprise and Education users, with a warmer personality coming soon.

03

Alibaba_Qwen_Qwen Chat Vision Understanding Major Upgrade

Qwen Chat's vision understanding capabilities have received a significant update, now featuring native 128K context and stronger performance across vision, video, and 3D tasks. Key upgrades include a substantial boost in math and reasoning, more accurate object recognition, OCR support for over 30 languages, enhanced 2D and 3D grounding, and major improvements in video understanding.

04

xywang626_OpenCUA: First Open-Source Computer-Use Agent Foundation Model Framework Released

Xinyuan Wang's team has released OpenCUA, the first from-scratch open-source foundation model framework for computer-use agents, along with the SOTA model OpenCUA-32B. This framework matches top proprietary models on the OSWorld-Verified benchmark and provides full open-source infrastructure and data. OpenCUA aims to address the lack of large-scale open desktop agent datasets and transparent pipelines. It offers a comprehensive stack for scalable data collection, effective data formulation, model training strategies, and reproducible evaluation, powering top open-source models like OpenCUA-7B and OpenCUA-32B.

05

_philschmid_Imagen 4 Officially Launched with Performance Boost and Pricing Details

Philipp Schmid announced that Imagen 4 is now generally available in AI Studio and the Gemini API. The service offers three models: Ultra, Standard, and Fast, with pricing starting at $0.02 per image. Imagen 4 boasts up to 10x faster generation compared to its predecessor, supports image output up to 2K resolution, and significantly improves spelling and typography for longer text strings. Users can generate between 1 and 4 images per prompt, with specific pricing at $0.06 for Ultra, $0.04 for Standard, and $0.02 for Fast.

06

omarsar0_OdysseyBench: Multi-Day Office Task AI Agent Benchmark

The omarsar0 team has introduced OdysseyBench, a novel benchmark and data-generation pipeline designed to evaluate AI agents' performance in realistic, multi-day office tasks. This benchmark spans applications like Word, Excel, PDF, Email, and Calendar, specifically targeting long-horizon, context-dependent workflows rather than atomic tasks. It provides a new paradigm for assessing agent capabilities in complex office environments.

GitHub

4 stories
01

FastAPI-MCP

FastAPI-MCP is an innovative Python library designed to seamlessly expose FastAPI API endpoints as Model Context Protocol (MCP) tools, complete with built-in authentication. This project adopts a FastAPI-native approach, distinguishing itself from mere OpenAPI converters, and supports zero or minimal configuration. It automatically preserves the schemas of request and response models, as well as Swagger documentation. Key advantages include efficient communication via FastAPI's ASGI interface and flexible deployment options, allowing it to function either as an extension to an existing FastAPI application or as a standalone service. By offering native dependency management and a unified infrastructure, FastAPI-MCP significantly streamlines the integration of existing FastAPI services into the MCP ecosystem, making it particularly suitable for developing and managing AI-powered tools.

02

SpatialLM

SpatialLM is an innovative 3D large language model specifically engineered to interpret and process complex 3D point cloud data. Its primary function is to generate highly structured 3D scene understanding outputs, which encompass detailed architectural elements like walls, doors, and windows, as well as precisely oriented object bounding boxes complete with their semantic categories. A key advantage of SpatialLM is its versatility in handling point clouds derived from a wide array of sources, including monocular video sequences, RGBD images, and LiDAR sensors, thereby overcoming limitations of previous methods requiring specialized equipment. This multimodal architecture effectively bridges the critical gap between raw, unstructured 3D geometric data and refined, structured 3D representations, offering a high-level semantic understanding of environments. Furthermore, SpatialLM 1.1 introduces advanced capabilities such as doubled point cloud resolution, a more powerful point cloud encoder (Sonata), and the ability to perform detection based on user-specified categories, leveraging the flexibility of LLMs. These features collectively enhance spatial reasoning, making SpatialLM highly valuable for cutting-edge applications in embodied robotics, autonomous navigation, and sophisticated 3D scene analysis tasks.

03

Magentic-UI

Magentic-UI is a cutting-edge research prototype of a human-centered interface, leveraging a sophisticated multi-agent system to automate complex web tasks while ensuring users retain full control. It excels at browsing and performing actions on the web, generating and executing code, and analyzing various file types, addressing scenarios from form filling to deep website navigation and data-driven code execution. Built on the AutoGen framework, Magentic-UI provides a transparent and highly controllable interaction paradigm, fostering efficient human-in-the-loop involvement. Its core functionalities include collaborative planning (Co-Planning), guided task execution (Co-Tasking), robust security measures (Action Guards), intelligent plan learning and retrieval from past runs, and efficient parallel task execution. This innovative approach significantly boosts human-agent interaction efficiency, making it an ideal solution for intricate web navigation, data extraction, and automated data processing challenges, as demonstrated by its performance on benchmarks like GAIA and AssistantBench.

04

Marker

Marker is an efficient and accurate document conversion tool, supporting various file formats such as PDF, images, Office documents, HTML, and EPUB, converting them into Markdown, JSON, chunks, or HTML. Its core features include intelligent recognition and formatting of complex elements like tables, equations, and code blocks, with support for all languages. Marker outperforms existing cloud services and open-source solutions in performance. Furthermore, it can significantly enhance conversion accuracy and structured data extraction capabilities by integrating Large Language Models (LLMs), particularly excelling in table recognition. The tool supports GPU/CPU/MPS operation, offering flexible API and CLI interfaces, making it widely applicable in document digitization, content management, and RAG data preparation.

wechat

6 stories
01

DINOv3 Arrives: 7 Billion Parameters, 1.7 Billion Images, Self-Supervised Learning Reaches New Milestone

Meta has launched DINOv3, a cutting-edge self-supervised learning (SSL) computer vision model. This new iteration scales unsupervised training to an unprecedented 7 billion parameters and leverages a massive 1.7 billion image dataset. DINOv3 marks a significant breakthrough by being the first single frozen visual backbone to outperform specialized solutions across multiple long-standing dense prediction tasks. Inspired by the success of large language models, DINOv3 emphasizes expanding model capacity and data scale. It introduces innovative techniques like Gram Anchoring to address performance degradation in dense tasks during extended training and incorporates high-resolution adaptation for enhanced feature utilization. The model demonstrates for the first time that self-supervised learning can broadly surpass weakly supervised models, achieving superior performance in both image classification and dense prediction benchmarks. Its real-world impact is already evident, for instance, in forest monitoring applications, signifying a major milestone for self-supervised learning in computer vision.

02

Google Open-Sources Gemma 3 270M, Outperforming Qwen 2.5 Equivalent Models

Google has officially released Gemma 3 270M, a compact language model with only 270 million parameters, specifically designed for fine-tuning on specialized tasks. Inheriting the advanced architecture of the Gemma 3 series, this model demonstrates robust instruction-following and text structuring capabilities, setting new performance benchmarks against equivalent models in IFEval benchmarks. Its key advantages include extreme energy efficiency (consuming only 0.75% battery for 25 conversations on a Pixel 9 Pro), support for production-ready INT4 quantization, and excellent instruction adherence. Gemma 3 270M aims to empower developers to build highly efficient, cost-effective, and offline-capable specialized AI systems. It is particularly suited for high-volume, low-latency, privacy-sensitive, and rapid iteration scenarios, fostering the widespread adoption of small, expert models.

03

Video-BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation

The Video-BLADE framework, jointly open-sourced by Zhejiang University and Huawei, significantly enhances the inference efficiency of video diffusion models by synergistically integrating Adaptive Block-Sparse Attention (ASA) with a data-free Trajectory Distribution Matching (TDM) distillation process. This innovative framework enables DiT models to achieve up to a 14x speedup while demonstrating superior or maintained generation quality, even surpassing traditional dense baselines on the VBench-2.0 benchmark. At its core, Video-BLADE dynamically prunes attention matrices to focus on critical spatiotemporal interactions and employs distribution-level alignment to ensure the student model efficiently learns the teacher model's generation trajectory. This approach substantially reduces computational costs while preserving high perceptual quality, establishing a new paradigm for efficient video generation.

04

StableAvatar: New Lip-Sync Technology for Infinite-Length, Zero-Frame-Drop Video Generation

Fudan University introduces StableAvatar, the first end-to-end video diffusion Transformer, specifically designed to overcome the limitations of current audio-driven avatar video generation models, which struggle to maintain natural audio synchronization and consistent identity over extended video durations. StableAvatar achieves this breakthrough through several key innovations: a timestep-aware audio adapter that suppresses error accumulation, an audio-native guidance mechanism that significantly improves audio-visual synchronization, and a dynamic weighted sliding window strategy to enhance overall video smoothness. These combined features enable the seamless generation of high-quality, infinite-length videos without the need for post-processing. Furthermore, StableAvatar demonstrates remarkable efficiency, reducing GPU memory consumption by approximately 50% and achieving a 10x inference speedup compared to leading competitors. Its superior performance in facial quality and precise lip-sync accuracy positions StableAvatar as a significant advancement in the field of long-duration avatar video generation.

05

WebWatcher: The First Open-Source Multimodal Deep Research Agent Surpassing Closed-Source Models

WebWatcher, the first open-source multimodal deep research agent, integrates various tools such as web browsing, image search, and code interpreters. It generates high-quality reasoning trajectories through a fully automated process, optimized by supervised fine-tuning and reinforcement learning. This agent is designed to tackle complex, cross-modal, multi-tool, and multi-step tasks. Its technical approach encompasses building high-difficulty multimodal data, optimizing reasoning trajectories, and applying reinforcement learning. Evaluated against challenging benchmarks like BrowseComp-VL, WebWatcher comprehensively outperforms leading closed-source models such as GPT-4o and Gemini in complex reasoning, information retrieval, knowledge integration, and information aggregation. This performance establishes WebWatcher's leading position as a new generation of open-source multimodal AI agents.

06

GPT-5 Surpasses Human Doctors: 24% Higher Reasoning and 29% Stronger Comprehension Than Experts

Recent research indicates that GPT-5 demonstrates exceptional performance in medical image reasoning and comprehension, with accuracy rates 24.23% and 29.40% higher than human experts, respectively. The model comprehensively outperforms GPT-4o and other variants across multiple standardized medical tests, including USMLE, MedXpertQA, and VQA-RAD. Its core capability enhancement stems from an architectural shift from text-dominant to native multimodal deep fusion, utilizing shared tokenization techniques and cross-modal attention mechanisms for seamless information processing. Although GPT-5 excels in ideal testing environments, researchers emphasize that its application in real-world clinical scenarios still requires further practical validation. Currently, in handling complex, unseen real-world cases, AI models, including GPT-5, still lag behind experienced radiologists.

huggingface

6 stories
01

We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities across various tasks, but still struggle with complex mathematical reasoning. Existing research primarily focuses on dataset construction and method optimization, often overlooking two critical aspects: comprehensive knowledge-driven design and model-centric data space modeling. In this paper, we introduce We-Math 2.0, a unified system that integrates a structured mathematical knowledge system, model-centric data space modeling, and a reinforcement learning (RL)-based training paradigm to comprehensively enhance the mathematical reasoning abilities of MLLMs. The key contributions of We-Math 2.0 are fourfold: (1) MathBook Knowledge System: We construct a five-level hierarchical system encompassing 491 knowledge points and 1,819 fundamental principles. (2) MathBook-Standard & Pro: We develop MathBook-Standard, a dataset that ensures broad conceptual coverage and flexibility through dual expansion. Additionally, we define a three-dimensional difficulty space and generate 7 progressive variants per problem to build MathBook-Pro, a challenging dataset for robust training. (3) MathBook-RL: We propose a two-stage RL framework comprising: (i) Cold-Start Fine-tuning, which aligns the model with knowledge-oriented chain-of-thought reasoning; and (ii) Progressive Alignment RL, leveraging average-reward learning and dynamic data scheduling to achieve progressive alignment across difficulty levels. (4) MathBookEval: We introduce a comprehensive benchmark covering all 491 knowledge points with diverse reasoning step distributions. Experimental results show that MathBook-RL performs competitively with existing baselines on four widely-used benchmarks and achieves strong results on MathBookEval, suggesting promising generalization in mathematical reasoning.

02

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

Prevailing autoregressive (AR) models for text-to-image generation either rely on heavy, computationally-intensive diffusion models to process continuous image tokens, or employ vector quantization (VQ) to obtain discrete tokens with quantization loss. In this paper, we push the autoregressive paradigm forward with NextStep-1, a 14B autoregressive model paired with a 157M flow matching head, training on discrete text tokens and continuous image tokens with next-token prediction objectives. NextStep-1 achieves state-of-the-art performance for autoregressive models in text-to-image generation tasks, exhibiting strong capabilities in high-fidelity image synthesis. Furthermore, our method shows strong performance in image editing, highlighting the power and versatility of our unified approach. To facilitate open research, we will release our code and models to the community.

03

ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing

Traditional cartoon and anime production involves keyframing, inbetweening, and colorization stages, which require intensive manual effort. Despite recent advances in AI, existing methods often handle these stages separately, leading to error accumulation and artifacts. For instance, inbetweening approaches struggle with large motions, while colorization methods require dense per-frame sketches. To address this, we introduce ToonComposer, a generative model that unifies inbetweening and colorization into a single post-keyframing stage. ToonComposer employs a sparse sketch injection mechanism to provide precise control using keyframe sketches. Additionally, it uses a cartoon adaptation method with the spatial low-rank adapter to tailor a modern video foundation model to the cartoon domain while keeping its temporal prior intact. Requiring as few as a single sketch and a colored reference frame, ToonComposer excels with sparse inputs, while also supporting multiple sketches at any temporal location for more precise motion control. This dual capability reduces manual workload and improves flexibility, empowering artists in real-world scenarios. To evaluate our model, we further created PKBench, a benchmark featuring human-drawn sketches that simulate real-world use cases. Our evaluation demonstrates that ToonComposer outperforms existing methods in visual quality, motion consistency, and production efficiency, offering a superior and more flexible solution for AI-assisted cartoon production.

04

PRELUDE: A Benchmark Designed to Require Global Comprehension and Reasoning over Long Contexts

We introduce PRELUDE, a benchmark for evaluating long-context understanding through the task of determining whether a character's prequel story is consistent with the canonical narrative of the original book. Our task poses a stronger demand for global comprehension and deep reasoning than existing benchmarks -- as the prequels are not part of the original story, assessing their plausibility typically requires searching and integrating information that is only indirectly related. Empirically, 88% of instances require evidence from multiple parts of the narrative. Experimental results highlight the challenge of our task: in-context learning, RAG and in-domain training with state-of-the-art LLMs, and commercial DeepResearch services, lag behind humans by >15%. A further human study reveals that models often produce correct answers with flawed reasoning, leading to an over 30% gap in reasoning accuracy compared to humans. These findings underscore the substantial room for improvement in long-context understanding and reasoning.

05

UI-Venus Technical Report: Building High-performance UI Agents with RFT

We present UI-Venus, a native UI agent that takes only screenshots as input based on a multimodal large language model. UI-Venus achieves SOTA performance on both UI grounding and navigation tasks using only several hundred thousand high-quality training samples through reinforcement finetune (RFT) based on Qwen2.5-VL. Specifically, the 7B and 72B variants of UI-Venus obtain 94.1% / 50.8% and 95.3% / 61.9% on the standard grounding benchmarks, i.e., Screenspot-V2 / Pro, surpassing the previous SOTA baselines including open-source GTA1 and closed-source UI-TARS-1.5.To show UI-Venus's summary and planing ability, we also evaluate it on the AndroidWorld, an online UI navigation arena, on which our 7B and 72B variants achieve 49.1% and 65.9% success rate, also beating existing models.To achieve this, we introduce carefully designed reward functions for both UI grounding and navigation tasks and corresponding efficient data cleaning strategies.To further boost navigation performance, we propose Self-Evolving Trajectory History Alignment \& Sparse Action Enhancement that refine historical reasoning traces and balances the distribution of sparse but critical actions, leading to more coherent planning and better generalization in complex UI tasks. Our contributions include the publish of SOTA open-source UI agents, comprehensive data cleaning protocols and a novel self-evolving framework for improving navigation performance, which encourage further research and development in the community. Code is available at https://github.com/antgroup/UI-Venus.

06

A Survey on Diffusion Language Models

Diffusion Language Models (DLMs) are rapidly emerging as a powerful and promising alternative to the dominant autoregressive (AR) paradigm. By generating tokens in parallel through an iterative denoising process, DLMs possess inherent advantages in reducing inference latency and capturing bidirectional context, thereby enabling fine-grained control over the generation process. While achieving a several-fold speed-up, recent advancements have allowed DLMs to show performance comparable to their autoregressive counterparts, making them a compelling choice for various natural language processing tasks. In this survey, we provide a holistic overview of the current DLM landscape. We trace its evolution and relationship with other paradigms, such as autoregressive and masked language models, and cover both foundational principles and state-of-the-art models. Our work offers an up-to-date, comprehensive taxonomy and an in-depth analysis of current techniques, from pre-training strategies to advanced post-training methods. Another contribution of this survey is a thorough review of DLM inference strategies and optimizations, including improvements in decoding parallelism, caching mechanisms, and generation quality. We also highlight the latest approaches to multimodal extensions of DLMs and delineate their applications across various practical scenarios. Furthermore, our discussion addresses the limitations and challenges of DLMs, including efficiency, long-sequence handling, and infrastructure requirements, while outlining future research directions to sustain progress in this rapidly evolving field. Project GitHub is available at https://github.com/VILA-Lab/Awesome-DLMs.