NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-06-10ENGLISH EDITION
This issue
—
All time
—

Twitter

6 stories
01

OpenAI_OpenAI_o3-pro Fully Rolls Out to Pro Users

OpenAI has announced that its latest model, o3-pro, is now fully rolling out to all ChatGPT Pro users and API developers. This update aims to enhance user experience and development efficiency, further expanding AI application scenarios. The launch of o3-pro is expected to bring more powerful features and stable services to professional users, solidifying OpenAI's leading position in the AI field.

02

ArtificialAnlys_OpenAI Significantly Reduces o3 Model Pricing

OpenAI has announced an 80% reduction in the pricing of its o3 model, making it price-competitive with Gemini 2.5 Pro and Claude 4 Sonnet, and eight times cheaper than Claude 4 Opus. The o3 model is now priced at $2/$8 per 1M input/output tokens, down from $8/$40. Additionally, it offers a 75% discount for cached input tokens, significantly enhancing its market competitiveness.

03

MistralAI_Launches First Reasoning Model Magistral

Mistral AI recently announced the launch of Magistral, its first reasoning model. This model is specifically designed to excel in domain-specific, transparent, and multilingual reasoning, aiming to provide more efficient and reliable AI inference capabilities. This release marks a new advancement for Mistral AI in the field of AI model development, expected to play a significant role in various application scenarios.

04

LangChainAI_Uber Uses LangGraph to Build AI Developer Agents for Code Fixes

LangChainAI shared how Uber leveraged the LangGraph framework to build AI developer agents. These agents are capable of generating thousands of code fixes daily, serving an organization of 5,000 developers working with hundreds of millions of lines of code, and have already saved over 21,000 hours of development time. This case highlights the significant potential of AI in enhancing software development efficiency.

05

Kling_ai_Pengfei Wan to Share Kling Video Generation Models at CVPR

Kling AI announced that Pengfei Wan, Head of Kling Video Generation Models, will deliver a keynote speech at CVPR, the leading annual computer vision event hosted by IEEE. Wan's presentation, titled "An Introduction to Kling and Our Research towards More Powerful Video Generation Models," will delve into Kling's latest breakthroughs and cutting-edge advancements in video generation technology. The event is scheduled for June 11, 2025, from 9:00-17:00 (UTC-5), aiming to provide exceptional value to students, researchers, and industry professionals.

06

summeryue0_Scale AI Releases LLM Red Teaming and System Safety Roadmap

Summer Yue and Zifan (Sail) Wang from Scale AI, representing SEAL and Red Team, have jointly published a significant position paper. This document outlines key learnings from their extensive experience in red teaming Large Language Models (LLMs), addressing what truly matters, what remains to be explored, and how model safety is intrinsically linked to broader system safety and continuous monitoring. The paper emphasizes crucial research priorities for red teaming frontier AI models and proposes a comprehensive roadmap for achieving robust system-level safety and effective AI monitoring, integrating diverse perspectives from practitioners to researchers.

GitHub

5 stories
01

🌟 Awesome LLM Apps

The "Awesome LLM Apps" GitHub repository presents a meticulously curated collection of practical large language model (LLM) applications. These applications are ingeniously built utilizing advanced techniques such as Retrieval-Augmented Generation (RAG), sophisticated AI Agents, collaborative Multi-agent Teams, Multi-Context Processing (MCP), and intuitive Voice Agents. The repository showcases the versatile integration of leading LLM providers like OpenAI, Anthropic, and Google, alongside powerful open-source models including DeepSeek, Qwen, and Llama, which can even be run locally. It illustrates how LLMs can address real-world challenges across diverse domains, from analyzing code repositories to managing email inboxes. This project's core objective is to offer tangible, innovative LLM application examples, thereby accelerating the practical deployment and advancement of large model technologies across various industries. Furthermore, it actively encourages community contributions, aiming to cultivate a vibrant and comprehensive open-source ecosystem for LLM-powered solutions.

02

Boltz

Boltz is a family of models designed for biomolecular interaction prediction, with Boltz-2 being the latest foundational model. It surpasses AlphaFold3 and Boltz-1 by jointly modeling complex structures and binding affinities. Boltz-2 is the first deep learning model to achieve accuracy comparable to physics-based free-energy perturbation (FEP) methods, while being 1000 times faster. This breakthrough makes accurate in silico screening practical for early-stage drug discovery, including hit-discovery and ligand optimization. The Boltz models and code are open-sourced under the MIT license, available for both academic and commercial use.

03

✨ YouTube Transcript API ✨

The YouTube Transcript API is a Python library designed to retrieve transcripts and subtitles for YouTube videos. It supports both automatically generated subtitles and multi-language translation, operating efficiently without the need for headless browsers, which significantly improves performance. The API also features proxy support to circumvent IP blocks, cookie authentication for age-restricted content, and offers various output formats like JSON and SRT. Additionally, it includes a command-line interface, greatly simplifying the process for developers and users to access and process YouTube video content.

04

Open-source Large Language Model Handbook

This project presents a dedicated open-source large language model (LLM) tutorial, specifically designed for beginners in China and optimized for Linux platforms. It provides comprehensive, full-process guidance encompassing essential skills such as environment configuration, local deployment, and efficient fine-tuning for a wide array of open-source LLMs. By simplifying the complex deployment, usage, and application workflows, the initiative aims to make advanced LLM technologies more accessible to a broader audience of students and researchers. The tutorial covers mainstream models like LLaMA, ChatGLM, and InternLM, offering practical instructions on command-line invocation, setting up online demonstrations, and integrating with frameworks like LangChain. Furthermore, it delves into advanced topics such as distributed full fine-tuning, LoRA, and P-tuning methods. This resource is crucial for fostering the adoption of open-source, free large models, enabling learners to seamlessly incorporate them into their studies and future professional endeavors.

05

MiniCPM

MiniCPM is an ultra-efficient series of large language models designed for edge devices, co-developed by ModelBest, Tsinghua University, and Renmin University of China. It achieves exceptional efficiency through innovative model architectures like InfLLM v2 sparse attention, efficient learning algorithms such as BitCPM 3-value quantization, and optimized inference systems like CPM.cu. While maintaining state-of-the-art performance for its size, MiniCPM models deliver over 5x generation speedup on typical edge chips. They also surpass similarly sized and even larger models in tasks like tool calling, code interpretation, and long-context processing, offering a powerful solution for edge AI applications.

wechat

5 stories
01

NVIDIA and HKU Jointly Innovate Visual Attention Mechanism! GSPN Accelerates High-Resolution Generation by Over 84x

The University of Hong Kong and NVIDIA have jointly introduced the Generalized Spatial Propagation Network (GSPN), revolutionizing visual attention mechanisms. GSPN employs 2D linear propagation combined with "stability-context conditions," reducing the computational complexity for high-resolution image processing from O(N²) or O(N) to O(√N) while fully preserving spatial coherence. This technology achieves significant improvements in both efficiency and accuracy across image classification and generation tasks. Notably, in the SD-XL model, GSPN accelerates 16K×8K image generation by over 84 times. GSPN inherently handles positional information without explicit embeddings and ensures stability, demonstrating immense potential for both academic and industrial applications. This innovation paves new directions for multimodal models and real-time visual applications.

02

Huawei Achieves New AI Computing Power Milestone: 98% Availability in Ten-Thousand-Card Cluster Training, Second-Level Recovery, Minute-Level Diagnostics

Huawei has achieved groundbreaking advancements in AI computing power, setting new industry benchmarks for large-scale AI clusters. Through the implementation of an innovative "3+3" dual-dimension technical system, the company has significantly enhanced the availability, recovery speed, and linearity of its ten-thousand-card AI clusters. This comprehensive system integrates three fundamental capabilities—full-stack observability, advanced fault diagnosis, and robust self-healing fault tolerance—with three critical business support capabilities: optimized cluster linearity, rapid training recovery, and rapid inference recovery. As a result, Huawei has demonstrated an impressive 98% training availability for its ten-thousand-card clusters, achieving fault recovery in mere seconds, diagnostics within minutes, and linearity exceeding 95%. These pioneering innovations not only substantially boost the efficiency and stability of large model training and inference but also effectively address long-standing critical challenges associated with deploying and managing large-scale AI applications.

03

Peking University and UC Berkeley Unveil IDA-Bench: A New Benchmark Exposing LLM Agents' Limitations in Iterative Data Analysis, Top Models Score Only 40%

Peking University and UC Berkeley have jointly introduced IDA-Bench, a novel benchmark designed to simulate real-world, iterative, and exploratory data analysis scenarios, specifically evaluating large language model agents' performance under multi-turn, dynamic instructions. Unlike traditional single-turn evaluations, IDA-Bench exposes the challenges of continuous interaction. Test results reveal that even top models like Claude-3.7 and Gemini-2.5 Pro achieve only a 40% task success rate, significantly below expectations. The research highlights current agents' profound difficulty in balancing strict instruction adherence with necessary autonomous reasoning, often exhibiting "overconfident" or "overcautious" behaviors that lead to critical task failures. This comprehensive evaluation underscores the critical need for substantial improvements in LLM agents' understanding, instruction following, and interactive capabilities to truly become reliable and effective data analysis assistants in complex, real-world settings.

04

Let AI Design Chips! Chinese Academy of Sciences Launches 'QiMeng' for Fully Automated Chip Design Flow

The Institute of Computing Technology, Chinese Academy of Sciences, in collaboration with the Institute of Software, has unveiled 'QiMeng,' a fully automated design system for processor chips and foundational software, powered by large models and other AI technologies. This system can autonomously complete chip hardware and software design, partially or entirely surpassing human expert levels. It has successfully designed RISC-V CPUs automatically, achieving performance comparable to ARM Cortex A53. 'QiMeng' employs a three-tiered architecture comprising domain-specific large models, intelligent agents, and an application layer. It addresses challenges such as data scarcity, correctness, and solution scale through an iterative evolution approach, promising to significantly enhance chip design efficiency, shorten development cycles, enable rapid customization, and fundamentally transform the paradigm of processor chip hardware and software design.

05

Are Large Language Models 'Observing the World from a Cave'? RL Expert Warns of LLM's Fatal Flaws

University of California, Berkeley reinforcement learning expert Sergey Levine argues that current Large Language Models (LLMs) do not learn directly from the world but rather indirectly "scan" the "projections" of human thought processes from internet text, akin to observers in Plato's Cave. He contends that LLMs' success stems from "reverse engineering" human cognitive processes, not from genuinely understanding the world or learning from direct experience. Levine questions the limitations of LLMs in physical world comprehension and autonomous skill acquisition, proposing that future AI development must explore new methods for acquiring representations directly from physical experience. This approach is crucial for achieving truly flexible and adaptive intelligence, moving beyond merely replicating the "shadows" of human minds. He contrasts this with the relative lack of success in video models, despite video containing richer real-world information. The core issue, according to Levine, is that LLMs are mimicking human intelligence by observing its output, not by developing their own understanding of reality. This perspective challenges the current trajectory of AGI research, advocating for AI systems that can learn and adapt from their own interactions with the physical world, rather than relying on human-mediated data.

huggingface

6 stories
01

Reinforcement Pre-Training

In this work, we introduce Reinforcement Pre-Training (RPT) as a new scaling paradigm for large language models and reinforcement learning (RL). Specifically, we reframe next-token prediction as a reasoning task trained using RL, where it receives verifiable rewards for correctly predicting the next token for a given context. RPT offers a scalable method to leverage vast amounts of text data for general-purpose RL, rather than relying on domain-specific annotated answers. By incentivizing the capability of next-token reasoning, RPT significantly improves the language modeling accuracy of predicting the next tokens. Moreover, RPT provides a strong pre-trained foundation for further reinforcement fine-tuning. The scaling curves show that increased training compute consistently improves the next-token prediction accuracy. The results position RPT as an effective and promising scaling paradigm to advance language model pre-training.

02

MiniCPM4: Ultra-Efficient LLMs on End Devices

This paper introduces MiniCPM4, a highly efficient large language model (LLM) designed explicitly for end-side devices. We achieve this efficiency through systematic innovation in four key dimensions: model architecture, training data, training algorithms, and inference systems. Specifically, in terms of model architecture, we propose InfLLM v2, a trainable sparse attention mechanism that accelerates both prefilling and decoding phases for long-context processing. Regarding training data, we propose UltraClean, an efficient and accurate pre-training data filtering and generation strategy, and UltraChat v2, a comprehensive supervised fine-tuning dataset. These datasets enable satisfactory model performance to be achieved using just 8 trillion training tokens. Regarding training algorithms, we propose ModelTunnel v2 for efficient pre-training strategy search, and improve existing post-training methods by introducing chunk-wise rollout for load-balanced reinforcement learning and data-efficient tenary LLM, BitCPM. Regarding inference systems, we propose CPM.cu that integrates sparse attention, model quantization, and speculative sampling to achieve efficient prefilling and decoding. To meet diverse on-device requirements, MiniCPM4 is available in two versions, with 0.5B and 8B parameters, respectively. Sufficient evaluation results show that MiniCPM4 outperforms open-source models of similar size across multiple benchmarks, highlighting both its efficiency and effectiveness. Notably, MiniCPM4-8B demonstrates significant speed improvements over Qwen3-8B when processing long sequences. Through further adaptation, MiniCPM4 successfully powers diverse applications, including trustworthy survey generation and tool use with model context protocol, clearly showcasing its broad usability.

03

CyberV: Cybernetics for Test-time Scaling in Video Understanding

Current Multimodal Large Language Models (MLLMs) may struggle with understanding long or complex videos due to computational demands at test time, lack of robustness, and limited accuracy, primarily stemming from their feed-forward processing nature. These limitations could be more severe for models with fewer parameters. To address these limitations, we propose a novel framework inspired by cybernetic principles, redesigning video MLLMs as adaptive systems capable of self-monitoring, self-correction, and dynamic resource allocation during inference. Our approach, CyberV, introduces a cybernetic loop consisting of an MLLM Inference System, a Sensor, and a Controller. Specifically, the sensor monitors forward processes of the MLLM and collects intermediate interpretations, such as attention drift, then the controller determines when and how to trigger self-correction and generate feedback to guide the next round. This test-time adaptive scaling framework enhances frozen MLLMs without requiring retraining or additional components. Experiments demonstrate significant improvements: CyberV boosts Qwen2.5-VL-7B by 8.3% and InternVL3-8B by 5.5% on VideoMMMU, surpassing the competitive proprietary model GPT-4o. When applied to Qwen2.5-VL-72B, it yields a 10.0% improvement, achieving performance even comparable to human experts. Furthermore, our method demonstrates consistent gains on general-purpose benchmarks, such as VideoMME and WorldSense, highlighting its effectiveness and generalization capabilities in making MLLMs more robust and accurate for dynamic video understanding. The code is released at https://github.com/marinero4972/CyberV.

04

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyVivid, a multi-subject video customization framework that enables flexible and identity-consistent generation. To establish accurate correspondences between subject images and textual entities, we design a VLLM-based text-image fusion module that embeds visual identities into the textual space for precise grounding. To further enhance identity preservation and subject interaction, we propose a 3D-RoPE-based enhancement module that enables structured bidirectional fusion between text and image embeddings. Moreover, we develop an attention-inherited identity injection module to effectively inject fused identity features into the video generation process, mitigating identity drift. Finally, we construct an MLLM-based data pipeline that combines MLLM-based grounding, segmentation, and a clique-based subject consolidation strategy to produce high-quality multi-subject data, effectively enhancing subject distinction and reducing ambiguity in downstream video generation. Extensive experiments demonstrate that PolyVivid achieves superior performance in identity fidelity, video realism, and subject alignment, outperforming existing open-source and commercial baselines.

05

GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior

Multimodal Large Language Models (MLLMs) have shown great potential in revolutionizing Graphical User Interface (GUI) automation. However, existing GUI models mostly rely on learning from nearly error-free offline trajectories, thus lacking reflection and error recovery capabilities. To bridge this gap, we propose GUI-Reflection, a novel framework that explicitly integrates self-reflection and error correction capabilities into end-to-end multimodal GUI models throughout dedicated training stages: GUI-specific pre-training, offline supervised fine-tuning (SFT), and online reflection tuning. GUI-reflection enables self-reflection behavior emergence with fully automated data generation and learning processes without requiring any human annotation. Specifically, 1) we first propose scalable data pipelines to automatically construct reflection and error correction data from existing successful trajectories. While existing GUI models mainly focus on grounding and UI understanding ability, we propose the GUI-Reflection Task Suite to learn and evaluate reflection-oriented abilities explicitly. 2) Furthermore, we built a diverse and efficient environment for online training and data collection of GUI models on mobile devices. 3) We also present an iterative online reflection tuning algorithm leveraging the proposed environment, enabling the model to continuously enhance its reflection and error correction abilities. Our framework equips GUI agents with self-reflection and correction capabilities, paving the way for more robust, adaptable, and intelligent GUI automation, with all data, models, environments, and tools to be released publicly.

06

SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems

Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled powerful autonomous agents capable of complex reasoning and multi-modal tool use. Despite their growing capabilities, today's agent frameworks remain fragile, lacking principled mechanisms for secure information flow, reliability, and multi-agent coordination. In this work, we introduce SAFEFLOW, a new protocol-level framework for building trustworthy LLM/VLM-based agents. SAFEFLOW enforces fine-grained information flow control (IFC), precisely tracking provenance, integrity, and confidentiality of all the data exchanged between agents, tools, users, and environments. By constraining LLM reasoning to respect these security labels, SAFEFLOW prevents untrusted or adversarial inputs from contaminating high-integrity decisions. To ensure robustness in concurrent multi-agent settings, SAFEFLOW introduces transactional execution, conflict resolution, and secure scheduling over shared state, preserving global consistency across agents. We further introduce mechanisms, including write-ahead logging, rollback, and secure caches, that further enhance resilience against runtime errors and policy violations. To validate the performances, we built SAFEFLOWBENCH, a comprehensive benchmark suite designed to evaluate agent reliability under adversarial, noisy, and concurrent operational conditions. Extensive experiments demonstrate that agents built with SAFEFLOW maintain impressive task performance and security guarantees even in hostile environments, substantially outperforming state-of-the-art. Together, SAFEFLOW and SAFEFLOWBENCH lay the groundwork for principled, robust, and secure agent ecosystems, advancing the frontier of reliable autonomy.