NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-08-12ENGLISH EDITION
This issue
—
All time
—

Twitter

6 stories
01

Skywork_ai_Launches Matrix-Game 2.0: First Open-Source Real-Time Long-Sequence World Model

Skywork_ai has launched Matrix-Game 2.0, claiming it to be the first open-source, real-time, long-sequence interactive world model. This model achieves a performance of 25 frames per second and supports minutes-long interactions. It aims to challenge the non-open-source nature of DeepMind's Genie 3, providing a fully open solution for the AI world model domain.

02

Thom_Wolf_Significant Progress in Open-Source Visual Reasoning with GLM-4.5V

Thomas Wolf expresses excitement over significant progress in visual reasoning for open-source models, emphasizing its critical importance for Vision-Language Models (VLMs) even more than for text models. He highlights Z.ai's introduction of GLM-4.5V, noting its breakthrough in open-source visual reasoning, achieving state-of-the-art performance and dominating across 41 benchmarks, signaling serious advancements in open-source VLMs.

03

Miles_Brundage_Claude Sonnet 4 Context Window Significantly Expanded

Anthropic's Claude Sonnet 4 model has received a significant context window upgrade on its API, now supporting 1 million tokens, a fivefold increase. This enhancement allows users to process over 75,000 lines of code or hundreds of documents in a single request, greatly improving the model's ability to handle long texts and complex tasks, providing developers with a more powerful tool.

04

allen_ai_Launches MolmoAct: An Open Action Reasoning Model for Physical World Actions

Allen AI has announced the release of MolmoAct, a new, fully open Action Reasoning Model (ARM). This model is designed to enable AI models that operate in the physical world to process and execute human instructions, facilitating more intelligent interactions and automation. The introduction of MolmoAct represents a significant advancement in the fields of embodied AI and robotics control, providing researchers with an open-source tool for further development.

05

sama_Sam Altman Details Compute Prioritization for GPT-5 and Future Expansion Plans

OpenAI CEO Sam Altman outlined the compute prioritization strategy for the coming months due to increased demand from GPT-5. He stated that the company will first ensure more usage for current paying ChatGPT users, then prioritize existing API demand, followed by improving the free tier of ChatGPT, and finally addressing new API demand. Altman also revealed plans to double OpenAI's compute fleet within the next five months to alleviate current compute constraints.

06

NaveenGRao_Improving AI Agent Reliability and New Generative Model Evaluation Method

Naveen Rao shared and commented on Jonathan Frankle's insights regarding generative model evaluation. Frankle highlights the limitations of current LLM judges, including their slow speed, high cost, non-determinism, and lack of calibration. He introduces PGRM as a novel evaluation method designed to overcome these drawbacks, ultimately contributing to the enhanced reliability of AI agents.

wechat

6 stories
01

Online Reinforcement Learning + Flow Matching Model! Flow-GRPO: The First Online RL-Driven Flow Matching Generative Model

Flow-GRPO marks a pioneering advancement by integrating online reinforcement learning into Flow Matching generative models. This innovative framework transforms deterministic Ordinary Differential Equation (ODE) samplers into equivalent Stochastic Differential Equation (SDE) samplers, thereby introducing the necessary stochasticity for online RL application and optimizing the denoising process. By incorporating the GRPO algorithm, Flow-GRPO effectively addresses the inherent challenges of Flow Models in handling complex scenarios such as multi-object composition, spatial relationships, and accurate text rendering. Experimental results demonstrate that Flow-GRPO significantly enhances the accuracy and human preference alignment of text-to-image (T2I) generation. Furthermore, its novel Denoising Reduction strategy substantially accelerates the training process without compromising output quality. This breakthrough opens new avenues for developing high-quality generative models capable of producing intricate and precise visual content.

02

Tencent Open-Sources Stand-In! A New Breakthrough in Lightweight Identity-Preserving Video Generation

Tencent has open-sourced Stand-In, a lightweight, plug-and-play framework designed to address the challenges of large training parameters and insufficient compatibility in high-fidelity identity-preserving video generation. By integrating a conditional image branch and a restricted self-attention mechanism, Stand-In achieves precise identity control and cross-branch information interaction with minimal training samples and negligible additional parameters. This framework demonstrates state-of-the-art performance in identity-preserving text-to-video generation, excelling in both video quality and identity consistency. Furthermore, Stand-In boasts strong compatibility, seamlessly extending to various applications such as topic-based generation, pose-guided video creation, stylization, and face swapping, showcasing its immense potential in the generative AI domain.

03

Physics' "AlphaGo Moment"? AI Solves 40-Year Unfinished Problem, Top Physicists Stunned

Artificial intelligence has achieved groundbreaking advancements in physics, marking what is being called the "AlphaGo moment for physics." AI successfully designed counter-intuitive and highly complex optical component layouts, significantly boosting the sensitivity of the LIGO gravitational wave detector by 10-15%, a remarkable feat that addresses a challenge puzzling physicists for decades. Furthermore, AI independently redesigned a quantum entanglement experiment, discovered a more accurate dark matter formula than those previously proposed by human scientists, and even independently rediscovered "Lorentz symmetry," a fundamental cornerstone of Einstein's theory of relativity, without prior physical knowledge. These compelling cases demonstrate that AI is rapidly evolving from a mere computational tool into an indispensable and powerful scientific collaborator, signaling the imminent arrival of an era where AI profoundly assists in discovering entirely new physics.

04

GitHub No Longer Independent: CEO Resigns, Microsoft Integrates into CoreAI, Signaling a Paradigm Shift for Developers

GitHub CEO Thomas Dohmke has announced his resignation, and GitHub will no longer operate independently, instead being fully integrated into Microsoft's newly formed CoreAI engineering group. This strategic move signifies GitHub's transformation from a mere code hosting platform into Microsoft's "AI agent factory" and "AI training ground," aiming to drive an "AI-first" software development paradigm through tools like Copilot. Microsoft's deep integration of GitHub into its AI division suggests a future where developers may increasingly supervise AI-generated code rather than writing it from scratch. This profound change is not merely a personnel shift but a complete reorientation of the software development model. Developers must adapt to GitHub's new role as a core component of Microsoft's AI arsenal. The article highlights GitHub's significant growth under Dohmke's leadership, with over 1.5 billion developers and 1 billion repositories, and Copilot's success as a leading AI coding assistant. Dohmke's departure to pursue new entrepreneurial ventures underscores the completion of his mission to lead GitHub into the AI era. This integration solidifies Microsoft's vision of embedding AI assistants across its product ecosystem, positioning GitHub at the forefront of this "Copilotization" strategy.

05

LLMs Overcomplicate Simple Tasks, Karpathy Frustrated: Some Tasks Don't Need That Much Thought

The article highlights a growing issue where large language models (LLMs), despite their advanced “deep thinking” capabilities enabled by reasoning and Chain of Thought, increasingly overcomplicate simple tasks. Experts like Andrej Karpathy observe that LLMs, by default, exhibit an excessive “agentic” tendency, engaging in lengthy reasoning processes for straightforward queries. This leads to slow responses and inefficiency, particularly evident in coding and image editing applications. The phenomenon is attributed to LLMs being heavily optimized for long-duration, complex task benchmarks, causing them to treat all tasks as high-stakes “exams.” The piece emphasizes the critical need for users to have a mechanism to explicitly specify the required depth of thought for a given task, thereby preventing unnecessary overthinking and enhancing the practical utility of LLMs. This adjustment is crucial for improving user experience and operational efficiency across various applications.

06

Zhipu Open-Sources GLM-4.5V, Unveiling Visual Reasoning Capabilities Rivaling OpenAI's

Zhipu AI has open-sourced its flagship visual reasoning model, GLM-4.5V, demonstrating exceptional performance across various visual tasks. The model notably defeated 99.99% of human players in the "GeoGuessr" game, showcasing its powerful visual inference capabilities from subtle cues. GLM-4.5V excels in advanced image recognition, long video understanding, GUI Agent applications, and complex document and chart interpretation. With 106 billion total parameters, it employs an advanced architecture and a three-stage training strategy, achieving state-of-the-art open-source performance across 41 public visual multimodal benchmarks. This open-sourcing initiative aims to shift the focus of AI development from benchmark competition to practical application, providing developers with a robust multimodal foundation model to collaboratively shape the future of AI.