NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-09ENGLISH EDITION
This issue
—
All time
—

AI Blog

6 stories
01

GPT-5.6 Delivers Frontier Intelligence and Scalable Capabilities

OpenAI has announced GPT-5.6, a new large language model designed to offer increased intelligence per token, improved performance per dollar, and scalable capability on demand for difficult workloads. The model aims to optimize the balance between cost-effectiveness and high-performance computing for developers and enterprise customers. In a related integration, OpenAI's GPT-5.6 has also become the preferred model powering Microsoft 365 Copilot, bringing these advanced language capabilities directly into productivity applications like Word, Excel, PowerPoint, Chat, and Cowork to enable faster execution and higher-quality outputs. (source: https://openai.com/index/gpt-5-6)

02

ChatGPT Work Introduced as an Agent for Ambitious Projects

OpenAI has introduced ChatGPT Work, an autonomous AI agent designed to partner with users on complex and long-horizon professional tasks. The system can take direct action across multiple software applications and files, allowing it to assist continuously over several hours to transform high-level goals into completed outputs. This release marks a shift from a conversational assistant to a functional workplace agent capable of executing complex workflows and managing technical details across diverse software environments. (source: https://openai.com/index/chatgpt-for-your-most-ambitious-work)

03

Meta Superintelligence Progress and the Frontier AI Landscape

SemiAnalysis published an analysis on Meta Superintelligence (MSL), detailing Meta's aggressive strategy to catch up with frontier AI leaders like OpenAI and Anthropic. Despite underwhelming initial results from its Muse Spark model compared to competitors like DeepSeek v4 Pro, Meta is investing heavily in talent, infrastructure, and team rebuilding. Key strategic moves include a $14.3 billion investment in Scale AI to secure specialized engineering talent, competitive compensation packages, and a new 'Tent' datacenter design, positioning Meta to potentially surpass Google in the frontier LLM landscape. (source: https://newsletter.semianalysis.com/p/the-future-of-meta-superintelligence)

04

Identifies Key Reliability Issues in SWE-Bench Pro Benchmark

OpenAI has published an analysis identifying critical reliability and accuracy issues within SWE-Bench Pro, a prominent benchmark used for evaluating AI coding models. The investigation demonstrates that noise within the benchmark can distort model performance metrics, making it difficult to distinguish true algorithmic progress from evaluation variance. The researchers highlight the urgent need for robust, noise-resistant evaluation frameworks to accurately measure agentic coding capabilities in future large language models. (source: https://openai.com/index/separating-signal-from-noise-coding-evaluations)

05

Launches Bio Bug Bounty Program for Frontier Models

OpenAI has launched the Bio Bug Bounty program, a specialized crowdsourcing initiative designed to identify potential biological risks associated with its frontier AI models. The program invites external researchers and security experts to find and report vulnerabilities where systems like GPT-4 could inadvertently assist in synthesizing or replicating hazardous biological materials. Participants are eligible for financial rewards based on severity, aiming to proactively mitigate biosecurity risks before models are widely deployed. (source: https://openai.com/index/bio-bug-bounty)

06

Former Federal Reserve Chair Ben Bernanke Joins Anthropic’s Long-Term Benefit Trust

Anthropic has appointed Nobel laureate and former Federal Reserve Chair Dr. Ben Bernanke to its Long-Term Benefit Trust (LTBT). As an independent governing body for Anthropic's Public Benefit Corporation structure, the LTBT has the authority to appoint board members and advise leadership on societal impacts. Dr. Bernanke will apply his economic expertise to guide Anthropic's research on how advanced artificial intelligence and large language models affect global workforces and economies. (source: https://www.anthropic.com/news/ben-bernanke)

Hacker News

8 stories
01

GPT-5.6

OpenAI has introduced information regarding its new GPT-5.6 model by releasing its deployment safety card and updated developer API guidelines. The documentation highlights the model's advanced reasoning capabilities and its alignment within OpenAI's latest deployment safety framework. This release indicates GPT-5.6 is prepared for developer workflow integration with upgraded mitigation protocols and system robustness standards, reflecting a deliberate focus on safe and scalable frontier model deployment. (source: https://openai.com/index/gpt-5-6/)

02

Muse Spark 1.1

Meta has introduced Muse Spark 1.1, representing the company's shift toward charging for its advanced artificial intelligence capabilities. Engineered as an agentic model, Muse Spark 1.1 is built to support multi-step reasoning workflows and autonomous task execution. Alongside the release, Meta published an evaluation report detailing technical benchmarks and a developer integration guide for the Muse Spark API, establishing a direct commercial competitor in the enterprise AI agent market. (source: https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/)

03

OpenAI faked inability to search training data, hid billions of logs, NYT says

The New York Times has alleged in its ongoing copyright lawsuit that OpenAI fabricated an inability to search its training data and hid billions of system logs. The publisher claims OpenAI systematically misled courts and researchers regarding its data indexing capabilities, which are vital for identifying copyright infringement in its foundation models. These allegations could significantly affect legal discovery and transparency requirements for generative AI developers. (source: https://arstechnica.com/tech-policy/2026/07/openai-faked-inability-to-search-training-data-hid-billions-of-logs-nyt-says/)

04

ChatGPT Work

OpenAI has officially launched ChatGPT Work, a specialized offering designed to integrate its language models into enterprise and team workflows. The solution provides workspace integrations, advanced reasoning, and data privacy protections, guaranteeing that customer data is not used for model training. The release aims to move conversational AI tools into secure, collaborative business operating systems. (source: https://openai.com/index/chatgpt-for-your-most-ambitious-work/)

05

DeepSeek aims to make its own AI chip

Chinese artificial intelligence firm DeepSeek is shifting toward designing proprietary AI chips to minimize dependency on foreign hardware and navigate US export restrictions. This custom silicon strategy aims to integrate hardware and software directly, lowering operational costs for training and running its large-scale models. The transition from an AI software developer to a full-stack hardware-software provider could significantly influence the global semiconductor supply chain. (source: https://www.proactiveinvestors.com/companies/news/1095178/deepseek-makes-pivot-that-should-put-silicon-valley-on-high-alert-1095178.html)

06

Launch HN: Context.dev (YC S26) – API to get structured data from any website

Context.dev has launched its API platform to convert unstructured web pages into clean, structured data packages optimized for AI agents and software applications. The YC S26-backed tool extracts markdown, rendered HTML, images, and screenshots alongside domain-level brand context such as logos, descriptions, and social links, simplifying web scraping workflows. (source: https://www.context.dev)

07

GLM 5.2 is nearly as accurate as a human book keeper

The performance evaluation of the GLM 5.2 language model reveals significant advancements in specialized, domain-specific tasks, particularly in financial bookkeeping and value-added tax (VAT) calculations. Benchmark testing indicates that the model operates at an accuracy level that nearly matches a professional human bookkeeper. This achievement showcases the potential for automating highly structured, precision-critical, and regulated financial workflows using large language models. (source: https://toot-books.pages.dev/blog/glm-5-2-vat-benchmark)

08

Show HN: Reverse-engineering web apps into agent tools

A new browser-based development tool has been introduced that reverse-engineers web applications into executable tools for AI agents. Operating directly inside authenticated web applications, the tool monitors internal API calls and translates them dynamically, serving as a self-updating Model Context Protocol (MCP) server. This approach simplifies enterprise software automation and expands the integration capabilities of AI agents beyond typical retrieval-augmented generation. (source: https://news.ycombinator.com/item?id=48847834)

Twitter

8 stories
01

OpenAI Launches Major ChatGPT Updates Including Desktop App and Hosted Sites

OpenAI CEO Sam Altman announced a series of major updates to the ChatGPT ecosystem. The release introduces ChatGPT Work, an enterprise-focused collaborative agent, alongside a dedicated ChatGPT desktop application. Additionally, the platform is expanding into integrated web development by adding direct site-hosting capabilities. These updates transition ChatGPT from a basic conversational interface into a versatile, cross-platform hub for collaborative content creation, site management, and professional workflows. (source: https://x.com/sama/status/2075264378962907597)

02

OpenAI Introduces Most Advanced Artificial Intelligence Model to Date

OpenAI announced the launch of its latest and most capable artificial intelligence model, detailing its architecture, reasoning advancements, and performance benchmarks in a comprehensive blog post. Alongside this release, OpenAI introduced the Sol and Terra models to prioritize operational efficiency, lowering token consumption and reducing dollars-per-task costs for enterprises. The updates collectively establish new industry benchmarks for reasoning, coding, cybersecurity, and scientific applications. (source: https://x.com/sama/status/2075266471316615436)

03

Prime Intellect Secures $130 Million Series A Funding for Open Superintelligence

Prime Intellect successfully secured $130 million in a Series A funding round to accelerate the development of its Open Superintelligence Stack. The investment round was led by Radical Ventures, with significant financial participation from semiconductor and technology industry leaders including NVIDIA and Intel Capital. The capital will be utilized to design an accessible, open-source infrastructure and modular stack architecture, aiming to democratize high-performance computing and scale advanced artificial intelligence capabilities. (source: https://x.com/Thom_Wolf/status/2075224964572049744)

04

GPT-5.6 Sol Sets New State of the Art Milestone on ARC-AGI-3 Benchmark

The GPT-5.6 Sol model has achieved a new state-of-the-art milestone by securing a 7.8% success rate on the ARC-AGI-3 benchmark, designed by François Chollet. This marks the first time a verified frontier AI model has successfully solved an ARC-AGI-3 task. Unlike traditional evaluations that measure pattern matching, the ARC benchmark evaluates true machine reasoning and generalization in highly abstract, novel scenarios. (source: https://x.com/fchollet/status/2075270979845276068)

05

Runway Launches Developer Platform Featuring Advanced Media Generation Models

Runway has launched Runway Dev, an enterprise-grade platform that centralizes the company's advanced generative media models in a single environment. Designed for developers and corporate users, the platform integrates Seed Audio 1.0 and Seedance Mini to streamline multimedia workflows at scale. Runway aims to provide robust infrastructure for high-performance audio and video synthesis, facilitating deployment into professional production pipelines. (source: https://x.com/c_valenzuelab/status/2075244261944000604)

06

Dr. Ben Bernanke Appointed To Anthropic Long-Term Benefit Trust

Anthropic announced the appointment of Dr. Ben Bernanke, former Chair of the Federal Reserve, as the newest member of its Long-Term Benefit Trust. The Trust acts as an independent governing body responsible for appointing Anthropic's board of directors. Dr. Bernanke's addition is intended to strengthen corporate oversight and ensure the company remains aligned with its safety standards and long-term benefit mission during rapid scaling. (source: https://x.com/AnthropicAI/status/2075257492716879967)

07

Google Cloud Announces General Availability Of AlphaEvolve

Google has announced the general availability of AlphaEvolve on Google Cloud. Developed in collaboration with Google DeepMind, the tool leverages Gemini-powered evolutionary algorithms to automate and scale computational discovery and optimization. By integrating DeepMind's research with Google's scalable cloud infrastructure, developers and enterprise clients can deploy sophisticated evolutionary computing models to solve highly complex industrial and scientific optimization problems. (source: https://x.com/Google/status/2075252213400649807)

08

Samaya AI Introduces FrontierFinance Benchmark for Agentic Systems

Samaya AI has launched FrontierFinance, a new benchmarking framework designed to evaluate the financial reasoning capabilities of autonomous, agentic AI systems. Recognized as a highly challenging evaluation standard in this domain, FrontierFinance measures how effectively autonomous agents process complex financial datasets, conduct market analyses, execute decision-making tasks, and generate precise financial reports under rigorous testing conditions. (source: https://x.com/ericschmidt/status/2075292806017659044)

huggingface

8 stories
01

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Researchers have introduced Single-Rollout Asynchronous Optimization (SAO), a reinforcement learning (RL) framework designed to optimize large language models for long-horizon agentic tasks. To address stability and off-policy challenges in asynchronous RL, SAO replaces traditional group-wise sampling with a single-rollout sampling strategy, utilizing one rollout per prompt. It incorporates value-model training designs and a strict double-side token-level clipping strategy. Evaluated on agentic coding and reasoning benchmarks like SWE-Bench Verified, BeyondAIME, and IMOAnswerBench, SAO consistently outperformed GRPO and its variants. The pipeline was successfully deployed in training the open GLM-5.2 model (750B-A40B). (source: https://huggingface.co/papers/2607.07508)

02

Infinite Worlds with Versatile Interactions

Researchers have launched LingBot-World 2.0 (also known as LingBot-World-Infinity), an advanced world modeling iteration designed for interactive, unbounded horizon video generation. The system leverages a causal pretraining paradigm and includes a distilled real-time variant capable of driving 720p video streams at 60 fps. It integrates an agentic harness where a pilot agent plans character behaviors and a director agent synthesizes new environment elements. The release features a primary 14B model alongside a lightweight 1.3B model suitable for single-GPU deployment. (source: https://huggingface.co/papers/2607.07534)

03

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Researchers have developed LingBot-Video, a Diffusion Transformer (DiT)-based video pretraining paradigm optimized specifically for embodied intelligence. To balance physical realism with modeling capacity, the authors employ a Mixture-of-Experts (MoE) architecture instead of a dense framework. The model is trained on a video engine combining standard web footage with robot-oriented datasets featuring manipulation, navigation, and egocentric viewpoints. A multi-dimensional reward system aligns outputs for physical validity. LingBot-Video is introduced as an open-source MoE video foundation model. (source: https://huggingface.co/papers/2607.07675)

04

Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation

Researchers have introduced LaMem-VLA, a vision-language-action framework designed to overcome long-horizon manipulation failures by integrating memory directly into the latent embedding space. It reconstructs historical experience into short-term and long-term latent tokens via a four-part system: a curator, a seeker, a condenser, and a weaver. By keeping memory within the native continuous latent space, these experience tokens actively participate in the VLA model's multimodal reasoning. Experiments on SimplerEnv and LIBERO benchmarks demonstrate improved performance. (source: https://huggingface.co/papers/2607.07608)

05

Automating the Design of Embodied Agent Architectures

Researchers have investigated Agent Architecture Search (AAS) for physical systems to automate modular designs of perception, memory, and planning. They introduced AgentCanvas, a typed-graph runtime to host embodied executors, alongside KDLoop, a search procedure cycling through proposal, critique, experiment, and distillation. Evaluating three AAS variants across four embodied executors revealed success-rate gains in vision-language navigation, embodied QA, and language-conditioned manipulation, while exposing critical challenges like noisy rollout signals and local edit basins. (source: https://huggingface.co/papers/2606.30111)

06

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

Researchers have released AgentLens, an open-source evaluation benchmark for interactive coding agents designed to measure whole-trajectory performance rather than binary success metrics. AgentLens pairs formal code verification with LLM-generated trajectory reviews and side-by-side comparisons to explain agent failures, tool usage, and recovery strategies. The benchmark helps developers diagnose model behavior and catch product regressions. The codebase has been made public on GitHub. (source: https://huggingface.co/papers/2607.06624)

07

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Researchers have introduced RoboDojo, a unified sim-and-real benchmark for evaluating generalist robot manipulation policies across simulation and physical deployments. RoboDojo contains 42 simulation tasks built in Isaac Sim alongside 18 real-world tasks evaluating memory, precision, long-horizon execution, and instruction following. It includes RoboDojo-RealEval, a standardized real-world setup with remote cloud access, standardized hardware, and reset protocols. The benchmark is integrated into XPolicyLab, where 30 policies have been evaluated. (source: https://huggingface.co/papers/2607.04434)

08

Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMs

Researchers have introduced Splash, a mask-isolated tactile alignment learning framework that adds physical tactile alignment to multimodal large language models without degrading vision-language capabilities. Splash isolates pretrained parameters into stable critical and trainable dormant subspaces. The dormant subspace is selectively updated to align tactile inputs with the language model while the critical subspace preserves general visual reasoning, avoiding catastrophic forgetting. Splash shows leading performance on visuo-tactile benchmarks TVL, SSVTP, and TacQuad. (source: https://huggingface.co/papers/2607.00302)