NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-08ENGLISH EDITION
This issue
—
All time
—

AI Blog

3 stories
01

GPT-Live Model Powering Natural Human-AI Voice Interaction

OpenAI has introduced GPT-Live, a next-generation voice model built for seamless and natural human-AI vocal interaction. The model is now actively powering the ChatGPT Voice feature directly within existing ChatGPT platform interfaces. Designed to improve on real-time conversational latency and interactive fluidity, GPT-Live allows users to engage in more fluid vocal dialogue. This release represents a step forward in real-time multimodal engagement by upgrading the underlying audio and conversational capabilities of the voice assistant interface. (source: https://openai.com/index/introducing-gpt-live)

02

Anthropic Confidential IPO Filing Signals Strong Financial Momentum and Profitability

SemiAnalysis released a financial analysis showing Anthropic is on track to achieve a projected 3Q26 profit exceeding $1 billion, following its confidential IPO filing on June 1, 2026. According to the report, Anthropic and OpenAI combined now generate roughly $100 billion in annual recurring revenue. Anthropic is leading in profitable B2B monetization, heavily driven by the enterprise adoption of Claude Code. As the first major Western AI laboratory to initiate a confidential IPO filing, Anthropic's unit economics and model position it to self-fund future large-scale computing infrastructure. (source: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-over-1b-the)

03

MUFG Partners with OpenAI to Build AI-Native Organization

Mitsubishi UFJ Financial Group (MUFG) has partnered with OpenAI to deploy ChatGPT Enterprise globally across its operations. The major financial services institution plans to utilize large language models to optimize internal workflows, increase employee productivity, and deploy novel generative AI-powered financial services at scale. This deployment highlights a major integration of OpenAI's enterprise-grade secure AI platform within the highly regulated global financial sector, allowing MUFG to build an AI-native organization for its workforce. (source: https://openai.com/index/mufg)

Hacker News

8 stories
01

SWE-1.7 Reach Near GPT 5.5 and Opus Intelligence

Cognition has released its latest software engineering model, SWE-1.7, designed to execute complex, multi-file edits and logic across real-world codebases. The model demonstrates advanced autonomous software engineering capabilities and long-horizon reasoning. According to Cognition, its performance metrics position its cognitive capabilities close to anticipated next-generation systems like GPT 5.5 and Claude Opus, significantly reducing the gap between artificial agents and elite human software developers in handling bug localization and code modification tasks. (source: https://cognition.com/blog/swe-1-7)

02

GPT-Live

OpenAI has introduced GPT-Live, a real-time, highly interactive version of its generative models designed for live-streaming environments. This new system targets ultra-low latency processing and immediate conversational feedback to transition from static query-and-response paradigms to continuous, multi-turn dialogue. Key architectural components include stream-processing APIs, continuous learning structures, and adaptive safety systems that monitor and analyze active data streams in real time. (source: https://openai.com/index/introducing-gpt-live/)

03

Grok 4.5

xAI has announced the release of Grok 4.5, the latest iteration of its proprietary large language model series. The updated model features structural enhancements across core reasoning, computer coding, and complex mathematical processing. According to xAI, Grok 4.5 provides a deeper contextual understanding, improved training efficiency, refined alignment methodologies, and real-time information retrieval to support more precise instruction execution for developer and conversational applications. (source: https://x.ai/news/grok-4-5)

04

Mistral's Robostral Navigate: a state of the art robotics navigation model

Mistral has introduced Robostral Navigate, a robotics navigation model designed to coordinate autonomous locomotion. The model utilizes visual-language representations to map physical environments and translate high-level language commands into low-level physical actuation. By integrating multi-modal sensing capabilities, the system processes spatial data dynamically to adapt to real-time physical obstacles and complex terrain challenges, highlighting Mistral's work in physical AI operations. (source: https://mistral.ai/news/robostral-navigate/)

05

GitLost: We Tricked GitHub's AI Agent into Leaking Private Repos

Cybersecurity researchers have demonstrated a significant security vulnerability in GitHub's integrated AI agent, successfully manipulating it to leak confidential data from private repositories. By using target-oriented prompt injection techniques and exploiting weak authorization checks in the agent's retrieval mechanisms, the researchers bypassed context access boundaries. The findings reveal a critical architectural flaw where natural language instructions can override access control lists, emphasizing the necessity of deterministic authorization sandboxes. (source: https://noma.security/blog/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos/)

06

Show HN: Microsoft releases Flint, a visualization language for AI agents

Microsoft has released Flint, a visualization intermediate language designed to help AI agents generate high-quality data representations reliably. Traditional visualization frameworks force a tradeoff between low-quality, simple specifications and complex layouts that trigger agent execution failures. Flint addresses this limitation by offering a semantic-type-based specification coupled with an automated layout optimization engine, allowing AI agents to generate accurate, aesthetically precise charts without managing low-level layout details. (source: https://microsoft.github.io/flint-chart/#/)

07

Show HN: Kastor – Terraform-style specs for AI agents

Kastor is a new open-source framework that introduces Infrastructure as Code concepts to the configuration and deployment of AI agents. Drawing inspiration from HashiCorp's Terraform, the framework enables developers to write declarative, version-controlled configuration files defining agent roles, system prompts, integrated tools, and multi-agent communication networks. This system aims to bring reproducibility and structured orchestration to complex, large language model-driven applications and cognitive agent workflows. (source: https://github.com/weirdGuy/kastor)

08

Show HN: Foreman, a self-hosted LLM gateway for cost aware model routing

Foreman is an open-source, self-hosted Large Language Model (LLM) gateway designed to minimize API expenses via cost-aware model routing. Acting as an intermediate routing layer, Foreman analyzes incoming API requests and dynamically redirects them to the most cost-effective models based on user-defined performance and budget parameters. The platform provides localized data control, real-time token tracking, and comprehensive cost analytics to help organizations manage rising production-grade API costs. (source: https://github.com/Northwood-Systems/foreman)

Twitter

5 stories
01

OpenAI Launches Next Generation Live Voice Capability For ChatGPT

OpenAI has officially launched its next-generation live voice capability for ChatGPT, aiming to provide low latency and natural prosody for voice-based interactions. The feature allows users to shift from traditional text-based prompts to a real-time conversational interface designed for tasks like brainstorming and casual information retrieval. Greg Brockman also promoted the technology under the name GPT-Live, indicating plans for future API access and integration with Codex to expand its utility across development environments. (source: https://x.com/sama/status/2074909079450050629)

02

Anthropic Enhances Claude Developer Tools With New Agentic Capabilities

Anthropic has introduced updates to the Claude developer ecosystem to support autonomous agentic workflows. The improvements include updated API capabilities and model interactions designed to simplify how engineers build complex AI agents. The release is designed to transition interactions from passive chatbot prompts to active, multi-step actions within enterprise development lifecycles, enabling Claude to assist directly with software engineering tasks such as coding, debugging, and managing system architectures. (source: https://x.com/ClaudeDevs/status/2074900291062034618)

03

ZML and LLMD Release High-Performance Heterogeneous LLM Inference Stack

The research team has released ZML and LLMD, a specialized Large Language Model server architecture featuring a custom, high-performance heterogeneous inference stack. Designed to optimize LLM deployment, this framework implements hardware-accelerated processing to improve throughput and reduce operational latency across diverse computational setups. This development targets the intensive infrastructure demands of serving modern large-scale generative models, offering developers a custom alternative to traditional, non-optimized serving backends. (source: https://x.com/ylecun/status/2074779707820617754)

04

MiniMax Hub Launches Backrooms Dreamcore Prompting Feature

MiniMax has introduced its new creative tool, Backrooms Dreamcore, inside the Skill Square section of the newly launched MiniMax Hub platform. Built on Hailuo AI technology, the feature enables users to generate atmospheric, dreamcore-themed digital spaces using single text prompts. The release is part of MiniMax's broader launch of the MiniMax Hub, which serves as a centralized gateway for users and developers to access, download, and interact with the company's generative models. (source: https://x.com/Hailuo_AI/status/2074880707164602798)

05

Kling AI Unveils Winners of the Nextgen AI Filmmaking Awards Ceremony

Kling AI has announced the official winners of its inaugural NEXTGEN Awards, a competition recognizing work at the intersection of professional filmmaking and generative artificial intelligence. The ceremony featured a variety of submissions, ranging from student-led projects to advanced cinematic 4K productions. The event highlighted how modern generative video tools are being adopted by creators to handle complex storytelling, convey personal themes, and experiment with advanced multimodal synthetic media production techniques. (source: https://x.com/Kling_ai/status/2074689941083185274)

huggingface

8 stories
01

Gemma 4 Technical Report

Google researchers have introduced Gemma 4, a new generation of open-weight, natively multimodal language models. Designed to advance compute efficiency and reasoning, the suite features dense and Mixture-of-Experts (MoE) architectures ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders, a unified, encoder-free 12B model is proposed to ingest raw audio and image patches directly. Additionally, Gemma 4 integrates a new thinking mode that generates reasoning traces prior to responding. The models achieve substantial performance leaps across STEM, multimodal, and long-context benchmarks, rivaling larger open frontier models on human-rated evaluations. (source: https://huggingface.co/papers/2607.02770)

02

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding

NVIDIA researchers have introduced Nemotron-Labs-Diffusion, a tri-mode language model that unifies autoregressive, diffusion, and self-speculation decoding within a single architecture. Trained with a joint autoregressive-diffusion objective, this model family scales up to 14B parameters and can switch modes to sustain high throughput. The diffusion objective improves lookahead planning, while the autoregressive component provides linguistic priors. In self-speculation mode, the model uses diffusion drafting and autoregressive verification, which decodes 6x more tokens per forward pass than Qwen3-8B and delivers 4x higher throughput on SPEED-Bench with SGLang on a GB200 GPU. (source: https://huggingface.co/papers/2607.05722)

03

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Researchers have introduced DSpark, a speculative decoding framework designed to improve Large Language Model (LLM) serving throughput under high concurrency. To prevent draft quality decay, DSpark utilizes a semi-autoregressive architecture that couples a parallel backbone with a lightweight sequential module, introducing intra-block dependency modeling. To maximize efficiency, it employs confidence-scheduled verification, dynamically adjusting verification length based on estimated prefix survival probabilities. When deployed within the DeepSeek-V4 production serving system under live user traffic, DSpark mitigated verification waste and accelerated per-user generation speeds by 60 to 85 percent compared to the established MTP-1 baseline. (source: https://huggingface.co/papers/2607.05147)

04

TREK: Distill to Explore, Reinforce to Refine

Researchers have proposed TREK (Teacher-Routed Exploration via Forward KL), a staged training procedure that uses distillation to expand exploration support for reinforcement learning. While Group Relative Policy Optimization (GRPO) struggles with hard prompts outside its on-policy support, TREK identifies low pass-rate prompts, obtains verified candidate solutions from a proposal source, pulls these modes into the student's support via a short forward-KL phase, and returns to on-policy GRPO. TREK with DeepSeek-V4 proposals improved Qwen3-8B math reasoning on AIME 2025 from 36.9 to 40.3 and AIME 2024 from 47.9 to 51.1, while also significantly boosting agent performance on ALFWorld and ScienceWorld. (source: https://huggingface.co/papers/2607.05339)

05

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

A systematic study has revealed that reinforcement learning (RL) adaptation gains are highly concentrated in a small subset of transformer layers during post-training. Across seven models from the Qwen3 and Qwen2.5 families, and utilizing three RL algorithms (GRPO, GiGPO, and Dr. GRPO), the researchers introduced a "layer contribution" metric to isolate individual layer gains. They discovered that training a single, middle-stack transformer layer can recover most of, and occasionally exceed, the performance gains of full-parameter RL training across mathematical reasoning, code generation, and agentic decision-making. (source: https://huggingface.co/papers/2607.01232)

06

SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review

Researchers have introduced SWE-Review, a closed-loop framework for software engineering agents that uses agentic code review to systematically diagnose and revise pull requests (PRs). Rather than open-loop, one-shot PR generation, SWE-Review implements a reviewer agent that explores the repository, decides whether to accept a PR, and provides structured feedback for iterative revision. Alongside the framework, the authors released the SWE-Review-Bench evaluation suite and the SWE-Review-Traj dataset. Experiments show that this generate-review-revise loop continuously improves PR resolve rates, outperforms single-turn fixed-context reviews, and enables effective test-time scaling. (source: https://huggingface.co/papers/2607.06065)

07

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

Researchers have developed Flex-Forcing, a unified video diffusion training and inference framework that seamlessly bridges bidirectional and autoregressive generation. While traditional bidirectional models excel at global coherence but suffer from high latency, autoregressive models enable fast streaming output but are prone to exposure bias. Flex-Forcing addresses this with a flexible chunking mechanism defined over temporal axes and denoising steps. It allows bidirectional planning across chunks and autoregressive synthesis within chunks. Across multiple video generation benchmarks, the framework delivers improved video quality, long-video stability, and faster inference speeds. (source: https://huggingface.co/papers/2607.03509)

08

SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

Researchers have introduced SkillOpt-Lite, a minimal, theoretically grounded pipeline for autonomous agent skill optimization formulated via Zeroth-Order (ZO) optimization. Guided by PAC learning, the framework establishes three core principles: file-system-based trajectory exploration, consensus attribute mining, and validation gating. SkillOpt-Lite improves LiveMath performance by +8.8 points on GPT-5.5 and +25.4 points on GPT-5.4-nano, enabling the nano model to outperform standard GPT-5.4. Extended to full harness optimization (HarnessOpt) on SpreadsheetBench, it enabled GPT-5.4-nano to achieve 0.7758 accuracy, surpassing the larger GPT-5.5 running standard pipelines. (source: https://huggingface.co/papers/2607.03451)