NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-08-04DEFAULT EDITION
This issue
—
All time
—

AI Blog

3 stories
01

Continuous Voice Interaction With GPT Live

OpenAI has announced GPT-Live, a continuous voice interaction system designed to enable natural, turnless conversations with AI. Developed over a six-month period, the system utilizes a turnless speech model architecture that eliminates the artificial turn-taking structure characteristic of traditional voice assistants. This design dramatically reduces conversational latency and supports more fluid, real-time communication. The research and development effort focused on optimizing speech modeling and system infrastructure to make voice interaction faster and more intuitive for everyday applications. (source: https://openai.com/index/continuous-voice-interaction-with-gpt-live)

02

New Education Plugins for ChatGPT Work and Codex

OpenAI has introduced new education-focused plugins for ChatGPT Work and Codex to assist K-12 teachers, college educators, and students in academic environments. These targeted integrations are designed to help users learn, teach, conduct research, and build projects within educational settings. By leveraging these specialized tools, educators can enhance classroom instruction while students gain access to tailored resources for academic study and development. (source: https://openai.com/index/learn-teach-chatgpt-work-codex)

03

Circles Integrates OpenAI Technology To Personalize Telecommunications Experience

Telecommunications company Circles has integrated the OpenAI API and Codex to power its new AI-native telecommunications experiences. The implementation of these generative models has led to a 22% increase in average revenue per user (ARPU) and a 9% reduction in subscriber churn. Additionally, incorporating Codex into engineering workflows has significantly improved software development efficiency. This collaboration illustrates the practical deployment and measurable business impact of generative AI technologies in the telecommunications sector. (source: https://openai.com/index/circles)

Hacker News

8 stories
01

Mistral's Shieldstral: 3B open-weights model for multimodal moderation

Mistral AI has released Shieldstral, a new 3B open-weights model engineered specifically for multimodal moderation tasks. The compact safety classifier is optimized for high-throughput, low-latency deployment, helping developers identify toxic, unsafe, or harmful content across both text and image modalities. Shieldstral aims to provide cost-effective, real-time guardrails for AI agents and collaborative applications. Benchmarks demonstrate its competitive performance against larger proprietary moderation solutions, offering the open-source community localized control over content pipelines. (source: https://mistral.ai/news/shieldstral/)

02

DeepSeek V4 Flash on a Single AMD MI300X

Developer Ryan Zhou has released a technical project demonstrating the performance optimization of DeepSeek V4 Flash on a single AMD Instinct MI300X accelerator. The repository provides specific deployment scripts, environment setup documentation, and configuration guidelines to maximize inference speed. By leveraging AMD's ROCm software platform, the project serves as a practical blueprint for developers looking to deploy large language models on non-NVIDIA enterprise hardware without suffering performance degradation. (source: https://github.com/ryanzhou/deepseek-v4-flash-mi300x)

03

Apple says more ex-employees may have taken confidential data to OpenAI

Apple has raised corporate security concerns, alleging that additional former employees may have transferred confidential proprietary data to OpenAI upon transitioning to the startup. This dispute highlights the rising competitive tension between established technology firms and generative AI companies over highly specialized talent and proprietary datasets. The allegations underscore the challenges of safeguarding trade secrets during industry transitions and could prompt stricter enforcement of non-disclosure agreements across the AI landscape. (source: https://techcrunch.com/2026/08/04/apple-says-more-ex-employees-may-have-taken-confidential-data-to-openai/)

04

The Warp Agent CLI

Warp has introduced the Warp Agent CLI, a terminal-based autonomous coding assistant designed to execute complex development pipelines directly in the command-line interface. The tool acts as an AI agent that integrates with the terminal environment, analyzing directory structures, generating context-aware code, and running local shell commands to debug system errors. This release addresses the demand for low-friction development assistance by bringing LLM capabilities directly into the terminal. (source: https://www.warp.dev/blog/introducing-the-warp-agent-cli-coding-agent)

05

Show HN: Fine-tune an 8B model on a 4 GB laptop GPU

Developer Makazhan Alpamys has launched Soup, an open-source framework designed to fine-tune 8B parameter large language models on consumer-grade hardware with only 4 GB of VRAM. By implementing memory optimization techniques like parameter offloading, gradient checkpointing, and quantization, the project lowers the barrier for running machine learning workloads locally. This tool enables individual developers to customize large language models on personal devices rather than relying on cloud instances. (source: https://github.com/MakazhanAlpamys/Soup)

06

Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research

Founders Rui and Michael have launched EdotEnv, a YC S26 company offering self-improving reinforcement learning environments based on quantitative trading workflows to evaluate large language models. The platform addresses benchmark saturation by using dynamic, self-correcting financial markets that naturally scale in difficulty. LLMs are assigned quantitative tasks and assessed on out-of-sample data, creating a robust framework designed to teach advanced research capabilities to AI systems. (source: https://edotenv.com/)

07

Why Large Language Models Fail at Tabular Prediction

A research paper has analyzed the performance bottlenecks of Large Language Models (LLMs) when applied to tabular prediction tasks. The study outlines why LLMs struggle compared to traditional tree-based algorithms, pointing to challenges in capturing complex numerical relationships, a lack of spatial awareness in serialized data formats, and standard tokenization methods that disrupt continuous numerical features. The findings offer architectural insights for researchers adapting foundation models to heterogeneous databases. (source: https://arxiv.org/abs/2608.02412)

08

Homebench – Benchmark local LLMs for speed, memory, and quality

Developer david-g-3654 has introduced Homebench, an open-source tool designed to benchmark local large language models. The framework measures performance across token generation speed, VRAM/RAM consumption, and response accuracy in a standardized local testing environment. The tool helps developers and enthusiasts identify the optimal balance between hardware constraints and quantization levels, simplifying the configuration and selection process for deploying models locally. (source: https://github.com/david-g-3654/homebench)

Twitter

6 stories
01

Sakana AI Deploys Namazu Model With Modal Infrastructure Support

Sakana AI has launched Namazu, a 1-trillion parameter large language model engineered for real-time web search and autonomous code execution. Tailored with Japanese cultural nuance and technical reasoning, the model is deployed on Modal's cloud computing platform to handle complex workflows and scale-out serverless workloads. This release highlights the shift towards domain-specific, agentic applications utilizing localized context. The model's infrastructure is optimized for high responsiveness and efficiency during active tool use and coding tasks. (source: https://x.com/hardmaru/status/2084448878657773810)

02

MiniMax H3 Video Generation Model Now Available With Low VRAM Support

MiniMax has released an optimization update for its H3 video generation model, allowing the generative AI tool to run locally on hardware with as little as 5GB of VRAM. Developed with contributions from @deepbeepmeep, the update lowers the barrier for high-quality video synthesis on consumer-grade graphics cards. The H3 model is also capable of producing 2K resolution videos from simple text prompts and has been integrated with tools like Midjourney V8.2 and platforms like Higgsfield AI for advanced cinematic, stylized, and rapid video content creation. (source: https://x.com/Hailuo_AI/status/2084538571185332345)

03

Adaption AI Launches AutoScientist API For Automated Model Training

Adaption AI has officially introduced the AutoScientist API, a new platform designed to automate the configuration and training of task-specific machine learning models. Built to simplify the model-building pipeline with minimal code, the technology demonstrates a 35% efficiency improvement over human-configured setups according to early performance benchmarks. Sarah Hooker also announced a strategic research partnership with AI Singapore to leverage these adaptive data techniques and autoscientist frameworks to accelerate automated scientific discovery. (source: https://x.com/sarahookr/status/2084511488061108626)

04

Pika Labs Introduces Monthly Membership Model For Cost-Effective API Access

Pika Labs has transitioned to a subscription-based pricing model offering API access to generative models for $10 per month. Subscribing users receive steep cost reductions on high-performance media tools, including up to 87% off Seedance 2.0, 50% off MiniMax H3, and 25% off GPT Image 2. This structure is designed to offer a competitive, affordable alternative to platforms like Fal and Runway for developer teams and power users requiring persistent, high-volume video and image generation. (source: https://x.com/pika_labs/status/2084711615249928266)

05

Unreal Fest Talk Explores MetaHuman Future and Markerless Motion Capture

Michael J. Black announced his Unreal Fest presentation detailing the five-year roadmap of the MetaHuman framework and the launch of the Markerless Motion Capture plugin for Unreal Engine 5.8. The integration provides developers with an efficient pipeline to generate high-fidelity character animations directly in virtual environments, eliminating traditional hardware constraints. The presentation outlines the technical roadmap to improve digital character generation and animation workflows. (source: https://x.com/Michael_J_Black/status/2084616472928629041)

06

Kling AI Highlights Enhanced Detail Precision in Video Generation Capabilities

Kling AI has deployed a technical update targeted at refining fine-grained detail and texture precision within its generative video platform. The release focuses on reducing resolution artifacts and improving micro-visual realism during complex scene reconstruction. By enhancing microscopic elements and textural synthesis, the update aims to provide content creators with higher-fidelity automated visual outputs for demanding media production environments. (source: https://x.com/Kling_ai/status/2084610301606170764)

huggingface

8 stories
01

DiffusionGemma Technical Report

Google researchers have introduced DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text. By iteratively refining blocks of 256 tokens in parallel, the model avoids the sequential decoding bottleneck of conventional autoregressive models. Fine-tuned from the mixture-of-experts Gemma 4 model (3.8B active, 25.2B total parameters) using a two-stage training pipeline, DiffusionGemma generates approximately 20 tokens per forward pass. It achieves about 1,500 output tokens per second on a single NVIDIA H100 GPU while retaining thinking mode, multimodal input support, and long context capabilities. (source: https://huggingface.co/papers/2608.00146)

02

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks

Researchers have proposed LongHorizon-Harness to improve long-horizon task execution for large language model agents by modeling execution as an explicit task-state management problem. The framework utilizes a Manage-Execute-Audit (MEA) loop to maintain task states outside execution, updating them using independently verified environmental facts. Using the lightweight AgentAdapter, the harness supports interchangeable backend models. In evaluations, LongHorizon-Harness improved Qwen 3.7-Plus from 51.8% to 80.7% on WeaveBench, from 69.7% to 77.2% on Terminal-Bench 2.1, and from 2.8% to 8.3% on OSWorld 2.0. Claude Opus 4.7 performance on an OSWorld 2.0 subset also increased from 20.0% to 34.3%. (source: https://huggingface.co/papers/2608.01964)

03

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation

To address the difficulty language model agents face in coordinating reusable procedural tools, researchers developed SKT, a verified data synthesis pipeline. SKT generates skill-grounded tasks and executable trajectories from large-scale collections. Utilizing 2,000 public skills, the pipeline constructed 4,000 task packages and 27,164 verified trajectories. Based on these configurations, the authors also built SkillEval, a held-out evaluation benchmark. Supervised fine-tuning on the generated trajectories consistently improved skill-use performance across multiple model backbones and agent harnesses, demonstrating the scalability of verified synthetic data generation. (source: https://huggingface.co/papers/2608.02287)

04

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

To evaluate how software engineering agents handle concurrent modifications in collaborative workspaces, researchers introduced SWE-Touch. The benchmark stress-tests agent adaptation using Counter-Edits, which are plausible edits made by a separate user patch generator that conflict with task completion. Across nine coding models evaluated on SWE-bench Verified, the inclusion of Counter-Edits decreased the average resolve rate by 7.7 percentage points. Trajectory analysis indicates that failures are linked to limited workspace awareness, as agents frequently failed to re-inspect codebases, validate edits, or detect conflicting external changes. (source: https://huggingface.co/papers/2608.02499)

05

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

A newly published study examines "deletion avoidance," a systematic bias where large language models avoid removing obsolete code during edits. Across five leading models on the SWE-bench Verified leaderboard, the highest deletion recall against developer patches was only 71.7%. Retrofitting 34 tasks with tests validating code removal caused the success rate of four frontier models to fall from 63.2% to 41.9%. To isolate this behavior, the authors curated CanItDelete, a benchmark of 200 pure deletion tasks, where smaller open-weight models dropped to an 18% success rate. (source: https://huggingface.co/papers/2607.28887)

06

Zero-Mem: Zero-Token Memory Operations for LLM Agents

To reduce the high token and latency costs of agent memory systems, researchers proposed Zero-Mem, a framework that executes memory operations without additional large language model generation steps. Zero-Mem stores interaction traces within an entity-context graph and a temporal hierarchy. For each query, the system retrieves relevant views from both structures and applies deterministic calibration to keep answers grounded, calling the LLM only for final question answering. Evaluated on long-memory and long-context benchmarks, Zero-Mem eliminated intermediate LLM calls and decreased memory-operation time by 57.6% compared to the fastest baseline. (source: https://huggingface.co/papers/2607.29377)

07

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step

To evaluate how well autonomous agents learn tool use purely through environmental feedback, researchers introduced ScrambleToolBench, a dynamic terminal benchmark. By omitting semantic tool descriptions and introducing obstacles like mapping drift and stochastic execution failures, the environment forces agents to reason via trial-and-error discovery. Evaluated models frequently failed to adapt to sudden changes like mapping drift, exhibiting belief inertia or resorting to expensive, brute-force searches instead of leveraging deductive strategies like cycle tracing. Increasing test-time reasoning computation only exacerbated exhaustive searching without resolving the underlying adaptation errors. (source: https://huggingface.co/papers/2608.02358)

08

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

Researchers have developed GradCuit (gradient through circuit), a test-time latent reasoning framework that optimizes continuous states at inference while keeping model parameters frozen. Unlike token-level optimization, GradCuit embeds optimizable latent states within Transformer layers, using causal self-attention to map gradients from the entire continuation directly back to the latents. Across five instruction-tuned backbones and three reasoning benchmarks, GradCuit achieved an average accuracy of 64.5%, outperforming chain-of-thought prompting by 6.6 percentage points. The method also demonstrates high robustness across different learning rates compared to previous test-time scaling baselines. (source: https://huggingface.co/papers/2608.02585)