NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-09-25ENGLISH EDITION
This issue
—
All time
—

AI Blog

3 stories
01

Proaction reports sales and time gains from OpenAI tools

OpenAI published a Proaction customer story reporting a 60 percent sales increase and more than 75 hours saved with Codex. The accompanying description says Proaction uses Codex, GPT-Live-1, and GPT-6 Astra to build, operate, and sell fleet-management products faster. This is a concrete business-workflow example linking coding and broader operations. The collected summary does not explain the measurement period or isolate each tool’s contribution, so the figures remain vendor-reported case-study outcomes.

02

SemiAnalysis maps China’s AI datacenter expansion

SemiAnalysis maps China’s AI datacenter expansion

SemiAnalysis introduced a China datacenter model covering more than 1,000 facilities across over 60 operators. Its summary describes infrastructure originally built for retail demand and increasingly redirected toward AI, alongside rapid construction and concentrated hyperscaler leasing. For model builders and infrastructure planners, the report offers a supply-side perspective on the compute market. The collected excerpt provides headline scope and claims, but not the underlying facility dataset or methodology needed to verify individual estimates.

03

invideo reports faster grading and custom effects with GPT-6 Astra

In a September 23 customer story, OpenAI describes invideo using GPT-6 Astra for more precise edit planning, color correction, and grading. The summary reports a threefold improvement in correction and grading and production of 50 custom effects in one day. This connects language-model planning with creative production tasks beyond text generation. The available excerpt does not specify the evaluation method or baseline, so these figures should be interpreted as reported customer outcomes.

Hacker News

1 story
01

Recurse presents a serverless harness for specialist agents

Recurse presents a serverless harness for specialist agents

Recurse introduced a coding-agent skill and serverless runtime aimed at specialist agents with clear input-output contracts. Developers provide requirements and evaluation criteria, while the coding agent tries prompt and Python-tool variants. A finite-state harness supports iterative refinement in verifiable domains. The team describes the system as operational, with documentation available but some examples still incomplete. Its central proposition is faster agent experimentation, including use of smaller models where evaluations can guide improvement.

Twitter

9 stories
01

Google highlights Gemini 3.8 speech models

Google highlights Gemini 3.8 speech models

Google highlighted Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS in a report-date recap of its audio releases. The company positions the pair around creative production and cost-efficient speech at scale. The announcement gives builders two named options to investigate for speech workflows, although this collected post does not provide pricing, latency measurements, or a comparative quality evaluation. It is a product update rather than independent benchmark evidence.

02

Claude adds a plugin submission and usage portal

Claude adds a plugin submission and usage portal

Claude developers announced a portal for submitting plugins, tracking review, and viewing usage. Plugins package MCP integrations and skills, connecting distribution with the components developers already use to extend Claude. The announcement also reports that MCP usage across Claude products has grown 110 times this year. That figure is a company-reported adoption claim; the concrete developer change is a more visible submission, review, and usage workflow for plugins.

03

Claude Code adds a graceful stop at usage limits

Claude Code adds a graceful stop at usage limits

Claude Code will try to reach a graceful stopping point when a user hits the five-hour usage limit during a task. Instead of cutting off immediately during an edit, it receives a small fixed allowance drawn from the weekly limit to finish what it can. This changes interruption behavior for ongoing coding work, but the announcement does not promise task completion or specify the size of the additional allowance.

04

Claude adds a calculator for Opus 5.5 task costs

Claude adds a calculator for Opus 5.5 task costs

Claude developers say Opus 5.5 costs 20 percent less per input and output token than Opus 5, with cache reads priced 60 percent lower. They also introduced a calculator accessible from the usage interface to help users examine task costs. The useful distinction is between token prices and actual workflow spending: the announcement supplies relative pricing claims and a planning tool, without establishing identical savings for every coding task.

05

Runway connects image and video generation to Claude through MCP

Runway connects image and video generation to Claude through MCP

Runway announced an MCP connection that lets users access image and video generation directly from Claude, highlighting Claude Opus 5.5. The post names Gen-4.5, Seedance 2.5, GPT Image 2, and Kling among the available models. This brings a collection of creative tools into an assistant workflow and reduces the need to move between interfaces. The announcement does not provide comparative output quality, pricing, or detailed permission behavior.

06

Runway introduces editable image layers

Runway introduces editable image layers

Runway introduced Layers, a feature that splits an image into editable components with a single click. The company describes background removal, text editing, and independent changes to individual elements within the same workspace. For creative teams, the stated benefit is a more direct path from a generated image to targeted revisions. The collected announcement does not quantify segmentation accuracy or show how consistently complex images separate into usable layers.

07

Pika connects Grok Bots to a multimodal generation API

Pika connects Grok Bots to a multimodal generation API

In a September 24 announcement, Pika described connecting Grok Bots to its API for image, video, and audio generation. The company says the integration provides access to more than 120 models through one API key, naming Seedance 2.5, Wan 3.0, and GPT-image-2. This is a practical example of assistants gaining creative production tools. The collected post does not include model-by-model availability, generation costs, or performance comparisons.

08

Claude resumes billing for selected blocked requests

Claude resumes billing for selected blocked requests

Claude developers announced on September 24 that billing would resume for requests blocked by safeguards before a response. The post limits the policy to categories described as having low false-positive rates, including biology, distillation attacks, and frontier language-model development. It also mentions recent coordinated attacks. For API users, the operational implication is that some rejected requests can still incur charges; the collected excerpt does not specify the complete billing calculation.

09

Runway adds a lower-cost Seedance 2.5 draft mode

Runway adds a lower-cost Seedance 2.5 draft mode

Runway announced a Draft mode for Seedance 2.5 on September 24. It describes faster generation with fewer credits during exploration, followed by an option to enhance preferred results to full quality. The feature separates early creative iteration from final rendering, potentially making it easier to compare ideas before committing more resources. The announcement does not specify exact credit savings, speed improvements, or how closely enhanced outputs preserve each draft.

GitHub

2 stories
01

Hindsight highlights learning memory for agents

Hindsight highlights learning memory for agents

Hindsight appeared in the collected GitHub Trending snapshot as a Python project described as agent memory that learns. Its latest recorded commit falls on the report date, and the snapshot records 29,789 stars. Those signals support including it as a project to inspect, rather than treating the snapshot as evidence of a new release. The available repository description does not establish its memory architecture, evaluation results, or production reliability.

02

Superpowers presents an agentic skills methodology

Superpowers presents an agentic skills methodology

Superpowers appeared in the collected GitHub Trending snapshot as a skills framework and software-development methodology for agents. The snapshot identifies Shell as its language, records 291,662 stars, and places its latest commit on the report date. This makes it relevant to the expanding ecosystem around reusable agent skills. The collected metadata does not describe individual skills or provide controlled productivity measurements, so the entry should be read as project discovery rather than a validated performance claim.

HF HuggingFace

12 stories
01

Transformers show evidence of linear superposition

Transformers show evidence of linear superposition

This paper studies whether a Transformer can process linearly combined text streams and produce a corresponding mixture of next-token distributions. The authors find evidence that the behavior is intrinsic to the architecture and diminishes during pretraining, while lightweight fine-tuning can restore it. They also introduce guided decoding to separate two coherent continuations from one forward pass. The result suggests an unusual inference direction, though the abstract does not establish broad deployment-ready efficiency gains.

02

World Action Agent rehearses robot actions before execution

World Action Agent rehearses robot actions before execution

World Action Agent gives vision-language models a visual workspace for robot manipulation, combining contact views, action rehearsal, and corrections tied to observed execution. Agents can preview and revise proposed actions, while reusable skills come from expert videos and human teaching. The authors report 75.6 percent average success on LIBERO-Pro with transferred skills. Training a smaller model on interaction traces also improves out-of-domain success, suggesting the harness can support both execution and data collection.

03

Rufus-Air publishes an eight-stage post-training recipe

Rufus-Air publishes an eight-stage post-training recipe

Rufus-Air presents an open post-training recipe for GLM-4.5-Air-Base, moving through eight stages from supervised fine-tuning to reasoning, coding, instruction following, agent capabilities, and RLHF. The paper documents data, rewards, infrastructure, ordering, and intermediate results. It emphasizes reliable rewards and appropriate difficulty as practical training principles, using public data and open components without new human annotation or an internal distillation teacher. The authors report improvements over the official post-trained base-family release.

04

IterSynth separates search planning from evidence synthesis

IterSynth separates search planning from evidence synthesis

IterSynth restructures deep-search agents around a Planner and a Synthesizer, using an evolving summary instead of an ever-growing search transcript as persistent state. Its training method assigns role-specific reinforcement-learning credit through final outcomes and turn-level rubric evaluations. Across five benchmarks, the authors report an average score of 50.7 for IterSynth-8B, exceeding the strongest prior agent in its size class. The design also provides a prompting approach for larger proprietary models.

05

Coding agents synthesize reusable robot planners

Coding agents synthesize reusable robot planners

This study asks coding agents to build programs for generalized task and motion planning using task descriptions and simulator access. Programs are frozen before testing on unseen instances, separating development from evaluation. Across 28 simulated environments and 98,000 evaluation episodes, the authors find strong performance from the evaluated Claude Code and Codex configurations. On environments with hand-engineered planners, reported mean success ranges from 56 to 95 percent, compared with 47 percent for those planners.

06

Qwen-Planner-Agent closes the mobile-agent development loop

Qwen-Planner-Agent closes the mobile-agent development loop

Qwen-Planner-Agent connects data production, training, and deployment through a shared action-feedback-verification contract. The framework includes human-gated data collection, supervised initialization, online agentic reinforcement learning, and coordinated adaptation of models and runtime tools. Its CARE method targets reasoning and tool-use costs while preserving task performance. The authors report the strongest overall results among the evaluated systems on MobilePA-Bench, with additional gains on non-mobile agent benchmarks and largely preserved general capabilities.

07

Jev evaluates alignment failures with calibrated decisions

Jev evaluates alignment failures with calibrated decisions

The paper evaluates Jev, a model trained with reinforcement learning for calibrated decisions, as a detector of ten kinds of alignment failure. Its benchmark spans 44 datasets and five target models, including prompt injection, deception, and reward hacking. A generic question reportedly achieves median AUROC of 0.886 without task-specific training. The authors find that supplied context matters more than question wording and report much lower scoring costs than generative judges in their evaluation.

08

AV-GRPO decouples reinforcement learning for audio and video

AV-GRPO decouples reinforcement learning for audio and video

AV-GRPO addresses joint audio-video generation by separating learning signals that otherwise make reward attribution difficult. Its modality-anchored rollouts, frozen-tower optimization, and adaptive objectives turn coupled preference learning into more focused subproblems. The accompanying 5DAV dataset controls training difficulty across five dimensions. On JavisBench and VABench, the authors report improvements over LTX-2.3 in generation quality, semantic alignment, and synchronization under both LoRA and full fine-tuning, with ablations supporting the proposed components.

09

ExplorationBench tests discovery in unfamiliar executable worlds

ExplorationBench tests discovery in unfamiliar executable worlds

ExplorationBench evaluates whether AI systems can discover unfamiliar rules rather than recall knowledge from training. It uses executable alien worlds with misleading manuals, feedback, and tool interfaces, making answers verifiable while deliberately conflicting with familiar assumptions. Two sandboxes each contain 70 tasks. Across ten evaluated systems, stronger models can learn and apply new rules, but longer exploration sometimes stalls or reverses gains, exposing a distinction between additional interaction and productive discovery.

10

WanPE turns video prompts into cinematic shot plans

WanPE turns video prompts into cinematic shot plans

WanPE is a 397-billion-parameter prompt-enhancement model trained on 1.05 million videos to plan multi-shot video generation. It combines video-grounded reverse construction with reinforcement learning designed to preserve user intent across shots and time. The authors introduce WanPEval with approximately 11,000 blind pairwise assessments. When used with Wan3.0, they report improved human preference over raw prompts, including a particularly large gain for 30-second sequences, highlighting text planning as a video-quality factor.

11

ViRDM simplifies few-step causal video post-training

ViRDM simplifies few-step causal video post-training

ViRDM replaces the teacher-and-critic setup commonly used for video diffusion distillation with generator-only post-training against a precomputed representation distribution. Its recipe addresses memory constraints and temporal dynamics through truncated supervision, a lightweight decoder, staged derivatives, and dynamics regularization. The authors report an official VBench score of 84.87 after 20 generator updates using 16 A100 GPU-hours. These results suggest a cheaper post-training route, while remaining specific to the evaluated generation settings.

12

Agent-Editing World Model revises contaminated task state

Agent-Editing World Model revises contaminated task state

Agent-Editing World Model targets unsupported assumptions and stale plans in agent histories instead of trying to predict complex tool responses. An Action Judge classifies decisions, while State Revision edits noisy reasoning and action continuations. The EditAct method combines those capabilities with real execution feedback. Across six benchmarks and three agent backbones, the authors report average gains of 3.2 to 6.7 points over the strongest baseline, suggesting that repairing internal task state can improve subsequent decisions.