NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-17ENGLISH EDITION
This issue
—
All time
—

AI Blog

1 story
01

CFO Introduces Scorecard to Measure AI Return on Investment

OpenAI Chief Financial Officer Sarah Friar has introduced a standardized scorecard framework to measure the return on investment for enterprise artificial intelligence implementations. The methodology addresses key economic and operational aspects of Large Language Model deployments by tracking metrics across four core pillars: useful work generated, the cost incurred per successful task, overall system dependability, and the return on compute resources. This framework is designed to help organizations transition from experimental AI pilots to scalable, value-driven operations by providing concrete metrics to justify technology expenditures and optimize resource allocation. (source: https://openai.com/index/a-scorecard-for-the-ai-age)

Hacker News

4 stories
01

Apple targets dozens of OpenAI employees with legal letters

Apple has issued formal legal warning letters to dozens of its former employees who have transitioned to OpenAI. The letters demand strict adherence to intellectual property agreements and non-disclosure obligations to protect Apple's proprietary trade secrets. This legal action highlights the intensifying competition for top-tier research and engineering talent between established consumer technology giants and rapidly growing generative artificial intelligence laboratories, which frequently hire key staff from competitors. (source: https://www.ft.com/content/1b8c9d52-88a9-426b-ba47-f1811f859166)

02

The state of open source AI

The State of Open Source AI report has released a comprehensive global analysis of the decentralized artificial intelligence ecosystem. The publication details the growth of open-access foundation models, open-weights architectures, and new fine-tuning methodologies. It outlines critical operational challenges for the open-source community, specifically around data licensing models, compute resource accessibility, and increasing global regulatory pressures. The report finds that open-weights alternatives are rapidly narrowing the capabilities gap with proprietary models. (source: https://stateofopensource.ai/)

03

Kimi K3, and what we can still learn from the pelican benchmark

Moonshot AI has released its new Kimi K3 large language model, which is evaluated against the pelican benchmark. The analysis demonstrates how Kimi K3 handles complex, long-context reasoning tasks and advanced information retrieval. The author notes that traditional academic tests often fail to measure these capabilities, advocating instead for dynamic, real-world benchmarks to accurately evaluate reasoning and instruction-following performance in next-generation systems. (source: https://simonwillison.net/2026/Jul/16/kimi-k3/)

04

VulnHunter: Capital One's agentic AI code security tool

Capital One has announced the open-source release of VulnHunter, an agentic artificial intelligence security tool designed to identify and remediate software vulnerabilities. VulnHunter utilizes advanced AI agent architectures and large language models to simulate developer workflows and attacker behaviors. This approach enables the tool to autonomously analyze codebases, trace execution paths, and detect complex, multi-step vulnerability chains, offering a significant reduction in false positives compared to legacy tools. (source: https://www.capitalone.com/tech/open-source/announcing-vulnhunter/)

Twitter

6 stories
01

New Course Launches on Building Fast LLM Applications With Cerebras Hardware

Andrew Ng has announced a new educational course developed in collaboration with Cerebras, focusing on building high-speed LLM applications. The curriculum addresses the memory-to-compute bottleneck and teaches developers to optimize inference by utilizing Cerebras' Wafer-Scale Engine hardware rather than standard GPUs. Students gain practical experience creating latency-sensitive tools, including live translation modules, real-time voice agents, and multi-step agentic workflows. (source: https://x.com/AndrewYNg/status/2078144569594761591)

02

Adaption AI Launches The AutoScientist Frontier Model Building Challenge

Adaption AI has officially launched the second part of its AutoScientist Challenge, a competitive event designed to incentivize frontier model building and automated scientific discovery. The initiative invites collaborative teams of researchers and developers to construct and optimize their own artificial intelligence models. Participants will compete for a total prize pool of $60,000, showcasing engineering workflows on the platform to push the boundaries of model adaptation. (source: https://x.com/sarahookr/status/2078185696905593240)

03

New Research Bridges Neuroscience Principles with Neural Network Training

Researchers have released a paper titled 'Diffusing Blame' that reconciles biological neural network dynamics with artificial deep learning architectures. The study introduces a novel routing methodology to train networks in accordance with Dale's principle, which dictates that biological neurons function as strictly excitatory or inhibitory. This framework aims to bridge computational neuroscience constraints with modern backpropagation-based optimization techniques. (source: https://x.com/hardmaru/status/2078156625479921847)

04

Inkling Model Shows High Potential in ARC AGI Benchmarks

The Inkling model is receiving increased industry attention due to its competitive performance on the ARC AGI benchmark, which evaluates abstract reasoning capabilities. Analysts report that a forthcoming, compact version of this architecture is planned for release, aiming to deliver high problem-solving efficiency within a smaller parameter layout. This development allows the research community to evaluate Inkling's specialized technical achievements against current generalized intelligence peers. (source: https://x.com/natolambert/status/2078177159856947430)

05

Google DeepMind Announces Significant Updates To Weather Lab AI Platform

Google DeepMind has launched a major update to Weather Lab, its digital platform designed to provide interactive public access to AI-powered climate and weather models. The update showcases deep learning applications solving complex geophysical and atmospheric forecasting tasks. This initiative aims to foster collaborative climate research and improve meteorological forecasting accuracy by validating neural-network-driven predictive structures. (source: https://x.com/GoogleDeepMind/status/2078150319016382762)

06

Kling AI Honors Next-Generation Creators at Seoul Awards Ceremony

Kling AI hosted its NextGen Awards Ceremony in Seoul to recognize creators leveraging its generative video platform for visual storytelling. The event highlighted creative projects such as 'The Well' by Jo Ryeongmi, which won the Best Storytelling award at the NEXTGEN Korea University Creative Challenge. These projects showcase the application of advanced generative architectures to execute complex creative workflows and synthetic video production. (source: https://x.com/Kling_ai/status/2078132620873929184)

huggingface

8 stories
01

RoboTTT: Context Scaling for Robot Policies

NVIDIA researchers introduced Test-Time-Training Robot Policies (RoboTTT), a training recipe and sequence model that scales robot visuomotor context to 8,000 timesteps. This architecture-aware method integrates Test-Time Training into Vision-Language-Action policies, using fast weights updated by gradient descent during both training and inference to bypass increased inference latency. In real-robot manipulation tasks, RoboTTT improves overall performance by 87% over a single-step context baseline, successfully completing a 5-minute, 10-stage assembly task. Scaling pretraining context from 1K to 8K timesteps yielded a 62% performance gain, suggesting context length is a viable scaling axis for robot foundation models. (source: https://huggingface.co/papers/2607.15275)

02

Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel

Researchers proposed own-position key-value (KV) cache grafting, a technique that improves the performance of frozen small language models while drastically reducing generation costs. By depositing verified knowledge as a byte-exact KV state artifact and restoring it into a fresh context, a frozen Gemma-4-12B model improved its AIME 2025 score from 80.0% to 93.3%. On recurring problems, the method reduced tokens processed by a factor of 6,574, leading to a corresponding 8,700-fold decrease in energy consumption. The approach also expanded the usable context window from 32,768 to 2,854,766 tokens without requiring extra accelerator memory. (source: https://huggingface.co/papers/2607.14431)

03

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

Researchers developed LongStraw, an architecture-aware execution stack designed to enable reinforcement learning post-training beyond 2 million tokens under a tight hardware budget. Implemented using Group Relative Policy Optimization (GRPO), LongStraw evaluates the shared prompt without autograd, saves only model-specific states, and replays short response branches sequentially to limit memory overhead. Tested on Qwen3.6-27B using eight H20 GPUs, LongStraw completed grouped scoring and response backwards at 2.1 million tokens. In separate stress tests, the system reached 4.46 million positions, demonstrating an execution path for a 2.1-million-token prompt across all 78 layers of GLM-5.2 on 32 H20 GPUs. (source: https://huggingface.co/papers/2607.14952)

04

SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning

Researchers proposed SEED (Self-Evolving On-Policy Distillation), an agentic reinforcement learning framework that converts completed trajectories into natural-language skills to guide intermediate token-level decisions. The policy is first trained to analyze its own outputs, generate hindsight skills, and re-score actions under ordinary and skill-augmented contexts. This converts skill-induced probability shifts into a dense, token-level distillation signal optimized alongside trajectory-level rewards. Evaluation across text- and vision-based tasks demonstrates that SEED consistently improves final performance and sample efficiency while fostering generalization to unseen environments. (source: https://huggingface.co/papers/2607.14777)

05

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

A systematic study investigated the training dynamics of on-policy distillation (OPD) in large language model post-training, identifying key pathologies. The research shows that OPD acts as an exploration catalyst but is highly sensitive to student-teacher distributional mismatches and length exploitation, where students truncate or pad responses to game token-level objectives. To regulate these issues, the authors implemented lightweight advantage clipping and log-scale compression. Experiments across seven benchmarks validated that these signal regulations eliminate length-dependent shortcuts and outperform standard OPD and reinforcement learning from visual feedback (RLVR) baselines. (source: https://huggingface.co/papers/2607.13399)

06

Hierarchical Denoising For Multi-Step Visual Reasoning

Researchers proposed Hierarchical Denoising for Visual Reasoning (HDR), a framework that organizes video latents into a tree hierarchy to perform coarse-to-fine reasoning in causal video generation. By utilizing coarse denoising layers to preserve planning hypotheses and finer layers to progressively resolve visual details, HDR maintains logical consistency. Testing on a multi-step visual reasoning benchmark covering six tasks—such as maze navigation and Sokoban—shows that HDR improves task success rates from 34.22% to 60.29%. HDR operates at 0.70 seconds per latent, delivering streaming outputs 54.2 times faster than bidirectional diffusion baselines. (source: https://hierarchical-diffusion-reasoning.github.io/)

07

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

To address repetitive search behaviors and progress tracking failures in agent teams, researchers developed SearchOS, a multi-agent collaboration framework. SearchOS models open-domain search as a relational schema completion task and relies on Search-Oriented Context Management (SOCM) to externalize execution states into an Evidence Graph, Coverage Map, and Failure Memory. Using a pipeline-parallel scheduler, SearchOS continuously executes sub-agents to resolve identified coverage gaps. On WideSearch and GISA benchmarks, SearchOS outperformed evaluated single- and multi-agent baselines in search efficiency and final output completeness. (source: https://huggingface.co/papers/2607.15257)

08

MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators

To optimize few-step generative models, researchers created MeanFlowNFT, a framework that applies forward-process reinforcement learning (RL) to average-velocity flow generators. Adapting DiffusionNFT's instantaneous velocity objective, MeanFlowNFT builds an induced instantaneous-velocity predictor while keeping sampling based on average velocities to preserve speed. The framework guarantees strict policy improvements and delivers fast sampling. On SD3.5-M, MeanFlowNFT outperformed prior RL-tuned few-step generators on 6 of 8 metrics. Additionally, a 4-step MeanFlowNFT model achieved a VBench score of 84.33 on Wan 2.1, exceeding the 50-step LongCat-Video RL baseline score of 82.57. (source: https://huggingface.co/papers/2607.15273)