NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-08-03DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

Architectural Deep Dive into Kimi K3 and Kimi Delta Attention

SemiAnalysis published a technical deep dive into the architectural design of the Kimi K3 open frontier model. To optimize memory and computation, Kimi K3 introduces a hybrid attention mechanism featuring Kimi Delta Attention (KDA). By eliminating the softmax operation, KDA converts quadratic computational complexity into linear complexity. This enables past key and value vectors to be compressed into a hidden state matrix functioning as associative memory. Additional mechanisms in the model include compressed memory, attention across depth, and latent expert routing. (source: https://newsletter.semianalysis.com/p/kimi-k3-the-manos-the-mythos-the)

Hacker News

8 stories
01

Qwen3.8-Max: A New Bar for Coding and Cowork

The Qwen team has released Qwen3.8-Max, a specialized conversational AI model engineered to handle complex programming, multi-step logical reasoning, and interactive pair programming. Optimizing for high-context understanding, the model aims to reduce developer friction during real-time code generation and debugging tasks. This release marks a significant step forward in code-centric assistant models that support human-in-the-loop developer cooperation, paving the way for next-generation software engineering workflows. (source: https://qwen.ai/blog?id=qwen3.8)

02

Ten advances in mathematics and theoretical computer science

OpenAI released a report highlighting ten developments at the intersection of mathematics and theoretical computer science where advanced computational models and deep learning are used to solve long-standing conjectures. This research is also discussed under reports of OpenAI's unreleased model codenamed 'Astra' resolving major math problems. These milestones demonstrate how machine learning can serve as a catalyst for proof verification, algorithmic efficiency, and symbolic reasoning. (source: https://openai.com/index/ten-advances-in-mathematics/)

03

MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video

ComfyUI announced immediate day-0 support for the newly released MiniMax H3 model within its node-based ecosystem. MiniMax H3 features open weights, enabling users to run generative workflows locally. The model provides advanced multimodal generation features, including native audio synthesis and high-resolution 2K video generation. By integrating these capabilities natively, ComfyUI allows creators to build complex media generation pipelines without relying on external cloud APIs. (source: https://blog.comfy.org/p/minimax-h3-day-0-support-in-comfyui)

04

AirLLM 70B inference with single 4GB GPU

The open-source AirLLM repository was released, allowing developers to execute massive 70B parameter models like LLaMA-2-70B on consumer-grade hardware with as little as 4GB of VRAM. It achieves this resource efficiency through advanced memory-saving techniques, including layered inference, block-by-block loading, and aggressive model quantization. This project democratizes access to large language models, enabling local, privacy-focused inference without requiring expensive multi-GPU setups. (source: https://github.com/lyogavin/airllm)

05

Show HN: Product analytics (and evals) for agent sessions on your MCP

Theodore and Louis launched Armature, a novel product analytics and evaluation platform designed specifically for Model Context Protocol (MCP) tool providers. By implementing a lightweight three-line software development kit (SDK) supporting TypeScript, Python, and Go, developers can reconstruct user sessions. The platform visualizes agent-user conversations directly within interfaces like Claude or ChatGPT, groups session data to reveal use cases, and highlights agent errors. (source: https://armature.tech/)

06

Launch HN: Hoplite (YC S26) – Effortlessly deploy cloud coding agents

Bence and Ryan launched Hoplite, a YC S26 backed platform designed to streamline the deployment of coding agents directly within cloud environments. Hoplite automates onboarding by porting local developer sessions, context memories, and Model Context Protocol (MCP) servers to cloud-hosted infrastructure. The platform assists developers with quality assurance, feature validation, and managing persistent, scalable AI-driven software development workflows in the cloud. (source: https://hoplite.sh)

07

Show HN: Nightcrawler – A local AI pentesting agent running on a smartphone

Nightcrawler, an open-source local AI penetration testing agent designed to run directly on a mobile smartphone, was launched on GitHub. The system allows cybersecurity professionals to execute autonomous network scans, exploit analysis, and system auditing tasks offline. By running inference fully locally on edge devices with limited computing resources, the tool prioritizes user privacy and ensures operational confidentiality without relying on external cloud APIs. (source: https://github.com/garagehq/nightcrawler/)

08

Smaller, faster, safer: running Kimi and GLM at scale

Cloudflare introduced new performance optimizations for deploying popular large language models, such as Kimi and GLM, across its global edge network. By utilizing advanced model quantization and memory-efficient runtime environments, the system reduces latency and computational overhead for edge-native inference. Security and privacy protections are built directly into the runtime environment to safeguard sensitive developer and user data during decentralized processing. (source: https://blog.cloudflare.com/smaller-faster-safer-models/)

Twitter

8 stories
01

Hailuo AI Releases H3 Multimodal Video Generation Model As Open-Weight

Hailuo AI has officially released its H3 multimodal video generation model as an open-weight release. This strategic shift from proprietary development to open weights aims to empower creators and developers by providing open access to powerful video synthesis technology. The updated H3 model features enhanced multimodal support enabling up to twelve reference files, including images, video, and audio. It introduces commercial-grade motion design capabilities, advanced title animations, and exceptional fidelity in animating static illustrations and rendering text within generated frames, representing a significant open-access milestone in the video generation landscape. (source: https://x.com/Hailuo_AI/status/2084284120625225750)

02

Sakana AI Introduces Upgraded Namazu Large Language Model API

Sakana AI has officially announced the launch of an upgraded version of its specialized large language model, Namazu, made available through an API interface. Designed with a distinct emphasis on Japanese linguistic nuances and cultural context, the model aims to provide developers with robust, localized Japanese-language processing capabilities. This release represents a significant technical advancement for the Tokyo-based laboratory, offering improved reasoning and natural language generation quality tailored specifically for high-performance localized applications in the expanding language-specific artificial intelligence market. (source: https://x.com/hardmaru/status/2084277077885452790)

03

Optimizing Inference Efficiency With Advanced Asari Agentic Models

Eric Schmidt highlighted the potential of the Asari approach in enhancing inference efficiency and lowering operational costs for data centers utilizing advanced AI agents. By leveraging these agentic frameworks, organizations can achieve superior performance levels compared to traditional methods. The shift emphasizes optimizing computational economics, inference latency, and resource consumption, rendering autonomous agentic systems more accessible for large-scale enterprise deployments and cloud infrastructure environments. (source: https://x.com/ericschmidt/status/2084310193731514782)

04

Kling AI Introduces MCP Tutorial for Automated Food Promo Video Creation

Kling AI has released a new tutorial demonstrating how to leverage its Model Context Protocol (MCP) to automate food promotional video creation. The system streamlines the entire production pipeline, managing artistic style analysis, storyboard development, and video synthesis within an automated creative workflow. This protocol integration highlights how companies can apply generative tools to meet specific marketing needs and optimize visual advertising pipelines. (source: https://x.com/Kling_ai/status/2084247919897780389)

05

Proposing New Impact-Based Benchmarks For Large Language Models

Cristobal Valenzuela has proposed shifting evaluation metrics toward impact-based benchmarks for large language models as standard benchmarks reach saturation. The proposal suggests that new model releases must demonstrate practical utility and tangible real-world contributions, such as discovering disease cures, solving complex mathematical problems, or accelerating general scientific discovery, rather than merely optimizing benchmark scores. (source: https://x.com/c_valenzuelab/status/2084307929264443649)

06

Energy Based Models and the Role of Optimization in Objective Driven AI

Yann LeCun discussed the foundational role of optimization during the inference phase in the context of Energy-Based Models (EBM). He emphasized that employing optimization at inference time is a core concept that underpins Objective-Driven AI architectures. This mechanism allows EBMs to differ from traditional models by focusing on energy surfaces to define behavior, improving goal-oriented reasoning and adaptability in machine learning systems. (source: https://x.com/ylecun/status/2084182218822332854)

07

New SMPL C++ Plugin For Unreal Engine Enhances Human Character Rendering

Joachim Tesch from the Max Planck Institute has introduced a new SMPL C++ plugin designed for Unreal Engine 5.8 to optimize human body model rendering. While highly realistic models like SMPL-X and SMPL+H traditionally impose substantial computational burdens on forward rendering passes, this integration addresses latency bottlenecks, enabling real-time skin deformation within game engine environments. (source: https://x.com/Michael_J_Black/status/2084261681811575284)

08

AutoScientist Platform Evolves Into Multimodal AI Capability

The AutoScientist platform has announced an architectural expansion integrating multimodal capabilities to handle heterogeneous data format processing. Moving beyond text-only inputs, the platform natively parses images, charts, and documents to automate complex research analysis. This update aims to bridge automated reasoning and visual data interpretation in scientific and corporate research settings. (source: https://x.com/sarahookr/status/2084279262128099389)

huggingface

8 stories
01

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Researchers proposed Reinforcement Learning with Self-Verifiable Rewards (RLSVR) and its multi-agent self-play instantiation, SpyRL, to scale LLM reinforcement learning beyond deterministic domains like math and coding. RLSVR transforms open-ended LLMs tasks into verifiable proxy environments where agent interaction outcomes generate internal reward signals. In experiments spanning text summarization, creative writing, and mathematical reasoning, SpyRL outperformed existing self-improvement methods on non-verifiable tasks while maintaining strong performance on verifiable reasoning challenges (source: https://huggingface.co/papers/2607.23802).

02

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers introduced Stable Advantage Fusion (SAF), a training framework designed to combine reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) without triggering entropy collapse. SAF uses a four-stage pipeline featuring a sparsify-then-compress mechanism for magnitude control and a warm-up-then-anneal mechanism for temporal control. Evaluating the framework using GRPO across seven mathematical reasoning and code generation benchmarks on Qwen3-1.7B/4B/8B models showed aggregate score improvements of 0.51% to 2.70% alongside more stable training (source: https://huggingface.co/papers/2607.29209).

03

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Researchers introduced Counterfactual Sensitivity Credit Reallocation (CSCR), an extension of the GRPO reinforcement learning framework designed to improve long-CoT reasoning in large language models. Investigating credit assignment mechanisms, the study revealed that privileged shifts in on-policy self-distillation fail to provide reliable answer-aligned directions, with magnitudes reflecting token-level sensitivity rather than reasoning value. By reducing credit for highly sensitive surface tokens and renormalizing token advantages, CSCR consistently outperformed GRPO baselines on long-CoT mathematical reasoning benchmarks under identical update budgets (source: https://huggingface.co/papers/2607.27888).

04

Enhancing Rubric-based RL via Self-Distillation

Researchers proposed Criterion-Distilled Policy Optimization (CriPO), a training method that enhances rubric-based reinforcement learning for large language models via on-policy self-distillation. CriPO addresses the core limitations of Unexplored Criteria and Suppressed Criteria without introducing a train-inference mismatch. It employs a criterion-injection self-teacher to introduce missing behaviors and a counterfactual self-teacher to reallocate positive advantages to useful tokens in negative rollouts. On science and medicine benchmarks, CriPO outperformed standard rubric-based RL, achieving stronger final performance with twice fewer optimization steps (source: https://huggingface.co/papers/2607.18082).

05

EMBL AI Librarian: Life-Sciences Knowledge Layer for AI Agents

Researchers developed the EMBL AI Librarian, a specialized knowledge layer designed to upgrade the Europe PMC interface for AI agents in the life sciences. The system uses a single LLM to orchestrate literature search by generating subqueries, searching over 40 million indexed records, and extracting precise evidence. On ScholarQABench, the Librarian improved Citation F1 by more than 16 points over strong baselines, and a GPT-5.4 agent utilizing this retrieval layer scored approximately 8 points higher on the open-form LitQA2 benchmark compared to standard web search (source: https://huggingface.co/papers/2607.28229).

06

Mental World Modeling

Researchers formulated Mental World Modeling (MWM), a theoretical framework that integrates physical and hidden mental states into predictive world models. The framework is instantiated in MENTIS, a training-free baseline that handles state parsing, action decomposition, and transition simulations. Evaluating eight LLM-based world models on a custom situated decision dataset of text, image, and sound-video stories demonstrated that explicitly tracking agents' internal beliefs, intentions, and desires is essential to predicting human decisions, exposing the limitations of purely physical world models (source: https://huggingface.co/papers/2607.27201).

07

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

LlamaIndex launched ExtractBench, an agentic evaluation benchmark designed for schema-guided extraction of enterprise documents. Comprising 4,869 pages across 370 documents, 8 business domains, and 67 document types, ExtractBench measures value accuracy, record completeness, grounding, and cost. Evaluation results show that commercial VLMs often struggle and truncate lists in longer documents, whereas coding agents maintain high accuracy at high costs. The LlamaExtract Agentic Plus model ranked first across all evaluated metrics, matching coding agent accuracy at a lower operational cost (source: https://huggingface.co/papers/2607.29677).

08

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Researchers introduced the SaliTrap Benchmark to study Salience Bias in large language models, where models prioritize explicit, irrelevant inputs over implicit commonsense prerequisites. Evaluating 12 state-of-the-art LLMs revealed that models suffer from knowledge suppression rather than knowledge absence when faced with distracting features. Stripping away the misleading task framing and utilizing context-free knowledge probes recovered over 90% of model compliance failures. Furthermore, light-weight, inference-time prompting was shown to close this reasoning gap without additional model training (source: https://huggingface.co/papers/2607.28478).