NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-05DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

Biodefense in the Intelligence Age

OpenAI has announced an action plan to leverage artificial intelligence for biological resilience and defense against emerging synthetic biology risks. The initiative focuses on building secure infrastructure, partnering with biosecurity experts, and establishing robust safety evaluations to safely develop dual-use technologies. By deploying advanced AI capabilities to identify and mitigate biological threats, the organization aims to establish proactive governance and global resilience in the era of advanced models. (source: https://openai.com/index/biodefense-in-the-intelligence-age)

Hacker News

8 stories
01

Transformers Are Inherently Succinct

Researchers have published a theoretical paper selected as one of three outstanding papers at ICLR 2026, proving that Transformer architectures are inherently succinct. The study mathematically demonstrates that Transformers can represent complex functions and execute intricate algorithmic tasks with far fewer parameters and lower computational complexity than previously assumed. By analyzing self-attention mechanisms, the authors establish new theoretical upper bounds for sequence-to-sequence mapping. This formal proof bridges the gap between the empirical success of deep learning and its theoretical foundations, offering a mathematical framework for designing optimized, scale-sensitive neural networks. (source: https://openreview.net/pdf?id=Yxz92UuPLQ)

02

Gemma 4 QAT models: Optimizing compression for mobile and laptop efficiency

Google has introduced the Gemma 4 QAT models, utilizing Quantization-Aware Training (QAT) to optimize model compression for mobile devices and laptops. By integrating quantization parameters directly into the training phase rather than applying post-training adjustments, these models maintain their original cognitive capabilities while achieving a significantly smaller memory footprint and faster local inference times. This technical advancement allows developers to deploy highly capable large language models directly onto consumer-grade edge hardware with constrained memory resources, improving on-device privacy, operational speed, and offline availability. (source: https://blog.google/innovation-and-ai/technology/developers-tools/quantization-aware-training-gemma-4/)

03

Sakana AI's Recursive Self-Improvement (RSI) Lab

Sakana AI has launched its Recursive Self-Improvement (RSI) Lab to develop and deploy autonomous AI agents capable of systematically enhancing their own capabilities. The initiative establishes a continuous feedback loop where AI models write, evaluate, and refine subsequent generations of code and algorithms with minimal human intervention. The RSI Lab serves as a foundational platform for analyzing safety boundaries, alignment paradigms, and scaling laws within self-improving codebases, aiming to accelerate machine learning research and resolve persistent challenges in autonomous system optimization. (source: https://sakana.ai/rsi-lab/)

04

Launch HN: General Instinct (YC P26) – Frontier models on edge devices

General Instinct has launched a platform designed to bring frontier artificial intelligence models to physical hardware systems and edge devices. Founded by robotics industry veterans Guanming and Bill, the startup addresses hardware bottlenecks by optimizing large models for devices with limited memory, low memory bandwidth, and variable network connectivity. As part of this launch, the team open-sourced InstinctRazor, a utility that streamlines model adaptation for local physical hardware. The platform enables low-latency, autonomous decision-making for embedded systems and robotics without relying on cloud-based computing infrastructure. (source: https://news.ycombinator.com/item?id=48414869)

05

Show HN: Lowfat – pluggable CLI filter that saved 91.8% of my LLM tokens

An open-source, pluggable command-line interface tool called Lowfat has been released to reduce token consumption for Large Language Model agents. Operating as a single binary shell wrapper, Lowfat filters verbose CLI outputs, such as those from Kubernetes, by stripping out structural noise before sending context to LLM endpoints. In testing over a two-month period, the utility achieved an average token reduction of 91.8 percent, with some commands reaching savings of 93.9 percent. This tool helps lower API operational costs, reduces context window load, and increases model response speed without sacrificing agent decision quality. (source: https://github.com/zdk/lowfat)

06

Open Code Review – An AI-powered code review CLI tool

Alibaba has released Open Code Review, an open-source, AI-powered command-line interface tool designed to automate developer code reviews. Utilizing Large Language Models, the tool analyzes code changes to identify bugs, detect security vulnerabilities, and evaluate adherence to formatting standards directly from the terminal. By integrating into continuous integration and continuous deployment pipelines, it automates early-stage code assessments to reduce manual human review overhead and accelerate pull request processing times. It aims to bridge static analysis and dynamic reasoning with context-aware recommendations. (source: https://github.com/alibaba/open-code-review)

07

Fine-tuning an LLM to write docs like it's 1995

A developer has executed a project focused on fine-tuning a Large Language Model to write technical documentation in a mid-1990s computing style. Using a specialized training dataset consisting of vintage software manuals, technical publications, and legacy guides from 1995, the author aligned the LLM's vocabulary, tone, formatting, and structural output. The project demonstrates how fine-tuning methodologies can be applied to capture highly specific historical personas and specialized writing aesthetics, offering a practical blueprint for developers pursuing niche domain adaptations. (source: https://passo.uno/fine-tuning-docs-llm/)

08

Programmers will document for Claude, but not for each other

An analysis piece explores a psychological shift in software engineering, noting that developers are increasingly writing comprehensive documentation specifically for AI assistants like Anthropic's Claude. While programmers historically neglect documentation for human colleagues due to delayed feedback, they provide detailed architectural context to LLMs to receive immediate dividends in the form of code generation and accurate debugging. This behavioral dynamic suggests a new software development paradigm where AI models act as the primary consumers of technical documentation, transforming it into an active tool for codebase maintenance. (source: https://blog.plover.com/2026/03/09/#documentation-wins-2)

Twitter

5 stories
01

Anthropic Demonstrates Claude's Capability In Advanced NMR Spectroscopy

Anthropic has released a new research report demonstrating that its Claude Opus 4.7 model can interpret complex Nuclear Magnetic Resonance (NMR) spectroscopy data to solve molecular structures. According to the publication, the model matches or surpasses dedicated, professional chemistry software on specific structural validation tasks. This marks a notable expansion of large language model capabilities into specialized laboratory workflows and chemical analysis. The update indicates progress in leveraging generative models as analytical scientific assistants to identify organic molecular compounds. (source: https://x.com/AnthropicAI/status/2062979607448682731)

02

Sakana AI Launches New RSI Lab In Tokyo For Adaptive AI Research

Sakana AI has announced the launch of its new RSI Lab in Tokyo, focusing on the research and development of adaptive, open-ended artificial intelligence. This initiative aims to build autonomous systems capable of building and optimizing other AI architectures, targeting recursive self-improvement. Building on two years of research, the laboratory represents a shift towards automated algorithm design and system discovery. This development explores next-generation agentic frameworks that can learn and adapt autonomously in complex conditions. (source: https://x.com/hardmaru/status/2062948594597208557)

03

Runway Showcases Fifty Crowns Fully AI-Generated In-Game Cinematic

Runway has debuted "Fifty Crowns", a narrative-driven in-game cinematic project generated entirely through its suite of generative AI tools. Developed by a single creator in under one week, the short film demonstrates automated animation, rendering, and rapid video production workflows. The release serves as a showcase for individual creators using generative video models to produce high-quality cinematic stories that historically required dedicated studio animation departments, highlighting a potential shift in the timeline and resource cost of game development. (source: https://x.com/runwayml/status/2062898193126302111)

04

Launch Of The AutoScientist Challenge With Fifty Thousand Dollars In Prizes

AutoScientist has launched the AutoScientist Challenge, an intense four-week AI competition offering a total prize pool of $50,000. Featuring 10 distinct categories, the challenge focuses on testing autonomous agents and automated scientific systems in solving complex research tasks. This initiative operates alongside plans to support researchers deploying frontier models on public platforms like Hugging Face and Kaggle to accelerate machine learning development cycles in domains ranging from medical research to underserved languages. (source: https://x.com/sarahookr/status/2062886117075308698)

05

Kling AI Launches Anniversary II Creation Showreel Contest

Kling AI has initiated its Anniversary II Creation Showreel Contest, running from June 3 to June 17, 2026, to showcase user projects generated with its multimodal AI video tools. Users are invited to submit creative videos under the themes of Anniversary Memories or Creation Showreel. The competition offers incentives including cash prizes, platform credits, and gift boxes, highlighting the growing utility and consumer adoption of generative video synthesis platforms for custom media production. (source: https://x.com/Kling_ai/status/2062927429359092005)

huggingface

8 stories
01

OPRD: On-Policy Representation Distillation

Researchers proposed On-Policy Representation Distillation (OPRD), a novel distillation framework that aligns student and teacher representations across intermediate hidden layers during rollouts. Unlike traditional on-policy distillation which works in output space, OPRD bypasses the language model head entirely. This design eliminates high sampling variance over large vocabularies and utilizes richer structural information. In evaluations on reasoning tasks like AIME 2024/2025 and AIMO, OPRD successfully closed the student-teacher performance gap where standard baselines plateaued, while delivering a 1.44x training speedup and reducing memory consumption by 54% compared to top-k on-policy distillation. (source: https://huggingface.co/papers/2606.06021)

02

Latent Reasoning with Normalizing Flows

Researchers introduced NF-CoT, a latent reasoning framework that models continuous intermediate thoughts using normalizing flows inside the large language model backbone. Traditional chain-of-thought methods force discrete verbalization, which is computationally expensive and rigid. By instantiating a TARFlow-style flow, NF-CoT generates continuous thoughts via a dedicated head while maintaining text generation through the standard head. This design supports native left-to-right decoding, causal KV-cache compatibility, and exact likelihood estimation. On coding benchmarks, NF-CoT improved pass rates over discrete chain-of-thought and prior latent baselines while significantly lowering intermediate reasoning costs. (source: https://huggingface.co/papers/2606.06447)

03

The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs

Researchers formulated inference-time budget allocation for Large Language Models as a global constrained optimization problem governed by economic principles. Modeling per-query reasoning utility with a shifted-surge function, they derived an optimal allocation policy based on a global shadow price that balances marginal utility. Under this theory, they introduced Constrained Latent-utility Equilibrium Allocation for Reasoning (CLEAR), which reallocates tokens from insolvent queries to solvable queries. CLEAR significantly improved the Pareto frontier of total token cost versus mean accuracy, achieving up to a 3x accuracy improvement over uniform allocation in resource-scarce regimes. (source: https://huggingface.co/papers/2606.03092)

04

World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis

Researchers proposed World-Language-Action (WLA) models, a new class of embodied foundation models that integrate language reasoning, future state prediction, and action synthesis. Built on an autoregressive Transformer backbone, the prototype WLA-0 uses 2 billion active parameters to jointly predict subtasks, subgoal images, and robot actions. It achieves 40 ms per inference on an NVIDIA RTX 5090. In evaluations, WLA-0 demonstrated strong multi-task capabilities, reaching a 92.94% success rate on RoboTwin2.0 Clean and a 56.5% success rate on RMBench, while showing potential to learn from cross-embodiment videos lacking action labels. (source: https://huggingface.co/papers/2606.05979)

05

Flash-WAM: Modality-Aware Distillation for World Action Models

Researchers introduced Flash-WAM, a step-distillation framework designed to enable real-time control in joint video-action generative models. Traditional consistency distillation fails on multimodal systems because video and action streams utilize mismatched noise schedules. Flash-WAM addresses this by employing a modality-aware architecture, utilizing linear-gradient-scaling for action streams and variance-preserving parameterization for video streams. This compression reduces inference to a single step, cutting latency on an NVIDIA L40S from 8.1 seconds to 348 ms. The model preserves performance with an 85.5% success rate on RoboTwin 2.0 and a 95.7% success rate on LIBERO. (source: https://huggingface.co/papers/2606.05254)

06

Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation

Researchers proposed a reinforcement learning (RL) approach to help large language models acquire the meta-skill of utilizing in-context linguistic knowledge for translating unseen, low-resource languages. Instead of relying on fine-tuning or manual context encoding which often overfit specific languages, the method uses a surface-level translation metric (chrF) as the reward to guide the model. Empirically, the RL-trained models successfully extracted and applied grammar rules from the context, outperforming traditional in-context learning and supervised fine-tuning baselines on entirely unseen test languages. (source: https://huggingface.co/papers/2606.06428)

07

LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing

Researchers developed LoomVideo, an efficient 5-billion-parameter unified architecture designed for multimodal video generation and editing. While previous frameworks rely on massive models and costly token concatenation, LoomVideo replaces the text encoder with a Multimodal Large Language Model and leverages a Deepstack feature-injection mechanism. It introduces a zero-overhead Scale-and-Add conditioning approach that adds clean source latents to noised target latents. This design eliminates sequence doubling, resulting in a 5.41x acceleration in inference speed compared to models of similar capability while achieving state-of-the-art results. (source: https://huggingface.co/papers/2606.06042)

08

EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management

Researchers introduced EvoDS, a self-evolving autonomous data science agent designed to expand its action set and manage long-term context using agentic reinforcement learning. EvoDS utilizes an Autonomous Skill Acquisition mechanism to synthesize and validate executable skills, alongside an Adaptive Context Compression strategy to prevent token-overflow issues in long-horizon pipelines. Evaluated across four benchmarks, EvoDS outperformed state-of-the-art open-source data science agents by an average of 28.9% while completely eliminating out-of-token failures. (source: https://huggingface.co/papers/2606.03841)