NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-19DEFAULT EDITION
This issue
—
All time
—

Hacker News

3 stories
01

John Jumper leaves Google to join Anthropic

Renowned scientist John Jumper, who led the development of Google DeepMind's AlphaFold, has left Google to join rival AI safety and research lab Anthropic. Jumper's prior work on AlphaFold solved the 50-year-old protein folding problem using advanced machine learning models. At Anthropic, he is expected to apply his expertise in deep learning, structural biology, and complex scientific systems to accelerate the organization's frontier model capabilities and steer safe artificial intelligence development. This high-profile hire highlights the intensifying industry competition for top-tier generative AI and research talent. (source: https://twitter.com/JohnJumperSci/status/2068001285173834106)

02

Is AI ruining our skills? Early results are in – and they're not good

Early empirical scientific studies reveal a measurable decline in critical thinking, coding proficiency, and problem-solving skills due to heavy reliance on generative AI tools and large language models. Research suggests that delegating cognitively demanding tasks to automated systems deprives users of the productive struggle needed for skill acquisition. This dynamic creates a feedback loop where human expertise degrades, potentially limiting innovation and the ability to audit AI errors. Experts call for structural changes in education and professional training. (source: https://www.nature.com/articles/d41586-026-01947-1)

03

Show HN: Music from every single country on Earth, generated in Esperanto

A new interactive global music map application showcases generative AI capabilities by synthesizing songs representing every country on Earth, with all lyrics written and performed in the constructed language Esperanto. The project leverages multimodal AI generation, localization, and computational linguistics to compose music across diverse regional genres while standardizing the vocal layer. This experiment demonstrates the capability of modern generative audio synthesis tools to produce cohesive cross-cultural content, highlighting automated regional style emulation and linguistic adaptation in low-resource environments. (source: https://elevenexperiments.com/around-the-world?lang=eo&genre=de)

Twitter

8 stories
01

Kling AI 3.0 Turbo Launches With Enhanced Video Generation Capabilities

Kling AI has released its 3.0 Turbo model, which is now integrated and available for use on the Selfyz AI web platform. This update focuses on enhancing generative video synthesis performance by delivering faster generation speeds, improved audio-video synchronization, and smoother motion sequences. This release targets creators requiring high-fidelity and rapid production cycles for AI-powered visual projects. The launch was accompanied by a thematic demonstration video showcasing the model's capabilities in realistic human motion and texture synthesis. (source: https://x.com/Kling_ai/status/2067986767199048086)

02

Impact of Anthropic and U.S. Government Controls on AI Model Sovereignty

Andrew Ng has raised critical concerns regarding Anthropic's usage restrictions on its Claude Fable 5 model and subsequent U.S. government export controls. Ng argues that limiting how developers deploy frontier models and implementing strict international access regulations hinders global competition under the guise of safety. These actions are driving nations to invest heavily in open-source projects and independent AI infrastructure to secure local AI sovereignty. He advocates for open research and international cooperation to build a collaborative foundation for future artificial intelligence development rather than closed control mechanisms. (source: https://x.com/AndrewYNg/status/2068039709126017356)

03

Claude Code Bug Affecting Usage Limits Successfully Resolved

Anthropic's development team has identified and resolved an operational bug in Claude Code affecting roughly 3% of Max and Pro users. The software error incorrectly displayed weekly usage limits and occasionally prevented users from sending messages. To mitigate the disruption, the engineering team executed a full reset of the five-hour and weekly usage limits for all affected developer accounts. Normal service operations have been fully restored and the team issued an apology for the service interruption. (source: https://x.com/ClaudeDevs/status/2067802163498352929)

04

Nobel Laureate John Jumper Reportedly Departs Google DeepMind

John Jumper, a 2024 Nobel Prize winner in Chemistry, is reportedly departing from Google DeepMind. Jumper is highly celebrated in the machine learning and computational biology communities for his pioneering contributions to protein structure prediction through the development of AlphaFold alongside Demis Hassabis. While specific details surrounding his departure remain limited, Jumper's exit represents a major change in research leadership at Google DeepMind as the organization continues its work applying advanced AI models to scientific discovery. (source: https://x.com/GaryMarcus/status/2068044403126759700)

05

Opus 4.8 Evaluated With Text-To-CAD Skills On CADGenBench

Researchers from Hugging Face have evaluated the Opus 4.8 large language model using the CADGenBench framework. The collaborative testing focused on benchmarking the model's capabilities in text-to-CAD generation, assessing how effectively natural language instructions can be converted into precise computer-aided design outputs. The evaluations showed significant spatial reasoning and technical drafting proficiencies. This project highlights the growing intersection of language models with technical design software and establishes new open-source standards for automated manufacturing and design workflows. (source: https://x.com/Thom_Wolf/status/2067959167164293358)

06

The Underrated Significance Of Supervised Fine-Tuning Methods

Nathan Lambert has highlighted a critical gap in machine learning literature regarding the lack of empirical focus on Supervised Fine-Tuning (SFT) methodologies. Despite SFT serving as a fundamental post-training pillar for large language models, there is a shortage of rigorous academic analysis examining fine-tuning protocols. Lambert argues that as model architectures scale and grow in complexity, a deep, experimental understanding of these SFT protocols will be essential for stabilizing training pipelines and achieving optimal model performance. (source: https://x.com/natolambert/status/2068046795310236047)

07

Cost and Instability Constraints Facing Reinforcement Learning Speedruns

Nathan Lambert has analyzed the current bottlenecks slowing down the adoption of reinforcement learning (RL) speedruns. High computational expenses serve as the primary hurdle, with model instability requiring researchers to run multiple training seeds to yield reliable outcomes. These iterative individual trials raise execution costs to approximately $100 per entry. Although these financial and technical constraints currently limit widespread developer participation, optimizing RL training workflows remains critical to improving stability and reducing operational resource consumption. (source: https://x.com/natolambert/status/2067977078201618660)

08

Leveraging OpenAI Models to Assist Families with Rare Genetic Diseases

OpenAI's o3 model has been utilized by researchers in the medical field to support families navigating rare genetic diseases. Greg Brockman highlights that the algorithmic capabilities of the o3 model successfully streamlined complex diagnostics and analysis. Given that this model is already over a year old, its deployment illustrates the rapid development cycle of large language models. This application emphasizes the growing utility of generative AI in clinical environments and points to the potential of current-generation architectures to handle advanced medical research. (source: https://x.com/gdb/status/2068016345451831480)

huggingface

8 stories
01

Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents

Researchers have proposed evaluating LLM agents using "predictive validity"—the correlation between in-sample and out-of-sample rank—to address rank instability in aggregate-score leaderboards. The study evaluates fourteen parallel implementations of an MCP-based industrial-agent benchmark across alternative orchestrations, reasoning modes, and infrastructure optimizations. It identifies a twelve-tier measurement apparatus to expose deployment-relevant dimensions that static leaderboards collapse, demonstrating that ranking configurations do not reliably transfer to out-of-distribution settings. (source: https://huggingface.co/papers/2606.19704)

02

FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines

Researchers introduced Fully Autonomous Prompt Optimization (FAPO), a framework that leverages Claude Code to automatically optimize multi-step LLM pipelines. FAPO diagnoses failures, proposes prompt edits, and shifts to changing chain structures when structural bottlenecks are detected. In evaluations across six benchmarks and three task models, FAPO outperformed the GEPA baseline in 15 out of 18 comparisons, yielding a mean gain of 14.1 percentage points, which rose to 33.8 percentage points in tasks requiring structural pipeline adjustments. (source: https://huggingface.co/papers/2606.19605)

03

Context-Aware RL for Agentic and Multimodal LLMs

Researchers proposed ContextRL, a context-aware reinforcement learning method designed to improve long-horizon reasoning and multimodal performance in LLMs. Instead of supervising only the final answer, ContextRL rewards models for matching a query-answer pair with the correct context from a contrastive pair. Evaluated across coding and multimodal domains, ContextRL achieved average gains of 2.2% over standard GRPO on five long-horizon benchmarks, and a 1.8% average improvement across twelve visual question answering benchmarks. (source: https://huggingface.co/papers/2606.17053)

04

LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents

Researchers developed LedgerAgent, an inference-time method designed to maintain observed task states for tool-calling agents in separate ledgers. By rendering states clearly in the prompt and checking state-dependent constraints before executing environment-changing tools, the method prevents policy violations. LedgerAgent was evaluated across four customer-service domains using various open- and closed-weight LLMs, demonstrating improved average pass-k performance over standard prompt-based tool-calling baselines, with particularly strong gains in strict multi-trial consistency. (source: https://huggingface.co/papers/2606.20529)

05

Current World Models Lack a Persistent State Core

A systematic diagnostic evaluation using the new WRBench benchmark revealed that current video world models fail to maintain a persistent state core when objects go unobserved. By assessing 9,600 generated videos across 23 models and four control paradigms, researchers found that models resume returning targets in the state they were abandoned rather than evolving them while out of view. The study highlights that world-state persistence remains a critical blind spot that scaling and geometric priors alone do not solve. (source: https://huggingface.co/papers/2606.20545)

06

Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe

A study on low-precision LLM training identified that non-uniform FP4 formats, such as E2M1, suffer from "Shrinkage Bias" due to geometric bin asymmetry, causing training instabilities. To address this, researchers introduced UFP4, a uniform 4-bit training recipe utilizing E1M2/INT4 grids with Random Hadamard Transform and targeted stochastic rounding. Evaluated on Dense 1.5B, MoE 7.9B, and MoE 124B models during long-run pretraining, the UFP4 recipe consistently reduced BF16-relative loss degradation compared to standard E2M1-based baselines. (source: https://huggingface.co/papers/2606.20381)

07

Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages

Researchers introduced Multi-LCB, a benchmark that extends the Python-restricted LiveCodeBench to evaluate LLMs across twelve programming languages. Multi-LCB translates LiveCodeBench competitive programming tasks into other languages while preserving original contamination controls. Evaluating 24 instruction and reasoning LLMs revealed Python-specific overfitting and language-specific contamination, establishing the new framework as a robust tool for assessing multilingual code-generation capabilities. (source: https://huggingface.co/papers/2606.20517)

08

Understanding the Behaviors of Environment-aware Information Retrieval

Researchers conducted the first systematic study of how LLMs learn to adapt their query formulation strategies to different retrievers using reinforcement learning. The empirical analysis demonstrated that different retrievers demand distinct, non-transferable optimal query styles (such as descriptive versus question-like formulations). By introducing a branching-based rollout technique to stabilize multi-step training, the team successfully trained LLMs to automatically customize queries based on specific retriever behaviors. (source: https://huggingface.co/papers/2606.16817)