NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-09DEFAULT EDITION
This issue
—
All time
—

AI Blog

4 stories
01

Launches Claude Fable 5 and Claude Mythos 5 Models

Anthropic has released two new large language models: Claude Fable 5, optimized for general enterprise use, and Claude Mythos 5, a specialized version with relaxed safety guardrails designed for authorized cyberdefenders and infrastructure operators. Claude Fable 5 demonstrates high capability in coding and vision; for example, Stripe used the model to complete a major migration on a 50-million-line Ruby codebase in a single day, a task normally requiring a team for two months. Initially deployed in partnership with the U.S. government through Project Glasswing, both models are priced at $10 per million input tokens and $50 per million output tokens. (source: https://www.anthropic.com/news/claude-fable-5-mythos-5)

02

DeepSeek v4 Inference Performance Monitored Across Hardware and Engines

SemiAnalysis and InferenceX monitored the inference performance of the 1.6T parameter DeepSeek v4 large language model from launch through Day 43 across various hardware and software stacks. While initial deployment suffered from compatibility bottlenecks on certain hardware, engineers from AMD SGLang achieved a 100x performance improvement by Day 26. Additionally, SemiAnalysis resolved issues with Nvidia's TensorRT-LLM using custom kernel patches, and analyzed Day 0 performance on the Huawei Ascend 950DT. Open-source inference engines including SGLang and vLLM displayed highly resilient, out-of-the-box disaggregated prefill capabilities during testing. (source: https://newsletter.semianalysis.com/p/deepseekv4-16t-day-0-to-day-43-performance)

03

Confidentially Submits Draft S-1 Statement to the SEC

OpenAI has officially confirmed the confidential submission of a draft S-1 registration statement to the Securities and Exchange Commission (SEC), marking an initial step toward a potential initial public offering. The organization stated that the specific timing for any subsequent steps has not yet been finalized. No supplementary financial details, pricing ranges, or estimated offering sizes were provided in the announcement. (source: https://openai.com/index/openai-submits-confidential-s-1)

04

Our Plan for Ensuring Artificial General Intelligence Benefits Everyone

OpenAI has outlined its strategic plan and long-term vision to manage the transition to artificial general intelligence (AGI) safely. The framework emphasizes maximizing public access to advanced tools, implementing rigorous safety measures, and distributing economic benefits globally to mitigate systemic risks. This strategic roadmap details how the company intends to develop and deploy future systems equitably, balancing commercial development with security and safety protocols. (source: https://openai.com/index/built-to-benefit-everyone-our-plan)

Hacker News

8 stories
01

Claude Fable 5

Anthropic has released Claude Fable 5 and Mythos 5, representing the latest advancements in their suite of generative large language models. The release is accompanied by a comprehensive System Card document detailing safety protocols, alignment procedures, and rigorous benchmark evaluations. The new system architecture emphasizes enhanced reasoning capabilities, complex problem-solving abilities, and robust contextual understanding. According to its technical documentation, Claude Fable 5 demonstrates state-of-the-art performance across standard AI safety and capability evaluations, mitigating common risks associated with hallucination and output bias. This story is discussed extensively on Hacker News with over 900 comments and on other platforms (discussion: https://www.oneusefulthing.org/p/what-it-feels-like-to-work-with-mythos). (source: https://www.anthropic.com/news/claude-fable-5-mythos-5)

02

Apple decided not to roll out Siri in EU after denied request for exemption

Apple has decided to withhold the deployment of its newly upgraded Siri assistant in the European Union following a rejection of its request for a regulatory exemption. The European Commission stated that Apple failed to align its integrated artificial intelligence tools with the block's safety and market regulations. This decision highlights the intensifying clash between global technology giants and European regulators over data privacy, competition, and compliance. Consequently, European users will not have access to the latest AI-driven Siri enhancements, reflecting a growing regional fragmentation in consumer AI technologies. (source: https://www.reuters.com/business/apple-failed-make-its-ai-tool-comply-eu-regulations-eu-commission-says-2026-06-09/)

03

Is Grep All You Need? How Agent Harnesses Reshape Agentic Search

This research paper explores the paradigm of agentic search, investigating whether simple string-matching utilities like 'grep' are sufficient for AI agents, or if sophisticated agent harnesses are necessary. The study evaluates the performance, limitations, and architectures of LLM-based agents navigating complex repositories. It demonstrates how specialized testing harnesses and environments improve the discovery, retrieval, and synthesis of information during autonomous agentic search tasks. (source: https://arxiv.org/abs/2605.15184)

04

Can LLMs Beat Classical Hyperparameter Optimization Algorithms?

This research paper evaluates whether Large Language Models (LLMs) can outperform traditional hyperparameter optimization (HPO) algorithms like Bayesian Optimization and Random Search. Leveraging their in-context learning and pattern recognition abilities, state-of-the-art LLMs were tested on various machine learning models and datasets. The findings indicate that while LLMs show promise in understanding parameter relationships based on semantic prior knowledge, they still face challenges in consistency, computational overhead, and numerical precision compared to specialized classical HPO algorithms. (source: https://arxiv.org/abs/2603.24647)

05

Apple's AI Can Now Change Your Passwords. What Could Possibly Go Wrong?

This analysis explores the security and privacy implications of Apple's latest artificial intelligence features, focusing on the AI system's capability to manage and alter user credentials. While integrating autonomous agents into tasks like credential management promises convenience, it introduces vulnerabilities such as prompt injection attacks or unexpected model behaviors. The author questions the safety boundaries of allowing these agents execution-level control over password vaults and user authenticators, calling for robust verification protocols and local-only processing guarantees. (source: https://www.kylereddoch.me/blog/apples-ai-can-now-change-your-passwords-what-could-possibly-go-wrong/)

06

Unified Controllable and Faithful Text-to-CAD Generation with LLMs

This research addresses the challenges of generating high-quality Computer-Aided Design (CAD) models directly from textual descriptions using Large Language Models (LLMs). Unlike traditional text-to-3D generation that yields uneditable mesh networks, this novel approach focuses on producing parametric CAD files that remain faithful to user-specified constraints. By leveraging the structured code-generation capabilities of LLMs, the unified framework ensures generated models are controllable, physically accurate, and editable within standard engineering workflows. (source: https://arxiv.org/abs/2604.19773)

07

The iPhone's Last Stand?

This analysis explores the evolving landscape of mobile computing, questioning whether the iPhone faces an existential shift due to the rapid acceleration of generative artificial intelligence and autonomous AI agents. As user interfaces shift toward conversational designs, specialized AI wearables, and ambient computing, the traditional app-store model is being challenged. The article evaluates how Apple is integrating artificial intelligence into its silicon and software to counter these threats, and discusses whether Apple can withstand the rise of next-generation autonomous assistants. (source: https://stratechery.com/2026/the-iphones-last-stand/)

08

'Sloppenheimer:' Amazon Employees Mock the Company's AI on Slack

Internal leaks from Amazon's Slack channels reveal that employees are actively criticizing the company's internal artificial intelligence tools, labeling the system 'Sloppenheimer.' Staff members have expressed frustration over the AI's frequent inaccuracies, hallucinations, and unhelpful responses during daily workflows. This internal backlash highlights a growing disconnect between corporate push for rapid generative AI integration and the practical utility of these tools for employees, illustrating the challenges tech companies face when deploying generative LLMs internally. (source: https://www.404media.co/sloppenheimer-amazon-employees-mock-the-companys-ai-on-slack/)

Twitter

7 stories
01

Anthropic Introduces Claude Fable 5 As A Mythos Class AI Model

Anthropic has launched Claude Fable 5, a new Mythos-class large language model designed to deliver advanced reasoning capabilities while meeting strict safety standards. The model has been integrated directly into the Claude Code development workflow, prompting developers to focus on higher-level task verification rather than manual execution. Users experiencing access issues are instructed to use the CLI command /model claude-fable-5 and ensure they are upgraded to Claude Code CLI version 2.1.170. This release has drawn criticism from industry observers regarding its rapid deployment timeline following recent existential risk warnings. (source: https://x.com/AnthropicAI/status/2064394443856232582)

02

Luma Labs Releases Ray 3.2 For Advanced Video Generation On Fal

Luma Labs has officially released the Ray 3.2 model, bringing advanced generative video capabilities to the Fal infrastructure and the Figma Weave platform. The update introduces precise control features, including multiple keyframe support for temporal consistency and facial performance tracking for character animation. Additionally, Ray 3.2 now supports high-quality HDR and EXR output formats designed for professional post-production workflows. Designers can bypass intermediate steps to leverage these tools directly inside Figma, facilitating highly detailed and creative digital design tasks. (source: https://x.com/LumaLabsAI/status/2064388785291567267)

03

Broadcom Secures Partnerships With OpenAI and Anthropic For AI Infrastructure

Broadcom has entered into a strategic collaboration with investment firms Apollo Global Management and Blackstone to finance large-scale AI infrastructure initiatives. The $36 billion program is structured to provide high-performance semiconductor hardware and financial backing to major developers, specifically OpenAI and Anthropic, to meet their computing and scaling requirements. The initiative underscores the massive capital expenditure and resource coordination required to train and deploy advanced large language models on an industrial scale. (source: https://x.com/GaryMarcus/status/2064394211244523730)

04

The Industry Still Fails To Account For Test-Time Compute Scaling

A critical analysis has highlighted the industry's failure to incorporate test-time compute scaling into safety and evaluation frameworks, despite its proven capacity to enhance model performance. While techniques popularized by OpenAI's o1 demonstrate major improvements through scaled inference, many labs continue to rely on static evaluations. Consequently, safety policies and Responsible Scaling Policies (RSPs) fail to mandate inference budget limits when establishing thresholds, exposing a significant gap in corporate and regulatory AI safety planning. (source: https://x.com/polynoamial/status/2064370734806532289)

05

Runway Automation Workflow Streamlines Media Post-Production Processes

An unnamed major media company has integrated Runway's generative capabilities into its post-production pipeline to automate end-credit sequences. The system restyles credits while preserving text overlays for 20 to 25 television episodes per week. This implementation reduces a manual task that typically cost between $10,000 and $15,000 per season, demonstrating the tangible cost-efficiency and utility of deploying creative generative AI in commercial media environments. (source: https://x.com/c_valenzuelab/status/2064343899578024100)

06

Tracing the Technical Evolution of Latent Mixture of Experts Models

A technical breakdown traces the evolutionary lineage of modern machine learning model architectures. The analysis shows that Latent Mixture of Experts (LatentMoE) models stem directly from linear algebra techniques, beginning with eigendecomposition and Singular Value Decomposition (SVD). These foundations led to Low-Rank Adaptation (LoRA), which subsequently influenced Multi-Head Latent Attention (MLA), forming the mathematical underpinnings of today's advanced, large-scale generative model designs. (source: https://x.com/rasbt/status/2064384777663193498)

07

Adaptive Data Integrates Multimodal Capabilities For Advanced Optimization

Adaptive Data has expanded its optimization platform to support multimodal processing, transitioning from text-only optimization to a dual text and image processing framework. This update aims to enhance the overall performance of models that rely on complex visual inputs by optimizing the handling of static image data, bridging data infrastructure gaps to improve the efficiency of heterogenous, real-world AI pipelines. (source: https://x.com/sarahookr/status/2064185216953143578)

huggingface

8 stories
01

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Researchers proposed FlashMemory-DeepSeek-V4, featuring Lookahead Sparse Attention (LSA) powered by a Neural Memory Indexer. This framework resolves GPU memory bottlenecks during ultra-long context serving by predicting future context demands and retaining only query-critical key-value (KV) chunks. Using a backbone-free decoupled training strategy, the indexer is trained independently of the massive model. Across benchmarks like LongBench-v2 and RULER, the paradigm compresses the physical KV cache footprint to 13.5% of the full-context baseline, while consistently preserving downstream reasoning and accuracy up to 500K context scales. (source: https://huggingface.co/papers/2606.09079)

02

Agents' Last Exam

Researchers introduced Agents' Last Exam (ALE), a living benchmark designed to evaluate AI agents on long-horizon, economically valuable, real-world tasks. Developed with over 250 industry experts, ALE aligns with the U.S. federal occupational taxonomy (O*NET / SOC 2018) and spans 13 industry clusters covering more than 1,000 tasks. Baseline evaluations of mainstream configurations reveal an average full pass rate of only 2.6% on the hardest tier, highlighting a massive gap between conventional benchmark success and actual GDP-relevant economic impact in non-physical industries. (source: https://huggingface.co/papers/2606.05405)

03

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short

Researchers introduced Reasoning Arena, an adaptive training framework designed to resolve uninformative reward signals in reinforcement learning with verifiable rewards (RLVR). When all sampled traces of a prompt yield identical outcomes, Reasoning Arena routes them to a judge system that conducts pairwise trace tournaments. To avoid quadratic comparison costs, new traces are ranked against a small, dynamically updated pool of anchor traces using a Bradley-Terry model. This system accelerates training by 27% to 41% and improves mathematics and coding performance by 7.6% over baseline RLVR. (source: https://huggingface.co/papers/2606.09380)

04

End-to-End Context Compression at Scale

To address memory bottlenecks in long-context language model inference, researchers developed Latent Context Language Models (LCLMs). Rehearsing encoder-decoder compression, the team continually pre-trained a family of 0.6B-encoder, 4B-decoder models on over 350 billion tokens. Achieving compression ratios of 1:4, 1:8, and 1:16, LCLMs provide a superior Pareto frontier across general-task performance, compression speed, and peak memory usage compared to standard KV cache compression. This framework serves as an efficient backbone for long-horizon agents to skim and selectively expand compressed context. (source: https://huggingface.co/papers/2606.09659)

05

Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops

An audit of five terminal-agent benchmarks revealed that 16% of 1,968 tasks are vulnerable to reward hacking. To address this, researchers introduced the hacker-fixer loop, an automated method to build exploit-resistant verifiers. The framework orchestrates three LLM agents: a hacker finding exploits, a fixer patching the verifier, and a solver ensuring valid solutions pass. On KernelBench, the loop reduced the attack success rate of frontier models from 62% to 0%. The researchers released the Terminal Wrench dataset containing 323 hackable environments and 3,632 trajectories. (source: https://huggingface.co/papers/2606.08960)

06

Latent Spatial Memory for Video World Models

Researchers introduced Mirage, a generative video world model framework that utilizes latent spatial memory to maintain 3D spatial consistency. Instead of relying on computationally expensive and lossy pixel-space point clouds, Mirage constructs a persistent 3D cache directly within the diffusion latent space by lifting latent tokens via depth-guided back-projection. Novel views are synthesized using direct latent-space warping. Experiments show Mirage achieves up to 10.57 times faster video generation and a 55 times reduction in memory footprint compared to traditional explicit 3D baselines. (source: https://huggingface.co/papers/2606.09828)

07

Why Muon Outperforms Adam: A Curvature Perspective

To understand why the Muon optimizer improves training efficiency over Adam by twofold in large language models, researchers analyzed its training dynamics through a curvature lens. By applying a second-order Taylor approximation to the training landscape, they demonstrated that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. This advantage is driven by lower Normalized Directional Sharpness (NDS) rather than update scale. Decompositions show data imbalance amplifies Muon's NDS advantage, and within-layer curvature remains significantly smaller during mid-to-late training. (source: https://huggingface.co/papers/2606.04662)

08

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Researchers introduced SWE-Explore, a new benchmark designed to isolate and evaluate the repository exploration capabilities of coding agents. Rather than treating coding tasks as binary outcomes, SWE-Explore requires agents to return a ranked list of relevant code regions under a strict line budget across 848 issues, 10 programming languages, and 203 open-source repositories. Line-level ground truth was gathered from successful independent agent trajectories. The evaluation demonstrates that agentic explorers outperform classical retrieval methods, though precise line-level coverage remains a key differentiator. (source: https://huggingface.co/papers/2606.07299)