NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-02DEFAULT EDITION
This issue
—
All time
—

AI Blog

8 stories
01

Submits Confidential Draft S-1 for Proposed IPO

Anthropic announced it has confidentially submitted a draft registration statement on Form S-1 to the U.S. Securities and Exchange Commission (SEC) for a proposed initial public offering (IPO) of its common stock. The submission gives the artificial intelligence safety and research company the option to list publicly following the completion of the SEC's standard review process. The timing, price range, and total number of shares to be offered have not yet been determined, as the IPO remains subject to market conditions. This announcement was made under Rule 135 of the Securities Act of 1933. (source: https://www.anthropic.com/news/confidential-draft-s1-sec)

02

Building the Infrastructure for the Intelligence Age in Michigan

OpenAI has officially broken ground on a new gigawatt-scale data center facility in Michigan as part of its ongoing Stargate infrastructure initiative. This development aims to construct advanced computing infrastructure to support the massive computational requirements of next-generation artificial intelligence systems. OpenAI projects that the site will expand technical capacity and foster local economic growth by creating new jobs within the region. The project is designed to scale hardware resources necessary to advance large language models and other intelligence capabilities. (source: https://openai.com/index/stargate-michigan-data-center)

03

Using AI Agents to Find Compiler Miscompiles

SemiAnalysis researcher Justin Lebar demonstrated how autonomous AI agents can automate compiler fuzzing to discover software bugs. Over a single afternoon, Lebar spent more than $10,000 running agents powered by OpenAI's Codex and Anthropic's Claude, revealing dozens of plausible bugs in LLVM and 80 miscompiled programs in NVIDIA's closed-source ptxas low-level compiler. Using ChatGPT 5.5 to write and adapt the fuzzer helped avoid duplicate bug traps, while preliminary tests with Anthropic's Claude Opus 4.8 and Claude Code "ultracode" mode reduced the cost of finding medium-to-high severity bugs by approximately 80%. (source: https://newsletter.semianalysis.com/p/finding-miscompiles-for-fun-not-profit)

04

Expands Project Glasswing to One Hundred Fifty New Organizations

Anthropic announced the expansion of Project Glasswing, its collaborative software security initiative, to approximately 150 new organizations across 15 countries. These partners span critical infrastructure sectors including healthcare, power, water, communications, and hardware. The initiative previously used the Claude Mythos Preview model with 50 partners to identify over 10,000 high- or critical-severity security flaws. Along with the expansion, Anthropic highlighted the release of Claude Security, a codebase scanning and patching tool utilizing Claude Opus 4.8, and began offering its vulnerability-finding tools to trusted security teams. (source: https://www.anthropic.com/news/expanding-project-glasswing)

05

Introducing Db9.ai: Serverless Postgres Built for AI Agents

Ed Huang introduced db9.ai, a serverless Postgres database designed specifically as a workspace for autonomous AI agents. Unlike traditional managed databases engineered for human developers, the platform combines SQL tables, files, vectors, jobs, and environment branches into a single unified workspace. This architecture allows AI agents to perform operations instantly via a command-line interface without complex cloud infrastructure setup. Notable features include fs9 for storing files alongside database tables, and environment-wide branching to safely isolate, test, and execute complex agent behaviors. (source: https://me.0xffff.me/what-is-db9-ai.html)

06

Anthropic Claude Growth On Amazon Bedrock Drives AWS Margins Higher

Amazon Web Services (AWS) saw its EBIT margins increase by 213 basis points quarter-over-quarter, driven primarily by customer spending growth on Anthropic's Claude models via Amazon Bedrock. According to an analysis by SemiAnalysis, this margin outperformance contrasts with flat or declining margins at rival cloud providers. The growth is rooted in AWS's dominant share of third-party model API spend, its unique token-as-a-service business model, rapid datacenter power procurement, and deep vertical integration utilizing custom Graviton and Trainium silicon chips. (source: https://newsletter.semianalysis.com/p/anthropic-growth-and-bedrock-mix)

07

Rise of Coding Agents Shifts Software Engineering toward Coordination Management

Lorin Hochstein analyzed how the rise of autonomous coding agents is transforming developers from active coders into managers of multiple parallel AI agents. Drawing parallels to high-performance computing network interconnect bottlenecks and thread pool performance degradation, the article highlights how coordination costs escalate rapidly as more agents are introduced. This shift in developer roles coincides with industry trends toward corporate downsizing in middle management, which increases the span of control for remaining engineers. Managing multiple coding agents alongside human colleagues is expected to introduce significant coordination overhead to daily engineering workflows. (source: https://surfingcomplexity.blog/2026/05/24/the-coming-coordination-calamity/)

08

Codex Transforms Productivity and Knowledge Work

OpenAI published the "Next Era of Knowledge Work" report detailing how its Codex model is transforming productivity across non-technical professional domains. The report outlines how Codex is being used for AI-powered research, data analysis, workflow automation, and content creation beyond its initial software development scope. By automating repetitive administrative and analytical tasks, Codex allows employees to focus on high-level problem-solving. This shift illustrates Codex's evolution from a developer-focused tool to a broader productivity driver for general business operations and digital workflows. (source: https://openai.com/index/codex-for-knowledge-work)

Hacker News

8 stories
01

Anthropic scales Claude Mythos to critical infrastructure in 15 countries

Anthropic has deployed its state-of-the-art AI model, Claude Mythos, to manage and optimize critical infrastructure systems across 15 countries. Designed for highly sensitive and regulated sectors, the model processes real-time telemetry data to predict system failures and automate maintenance schedules in energy grids, water treatment facilities, and public transportation networks. This rollout integrates advanced natural language understanding and real-time reasoning safety guardrails into international cybersecurity frameworks and public service operations, advancing the deployment of agentic AI workflows. (source: https://techcrunch.com/2026/06/02/anthropic-scales-claude-mythos-to-critical-infrastructure-in-15-countries/)

02

Microsoft announces Scout, an autonomous AI agent built on OpenClaw

Microsoft has announced the release of Scout, an autonomous artificial intelligence agent built on the open-source OpenClaw framework. Designed to run persistently in the background as an always-on coworker, Scout marks a shift from passive chatbots to active systems executing multi-step workflows. While marketed for workplace productivity, leaked internal documents highlight strategic goals to foster high user engagement. Widely discussed on Hacker News, the release signals an aggressive expansion into the autonomous agent market. (source: https://www.computerworld.com/article/4180103/microsoft-unveils-scout-an-autonomous-ai-agent-built-on-openclaw.html) (discussion: https://news.ycombinator.com/item?id=48374503)

03

Trump signs downsized AI order after weeks of reversals

President Trump has signed a downsized executive order focusing on artificial intelligence innovation and national security following weeks of administrative policy reversals. The downsized directive aims to accelerate the deployment of cutting-edge AI technologies and high-performance computing capabilities while establishing safety guidelines to secure infrastructure. By narrowing its policy scope, the order establishes a regulatory framework that seeks to foster public-private partnerships without placing overly restrictive regulations on the domestic technology sector. (source: https://www.politico.com/news/2026/06/02/trump-signs-downsized-ai-order-00946389)

04

MAI-Thinking-1

Microsoft has released MAI-Thinking-1, a reasoning-intensive AI initiative accompanied by the launch of seven new MAI models designed to run as a specialized hillclimbing machine. This architecture focuses on iterative optimization and computational reasoning processes to dynamically improve system performance. In addition, the release includes MAI-Code-1-Flash, a model optimized for rapid code generation, autocomplete, and debugging in developer workflows. The suite targets complex problem-solving and automated optimization with lower latency. (source: https://microsoft.ai/news/introducing-mai-thinking-1/) (discussion: https://microsoft.ai/news/introducingmai-code-1-flash/)

05

Expanding Project Glasswing

Anthropic has announced the expansion of Project Glasswing, an initiative focused on improving the interpretability, safety, and transparency of frontier artificial intelligence models. As deep learning networks scale, understanding their internal mechanisms remains a critical bottleneck. The project deploys mechanistic interpretability techniques to map and control neural pathways. By scaling up research engineering resources, Anthropic aims to discover new methodologies in neural decoding and representation engineering, establishing auditing benchmarks for autonomous agents. (source: https://www.anthropic.com/news/expanding-project-glasswing)

06

Rethinking Search as Code Generation

Perplexity Research has introduced a framework that conceptualizes the web search process as a dynamic code generation task rather than a static retrieval task. By leveraging large language models to write and run programmatic instructions, the system dynamically fetches, filters, and processes web-scale data in real-time. This method enables multi-step reasoning and precise data manipulation to improve accuracy and retrieval relevancy, establishing a new paradigm for next-generation search agents. (source: https://research.perplexity.ai/articles/rethinking-search-as-code-generation)

07

Bringing Up DeepSeek-V4-Flash on AMD MI300X

This technical report documents the deployment and optimization of the DeepSeek-V4-Flash large language model on AMD's MI300X enterprise GPU hardware. The testing outlines compatibility within AMD's ROCm software stack as an alternative to NVIDIA's CUDA, showing how to resolve library dependencies and configure memory allocation. The results demonstrate that the AMD MI300X is a viable hardware accelerator for running large-scale language models in low-latency production environments. (source: https://fergusfinn.com/blog/deepseek-v4-flash-mi300x/)

08

GitHub Copilot App

GitHub has unveiled the preview of the GitHub Copilot App, an integrated developer application that embeds generative artificial intelligence capabilities directly into programming environments. Acting as an intelligent companion, the tool utilizes advanced generative AI models trained on public code repositories to provide real-time suggestions, contextual assistance, and code auto-completions. The app aims to minimize repetitive boilerplate coding and optimize modern software development pipelines through deep ecosystem integration. (source: https://github.com/features/preview/github-app)

Twitter

5 stories
01

Anthropic Expands Claude Mythos Preview To Global Organizations

Anthropic has officially expanded its Project Glasswing initiative, granting access to the Claude Mythos Preview model to approximately 150 additional organizations across more than fifteen countries. This strategic rollout allows a diverse range of global entities to integrate, test, and experiment with these advanced preview model capabilities. The expansion marks a key milestone in scaling Claude's infrastructure and gathering international developer feedback. (source: https://x.com/AnthropicAI/status/2061796327986454883)

02

Google DeepMind Unveils Co-Scientist For Advanced Research Breakthroughs

Google DeepMind has introduced Co-Scientist, a new specialized artificial intelligence agent designed to serve as a collaborative research partner for scientific discovery. The system integrates advanced machine learning to automate complex analytical tasks, navigate vast datasets, and streamline experimental design alongside human researchers. This initiative represents a milestone in leveraging agentic research models for systematic scientific innovation. (source: https://x.com/vivnat/status/2061868456912568672)

03

Implementing Self-Correction Feedback Loops for Claude Code

Claude Developers have released a walkthrough exploring self-correction feedback methodologies to enable Claude Code to perform internal validation of its own output. By encoding manual verification steps directly into agentic workflows, developers can build closed-loop feedback mechanisms that improve the reliability and accuracy of autonomous code execution without constant human intervention. (source: https://x.com/ClaudeDevs/status/2061900434722496604)

04

OpenAI Models Become Available Through Amazon Bedrock Infrastructure

OpenAI has integrated its advanced large language models into the Amazon Bedrock infrastructure. This partnership allows enterprise developers to deploy OpenAI models directly through their AWS cloud environment, leveraging Amazon's scalability, security, and governance standards to simplify model deployments, maintain strict data compliance, and reduce operational latency. (source: https://x.com/gdb/status/2061603059781054878)

05

Kling AI Introduces Interactive World Cup Themed Video Generation Features

Kling AI has launched a new interactive generative video feature designed for the World Cup, enabling users to generate personalized cheering and dancing videos. The release leverages Kling's advanced generative video synthesis technology to showcase accessible multimodal creation, allowing fans to interact dynamically with tournament themes via automated animation tools. (source: https://x.com/Kling_ai/status/2061825163440660891)

huggingface

8 stories
01

Multi-Agent Computer Use

Researchers have proposed Multi-Agent Computer Use (MACU), a framework designed to overcome the limitations of single-agent computer use systems. MACU uses a manager model to decompose complex tasks into a directed acyclic graph (DAG) and dispatches subagents to execute nodes in parallel while continuously updating the plan. Tested across OSWorld, Online-Mind2Web, WebTailBench, and Odysseys, the MACU framework improves task completion rates by 3.4% to 25.5% compared to single-agent baselines, and it accelerates execution by approximately 1.5 times on the long-horizon Odysseys benchmark. (source: https://huggingface.co/papers/2606.01533)

02

OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents

Researchers have developed OpenWebRL, an open framework designed to train visual web agents directly on live websites using online multi-turn reinforcement learning (RL). The framework includes live-browser infrastructure, supervised initialization, and trajectory-level success evaluation. Using this setup, the team trained OpenWebRL-4B with just 400 initialization trajectories and 2,200 open-ended RL tasks. OpenWebRL-4B established a new open-source state of the art, achieving 67.0% success on Online-Mind2Web and 64.0% on DeepShop, outperforming previous open agents and remaining competitive with proprietary systems. (source: https://huggingface.co/papers/2606.02031)

03

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

To address the lack of specialized evaluations for personalized tools, researchers introduced MCP-Persona, the first benchmark for evaluating language model agents using the Model Context Protocol (MCP) in simulated personal applications. Covering platforms such as Reddit, Xiaohongshu (Rednote), Lark (Feishu), and Slack, MCP-Persona evaluates how agents handle local databases and individual accounts. Experimental results indicate that current state-of-the-art models struggle with personalized tool interactions, highlighting a significant performance gap in custom, user-centric environments. (source: https://huggingface.co/papers/2606.02470)

04

ESPO: Early-Stopping Proximal Policy Optimization

Researchers proposed ESPO (Early-Stopping Proximal Policy Optimization) to mitigate the computational waste and noise of standard reinforcement learning in language models. ESPO monitors trajectories on-the-fly using surrogate regret calculated from sampling logits, terminating rollouts early when failure is detected. This technique concentrates negative temporal-difference errors near the point of failure without needing an external reward model. Evaluated on mathematical reasoning using DeepSeek-R1-Distill-Qwen-7B, ESPO outperformed standard PPO on AIME 2024 (46.28% vs 45.25%) and MATH-500 (87.42% vs 85.43%) while reducing rollout token usage by over 20%. (source: https://huggingface.co/papers/2605.29860)

05

LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation

To combat identity drift and accumulated errors in autoregressive video synthesis, researchers developed LongLive-RAG, a retrieval-augmented generation framework. Instead of relying solely on a localized sliding active window, LongLive-RAG treats previously generated latents as a content-addressable history, retrieving relevant contexts via query embeddings. It incorporates a Window Temporal Delta Loss to suppress redundant local similarities. Across several autoregressive backbones, LongLive-RAG reduces error accumulation during long video generation and achieves the top rank on the VBench-Long benchmark. (source: https://huggingface.co/papers/2606.02553)

06

LVSA: Training-Free Sparse Attention for Long Video Diffusion

Researchers introduced Long Video Sparse Attention (LVSA), a training-free, block-sparse attention mechanism designed for video diffusion transformers to bypass the quadratic memory costs of dense attention. LVSA combines a structured window pattern with rotating global anchors to eliminate long-range temporal loop artifacts. Built with a FlashInfer kernel, it yields up to 3.17x compute reductions on Wan 2.1 1.3B and enables out-of-memory generations on HunyuanVideo 1.5 at a 2x horizon on a single GPU. The authors also released VQeval to properly evaluate loopy video failures. (source: https://huggingface.co/papers/2605.31057)

07

Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs

A study has demonstrated a major vulnerability in text watermarking schemes, showing that linear ensembles of multiple language models can trivially eliminate statistical signatures. The authors proposed WASH (Watermark Attenuation via Statistical Hybridisation) to address vocabulary misalignment and tokenization differences across models. Testing across six watermarking schemes and three LLMs showed that averaging outputs from only three to five models suppressed detection z-scores from up to 300 down below the detection threshold, showing that watermarking requires unprecedented coordination among model providers to remain viable. (source: https://huggingface.co/papers/2605.30501)

08

FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search

Researchers have proposed FineVerify, a fine-grained self-verification framework designed to scale test-time compute for search agents. FineVerify decomposes complex questions into checkable sub-questions, scores sampled candidate trajectories against these specific criteria, and selects the highest-scoring candidate. Across four agentic search benchmarks, FineVerify consistently improved the average accuracy of GPT-5-mini by 8.2 points and Gemini-3-flash by 5.6 points with only four sampled trajectories. Additionally, it enabled GPT-5-mini to surpass the frontier GPT-5 model on the BrowseComp-Plus benchmark when using 12 samples. (source: https://huggingface.co/papers/2606.00660)