NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-04ENGLISH EDITION
This issue
—
All time
—

AI Blog

6 stories
01

GPT Rosalind Enhances Biological Reasoning and Life Sciences Research Capabilities

OpenAI announced the introduction of new capabilities to its GPT-Rosalind model to advance life sciences research. The updated model offers enhanced biological reasoning, medicinal chemistry expertise, and genomics analysis. Additionally, GPT-Rosalind introduces experimental workflow capabilities designed to assist researchers in streamlining laboratory and computational procedures. This update focuses on expanding the utility of large language models within specialized scientific domains and providing advanced intelligence for complex biological challenges. (source: https://openai.com/index/introducing-new-capabilities-to-gpt-rosalind)

02

ChatGPT Introduces New Memory System to Keep Context Across Conversations

OpenAI introduced a new memory system for ChatGPT designed to remember user preferences and maintain context across separate conversations. This capability allows the large language model to retain important details and user-specified guidelines over time, eliminating the need for users to repeat information in new chats. This update focuses on streamlining long-term interactions and improving the continuity of the assistant's helpfulness across recurring tasks and ongoing projects. (source: https://openai.com/index/chatgpt-memory-dreaming)

03

Endava Redesigning Software Delivery Around AI Agents

IT services firm Endava is deploying autonomous AI agents across its enterprise to accelerate its software delivery lifecycle and automate internal workflows. By leveraging ChatGPT Enterprise and Codex, the firm is building an AI-native organizational culture designed to enhance developer productivity. The integration focuses on deploying agentic workflows to handle complex programming tasks, assist with code generation, and optimize delivery pipelines while maintaining high standards of code quality and security. (source: https://openai.com/index/endava-frontiers)

04

Wasmer Builds Edge Node.js Runtime Using OpenAI Codex

Wasmer utilized OpenAI Codex powered by GPT-5.5 to develop and build a new Node.js runtime tailored for edge environments. By integrating the model into their development workflow, Wasmer successfully accelerated their engineering velocity by 10x to 20x. This efficiency boost enabled the team to ship the completed edge runtime product within weeks rather than months, highlighting the impact of specialized generative AI coding tools on software development. (source: https://openai.com/index/wasmer)

05

Andreessen Horowitz Launches New Global Initiatives and Leadership Appointments

Andreessen Horowitz announced a set of global initiatives to drive partnerships with allied nations and help growth-stage companies scale internationally. Former NSA and White House official Anne Neuberger joins as General Partner and Head of Global Affairs to focus on AI, robotics, and defense modernization. This consolidation of global policy, international go-to-market strategies, and partnerships connects portfolio founders with sovereign, institutional, and strategic capital. (source: https://a16z.com/a16zs-global-mission/)

06

The Next Frontier of Visual AI Is Code

Andreessen Horowitz investor Yoko Li details how visual AI is shifting from pixel-native generation to code-native generation. Instead of generating uneditable pixels via diffusion models, modern visual AI tools increasingly output underlying source code such as SVGs, HTML/CSS, React components, and Blender scripts. This programmatic representation allows designers, engineers, and AI agents to continuously iterate and edit the output within production workflows. (source: https://a16z.com/the-next-frontier-of-visual-ai-is-code/)

Hacker News

7 stories
01

When AI Builds Itself: Our progress toward recursive self-improvement

Anthropic has released a report detailing progress and safety boundaries regarding recursive self-improvement in AI systems. The document outlines technical pathways where current large language models assist in code generation, reinforcement learning feedback, and automated dataset curation to optimize subsequent generations. It highlights the safety protocols, alignment challenges, and evaluation frameworks required to manage self-improving systems responsibly and mitigate intelligence explosion risks. (source: https://www.anthropic.com/institute/recursive-self-improvement)

02

The ways we contain Claude across products

Anthropic's engineering team detailed the architectural strategies and safety boundaries used to isolate and contain Claude across its product environments. To protect systems from untrusted outputs and malicious code execution exploits, Anthropic utilizes a containment pipeline featuring secure sandboxing, microsegmentation, and input/output sanitization. This secure infrastructure allows Claude to run code and interact with tools autonomously while maintaining system-level security. (source: https://www.anthropic.com/engineering/how-we-contain-claude)

03

I built a vulnerable app and spent $1,500 seeing if LLMs could hack it

An independent researcher conducted an experiment spending $1,500 in API costs to test whether modern large language models could autonomously discover and exploit vulnerabilities in a custom-built web application. While the LLM-based autonomous agents recognized common vulnerability patterns quickly, they struggled to execute complex, multi-step exploits and handle context-specific application logic, demonstrating current limitations in fully replacing human penetration testers. (source: https://kasra.blog/blog/i-spent-1500-seeing-if-llms-could-hack-my-app/)

04

KVarN: Native vLLM backend for KV-cache quantization by Huawei

Huawei's Computing Systems Laboratory has open-sourced KVarN, a native vLLM backend engineered specifically for KV-cache quantization. KVarN optimizes Key-Value cache management directly within the vLLM inference engine, reducing the memory footprint of large language models during deployment. This allows for higher throughput and larger batch sizes, making LLM inference more efficient on constrained hardware resources within enterprise and cloud-native environments. (source: https://github.com/huawei-csl/KVarN)

05

Show HN: Boxes.dev: ditch localhost; run Claude Code and Codex in the cloud

Two former Gem engineers launched Boxes.dev, a cloud-only agentic development platform designed to run Claude Code and Codex agents in persistent, dedicated cloud containers. The service aims to eliminate the limitations of localhost setups, such as heavy git worktrees, resource constraints during parallel testing, and local hardware dependency, enabling developers to run parallel AI coding agents continuously and access environments from mobile devices. (source: https://boxes.dev)

06

Failing grades soar with AI usage, dwindling math skills in Berkeley CS classes

UC Berkeley professors are reporting a sharp rise in failing grades in undergraduate computer science courses, attributing the decline to students' over-reliance on generative AI tools for programming assignments and a regression in foundational math skills. Faculty observe that AI assistance prevents students from developing critical debugging and problem-solving skills, leading to poor performance in controlled, AI-free exam environments. This trend is prompting educators to reconsider course prerequisites and exams. (source: https://www.dailycal.org/news/campus/academics/failing-grades-soar-as-professors-see-greater-ai-usage-dwindling-math-skills-in-uc-berkeley/article_16fad0bf-02cb-4b8c-8d88-888ffd9f8608.html)

07

Show HN: Cost.dev (YC W21) – making agents cost-aware and cheaper to call

The creators of Infracost have launched Cost.dev, a new command-line interface engineered to make autonomous coding agents cost-aware. Rebuilt from the ground up to support AI callers, the tool optimizes output formatting to minimize token consumption for large language models like Anthropic's Claude. This allows agents to evaluate cloud infrastructure cost impacts in pull requests efficiently while keeping LLM operational costs low. (source: https://cost.dev/)

Twitter

8 stories
01

Nemotron 3 Ultra Released With Mamba-2-Attention Hybrid Architecture

Nvidia has released the Nemotron 3 Ultra, a new open-weight large language model emphasizing a strong capability-to-efficiency ratio. The model architecture integrates a Mamba-2-attention hybrid stack with the LatentMoE framework to scale performance while optimizing computational efficiency. Expanding upon the parameter size of the previous Super variant, this release is designed to help developers build advanced generative artificial intelligence applications with specialized attention routing mechanism designs. (source: https://x.com/rasbt/status/2062575436836536377)

02

Sakana AI Develops Japan’s First One Trillion Parameter Agentic Model

Sakana AI has announced an initiative to build Japan's first one-trillion-parameter, agent-native AI model. Supported by the Ministry of Economy, Trade and Industry (METI) via the GENIAC program, this project focuses on developing an autonomous system optimized specifically for long-horizon deep research tasks. By scaling model parameters to this magnitude, the company aims to improve multi-step, autonomous workflows within the domestic AI landscape. (source: https://x.com/hardmaru/status/2062450123121262624)

03

Lindy Transitions Entire Traffic Load To DeepSeek V4 Model

Lindy has migrated 100% of its production-grade traffic to the DeepSeek V4 large language model, replacing its previous reliance on Anthropic models. The platform estimates this transition will save millions of dollars in infrastructure expenses. This strategy highlights a broader enterprise trend of benchmarking proprietary models to reduce operational and inference overhead without sacrificing functional capabilities. (source: https://x.com/Thom_Wolf/status/2062558091279818809)

04

Nvidia Adopts Multi-Teacher On-Policy Distillation For Model Training

Nvidia has integrated Multi-Teacher, On-Policy Distillation into its model training pipeline to optimize post-training processes. This methodology guides the fine-tuning and reinforcement learning phases of large language models by leveraging multiple teacher models, building on techniques popularized by Microsoft and DeepSeek R1. Industry observers expect this distillation methodology to establish a new benchmark for upcoming models. (source: https://x.com/natolambert/status/2062528878997029030)

05

Runway Experiences Rapid Growth in Enterprise Adoption and User Engagement

Runway has reported significant growth in enterprise adoption for its creative generative AI platform over a six-week period. The company recorded a 50% increase in token consumption, a 140% growth in power users, and achieved an inflection point with Net Dollar Retention surging to 300%. These metrics highlight the deeper integration of generative video synthesis in professional workflows. (source: https://x.com/c_valenzuelab/status/2062614359747055618)

06

Anthropic Internal Data Highlights Claude's Role In AI Development Speed

Anthropic has released internal data demonstrating that its Claude model accelerates AI development workflows. The findings suggest that Claude's current reasoning and coding capabilities provide a viable pathway toward recursive self-improvement and increased autonomy in AI systems. The dataset serves as a milestone in evaluating how modern language models catalyze their own progress and architectural development. (source: https://x.com/ch402/status/2062573658888081769)

07

Andrew Ng Launches New Course On Efficient Large Language Model Serving

Andrew Ng has launched a new educational course focused on efficient serving of large language models, created in collaboration with Red Hat. Instructed by Cedric Clyburn, the course teaches techniques for managing memory usage, optimizing the KV cache, quantizing models to reduce memory footprint, and utilizing vLLM for high-performance concurrent deployment. (source: https://x.com/AndrewYNg/status/2062576164657664469)

08

Significant Enhancements Announced for ChatGPT Memory Capabilities

Greg Brockman has announced significant enhancements to ChatGPT's memory capabilities, focused on improving the conversational system's contextual retention across multiple user sessions. The update is designed to make the storing and retrieving of interaction history more robust, reducing the need for repetitive user instructions and streamlining long-term, multi-session workflows. (source: https://x.com/gdb/status/2062608071411540196)

huggingface

8 stories
01

Cosmos 3: Omnimodal World Models for Physical AI

We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture. By supporting highly flexible input-output configurations, Cosmos 3 seamlessly unifies critical modalities for Physical AI -- effectively subsuming vision-language models, video generators, world simulators, and world-action models into a single framework. Our evaluation demonstrates that Cosmos 3 establishes a new state-of-the-art across a diverse suite of understanding and generation tasks, demonstrating omnimodal world models as scalable, general-purpose backbones for embodied agents. (source: https://huggingface.co/papers/2606.02800)

02

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

We propose ThoughtFold, a framework that leverages fine-grained preference learning to mitigate redundant explorations for efficient reasoning. ThoughtFold employs an introspective strategy to identify redundancy within each correct trajectory, which yields a spectrum of candidate sub-trajectories. Leveraging this spectrum, we introduce a masked preference optimization objective that explicitly penalizes redundant explorations and encourages the model to directly bridge essential reasoning segments, effectively folding its reasoning chains into a more concise path. Extensive experiments show that ThoughtFold significantly enhances efficiency. It reduces the token usage of DeepSeek-R1-Distill-Qwen-7B by approximately 56% while maintaining state-of-the-art accuracy. (source: https://huggingface.co/papers/2606.03503)

03

Streaming Communication in Multi-Agent Reasoning

We introduce StreamMA, a multi-agent reasoning system that streams each reasoning step to downstream agents as soon as it is generated, pipelining adjacent agents and thus reducing latency. Surprisingly, this pipelining also improves effectiveness: because multi-step reasoning quality is non-uniform and early steps are more reliable than later ones, working with these reliable early steps instead of the full chain prevents error-prone late steps from misleading downstream agents. We formalize both advantages with the first closed-form joint analysis of stream, serial, and single protocols, deriving the effectiveness ordering, speedup upper bound, and cost ratio. (source: https://huggingface.co/papers/2606.05158)

04

Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation

We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length history at constant cost. Existing methods mainly curate memory with predefined KV-cache schedules, fixed-ratio heuristic compression, or inference-time RoPE adaptation. These designs inevitably lose historical information and amplify compounding errors due to their limited cache window and ignorance of autoregressive generation noise. Inspired by human memory consolidation, Echo-Infinity replaces handcrafted memory curation with learnable Memory Query, which are updated by attention and a gating mechanism when past frames are evicted from the local window. (source: https://huggingface.co/papers/2606.04527)

05

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

To address this gap, we introduce AutoLab, a new benchmark for ultra long-horizon closed-loop optimization. AutoLab consists of 36 realistic, expert-curated tasks spanning four diverse domains: system optimization, puzzle & challenge, model development, and CUDA kernel optimization. Each task begins with a correct but deliberately suboptimal baseline and challenges agents to improve it within a strict wall-clock budget. Evaluating 17 state-of-the-art models reveals the dominant predictor of success is not the quality of an agent's initial attempt, but its persistence in repeatedly benchmarking, editing, and incorporating empirical feedback. (source: https://huggingface.co/papers/2606.05080)

06

Large Language Models Hack Rewards, and Society

We observe that societal regulations are structurally similar to reward functions. They define measurable outcomes, thresholds, and exceptions, while often leaving institutional intent only partially specified. We hypothesise that the RL training process may exploit these gaps and therefore ask whether models' well-known tendency to hack reward functions during RL can scale into a more consequential failure mode named societal hacking: discovering loopholes in the rules society runs on. To study this phenomenon, we introduce SocioHack, a sandbox of 72 societal environments, and find that within these environments, reward hacking naturally emerges and leads to regulatory loophole discovery. (source: https://huggingface.co/papers/2606.04075)

07

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

Inspired by Friedrich Hayek's economic theory of decentralized coordination in markets, we study this question through an agent economy in which agents compete via auctions for the right to act, exchange payments, and accumulate wealth from environmental rewards. These simple economic signals induce decentralized credit assignment, driving planning without global orchestration or explicit communication protocols. The population evolves through economic selection: effective agents accumulate wealth and are mutated via exploitation, while ineffective ones go bankrupt and are replaced via exploration. We show that, initialized with weak agents, the economy produces emergent multi-step reasoning strategies. (source: https://huggingface.co/papers/2606.02859)

08

Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents

This paper presents Agent libOS, a library-OS-inspired runtime substrate for LLM agents. Agent libOS runs above a conventional host operating system; it does not implement hardware drivers, kernel-mode isolation, or a POSIX-compatible operating system. Instead, it treats an agent as an AgentProcess: a schedulable execution subject with process identity, parent-child lineage, lifecycle state, a tool table derived from an AgentImage, typed Object Memory, explicit capabilities, human queues, checkpoints, events, and audit records. Its central design rule is tools are libc-like wrappers; runtime primitives are the authority boundary. (source: https://huggingface.co/papers/2606.03895)