NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-30ENGLISH EDITION
This issue
—
All time
—

AI Blog

3 stories
01

Introduces Claude Sonnet 5 with Enhanced Agentic Capabilities

Anthropic released Claude Sonnet 5, an agentic large language model engineered for autonomous workflows, browser and terminal navigation, and planning. Benchmarked on BrowseComp and OSWorld-Verified, Sonnet 5 shows significant improvements in coding, reasoning, and tool use over Sonnet 4.6 while closing the performance gap with Opus 4.8. The model is now active as the default option for Free and Pro users, and is accessible across Max, Team, Enterprise, Claude Code, and Claude API tiers. Introductory pricing is set at $2 per million input tokens and $10 per million output tokens through August 31, 2026, before rising to standard rates of $3 and $15. (source: https://www.anthropic.com/news/claude-sonnet-5)

02

Introduces Claude Science AI Workbench for Researchers

Anthropic announced Claude Science, a specialized AI workbench designed to assist scientists with complex research workflows in genomics, structural biology, and proteomics. The platform features a generalist coordinating agent equipped with over 60 pre-configured skills, alongside specialized verification and reviewer agents to validate calculations and source citations. Claude Science enables reproducible artifact generation and manages compute environments across local machines, HPC clusters, and on-demand GPUs. The system is currently available in a beta release for Claude Pro, Max, Team, and Enterprise users. (source: https://www.anthropic.com/news/claude-science-ai-workbench)

03

Enterprises Shift from Tokenmaxxing to Token Budgeting

SemiAnalysis released research on enterprise AI spending showing a structural shift from unconstrained token consumption to strict token budgeting. Following massive utilization in early 2026, such as Uber exhausting its annual Claude Code and Codex budget within four months, enterprises are implementing spending caps starting at $250 to tens of thousands of dollars per month. This shift is driving organizations to downgrade default models and adopt cheaper tiers. Analysis of customer data reveals a vast divide, with the 99th percentile tech-forward customers spending $90,000 per employee annually, while the median Ramp customer spends just $136. (source: https://newsletter.semianalysis.com/p/tokenbudgeting-our-conversations)

Hacker News

8 stories
01

Claude Sonnet 5

Anthropic has announced the release of Claude Sonnet 5, their latest advanced large language model. This new iteration delivers significant improvements in logical reasoning, mathematical computation, and sophisticated code generation. Claude Sonnet 5 excels at complex instruction following, multi-turn dialogue management, and advanced contextual synthesis, demonstrating state-of-the-art benchmarks in standard natural language processing and agentic workflow tasks. Designed to enhance developer productivity and enterprise automation, the model establishes a new paradigm for efficient, safe, and highly capable artificial intelligence systems. (source: https://www.anthropic.com/news/claude-sonnet-5)

02

LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

The LongCat team has announced LongCat-2.0, a large-scale Mixture of Experts (MoE) language model featuring 1.6 trillion total parameters. To maintain high computational efficiency, the model utilizes 48 billion active parameters during inference, balancing high-capacity representation with practical deployment constraints. This design selectively activates routing paths, allowing the architecture to manage massive language tasks without requiring prohibitive computing power. This release reflects the ongoing industry trend toward highly optimized sparse architectures for cost-effective, high-performance runtime operations on modern workloads. (source: https://longcat.chat/blog/longcat-2.0/)

03

Claude Code is steganographically marking requests

A technical analysis has revealed that Claude Code, Anthropic's developer-focused command-line interface, is embedding hidden steganographic markers within its outgoing prompts and API requests. The tool utilizes subtle encoding techniques within the whitespace or formatting of system instructions to covertly flag its traffic as agent-generated. This discovery raises critical questions regarding security, data privacy, and transparency within the developer community. The practice enables external servers to identify and treat Claude Code requests differently from standard human-generated API traffic, highlighting growing tensions in proprietary AI tooling behaviors. (source: https://thereallo.dev/blog/claude-code-prompt-steganography)

04

Claude Desktop is now available on Linux (in beta)

Anthropic has released the beta version of its native Claude Desktop application for Linux operating systems, bringing consistent cross-platform availability to software engineers, system administrators, and open-source advocates. The client provides a dedicated local workspace that matches the existing functionality of the macOS and Windows platforms. Key features of the desktop application include integration with local system resources, faster performance, and customized keyboard shortcuts that bypass the browser experience. This release expands Anthropic's multi-platform strategy to better support highly technical developers in their local environments. (source: https://code.claude.com/docs/en/desktop-linux)

05

Tell HN: Installing Cursor on iOS irreversibly changes your privacy settings

A developer reported on Hacker News that logging into the Cursor AI editor mobile app on iOS silently and irreversibly migrated their account-wide privacy settings. The user's account was automatically switched from the strict 'Privacy Mode (Legacy)' setting—which guaranteed that proprietary code was not stored on Cursor's servers—to a newer, more permissive 'Privacy Mode' that allows data storage for background agents. Following the migration, the legacy option disappeared entirely from the interface menus, raising significant concerns about dark patterns and data sovereignty when using AI development platforms. (source: https://news.ycombinator.com/item?id=48737226)

06

Nano Banana 2 Lite

Google DeepMind has developed Gemini Flash-Lite, a multimodal model optimized for fast, cost-effective image generation and computer vision tasks. Built to reduce latency and computational overhead, the model enables real-time processing and low-latency workflows directly on standard cloud infrastructure. This streamlined, distilled variant helps developers embed multimodal capabilities into web applications without the high computational resources typically demanded by larger-scale models. The release highlights an industry-wide trend toward deploying highly efficient versions of flagship AI systems to expand practical developer access. (source: https://deepmind.google/models/gemini-image/flash-lite/)

07

Claude Science

Anthropic has announced Claude Science, a specialized initiative aimed at leveraging the Claude platform to accelerate advanced scientific research and computational workflows. The platform integrates state-of-the-art reasoning, data analysis, and technical document comprehension to help researchers analyze complex datasets and formulate hypotheses. By synthesizing quantitative data and interpreting complex academic publications, the tool acts as an intelligent assistant to streamline biotechnology, physics, chemistry, and environmental science research. This release represents a targeted effort to deploy large language models for structured, highly rigorous scientific discovery. (source: https://claude.com/product/claude-science)

08

Zluda 6 release (run unmodified CUDA applications on non-Nvidia GPUs)

The ZLUDA project has released version 6 of its high-performance translation layer, which enables unmodified CUDA applications to run seamlessly on non-Nvidia GPU hardware. By targeting alternative architectures like AMD and Intel GPUs, this release addresses vendor lock-in in deep learning and parallel processing ecosystems. The update introduces critical compatibility and performance optimizations for complex GPGPU and high-performance computing workloads. This development helps decouple software dependencies from specific hardware, promoting a more diverse and competitive hardware landscape for machine learning developers. (source: https://vosen.github.io/ZLUDA/blog/zluda-update-q1q2-2026/)

Twitter

8 stories
01

Anthropic Announces Claude Sonnet 5 With Enhanced Agentic Capabilities

Anthropic has officially introduced Claude Sonnet 5, introducing significant advancements in autonomous agentic capabilities. The updated model is designed to operate with greater autonomy, generating complex operational plans and executing tasks across digital environments via integrated tools such as web browsers and command-line interfaces. This launch aims to bridge the gap between logical reasoning and direct tool execution, streamlining automated task management for developers and enterprise teams. While related computer automation capabilities remain in active development, this release establishes a robust foundation for building interactive, tool-using AI agents. (source: https://x.com/AnthropicAI/status/2072032717550833796)

02

Google Unveils Gemini Flash For Video Generation And Conversational Editing

Google has officially launched Gemini Flash, a highly efficient multimodal model engineered for high-performance and cost-effective operations. The model is specifically optimized to manage complex workflows, focusing on advanced video generation capabilities and real-time conversational media editing. By reducing operational latencies and serving costs associated with large generative AI workloads, Google aims to provide developer ecosystems with accessible tools for media creation, interactive storytelling, and dynamic content manipulation without compromising processing speed or output quality. (source: https://x.com/Google/status/2071998701845774786)

03

Nano Banana 2 Lite Launches With Rapid Four Second Text To Image Generation

Google DeepMind has launched Nano Banana 2 Lite, a specialized generative AI model engineered to significantly accelerate text-to-image production pipelines. The optimized system is capable of delivering high-quality visual outputs in four seconds, facilitating rapid creative ideation and streamlined design workflows. This release targets low-latency generation requirements, enabling designers and developers to transition from textual prompts to finished image outputs with minimal friction, thereby improving iteration speed in digital media and professional content creation pipelines. (source: https://x.com/Google/status/2071996507608260678)

04

Luma Labs Launches Dream Machine 2.0 Mini For Advanced Video Generation

Luma Labs has released Seedance 2.0 Mini, a specialized video generation model integrated within the Luma platform canvas. The release features faster generation speeds and refined iterative editing controls, allowing creators to produce high-quality video assets from text prompts. This update streamlines production by combining creative ideation and rendering within a unified workspace. By improving direct interface adjustments and feedback loops, the platform seeks to establish efficient text-to-video production pipelines suitable for professional and experimental filmmakers. (source: https://x.com/LumaLabsAI/status/2072040831041556824)

05

Mastering Loop Engineering For AI-Driven Software Development

Andrew Ng has outlined "loop engineering," a structured methodology designed for building software applications with autonomous AI agents. The framework categorizes the agentic development lifecycle into three functional layers: an agentic coding loop that iterates and tests code independently, a developer feedback loop that maintains human contextual oversight, and an external feedback loop driven by real-world user data. The model emphasizes that while AI agents increasingly automate technical implementation, human judgment remains necessary to guide product requirements. (source: https://x.com/AndrewYNg/status/2071988145667928442)

06

Kling AI Projects Secure Three Prestigious Awards At Cannes Lions 2026

Kling AI has announced that media projects produced using its generative video synthesis tools have received three awards at the Cannes Lions 2026 International Festival of Creativity. The accolades include a Silver Lion in the Film: Consumer Goods category, alongside Bronze Lions in both the Film: B2B and the newly introduced AI Craft categories. This marking demonstrates the increasing commercial adoption of advanced text-to-video tools by professional production teams to generate broadcast-quality marketing assets and cinematic narratives within mainstream advertising frameworks. (source: https://x.com/Kling_ai/status/2071792939706007939)

07

Innovative Long Context Handling Techniques Emerging For Large Language Models

Researchers have introduced advanced sequence processing methodologies optimized for long-context Large Language Models (LLMs). These techniques expand model context windows, allowing systems to ingest, retain, and synthesize significantly larger datasets. By overcoming physical memory constraints, the new architectures resolve long-range dependencies and enable deep document parsing. These optimizations are critical for building reliable AI agents capable of maintaining conversational consistency and processing vast data payloads over extended runtime cycles. (source: https://x.com/natolambert/status/2071970640144556191)

08

First Global Summit Dedicated To World Models Coming To San Francisco

Cristobal Valenzuela has announced the inaugural summit dedicated specifically to the development and study of world models, set to take place in San Francisco this September. The technical event will convene AI researchers, engineers, and industry leaders to discuss breakthroughs in spatial reasoning, environmental simulation, and predictive modeling. As world models represent a key frontier for internalizing and simulating virtual and physical physics, this summit aims to foster collaboration on advanced architectures that transcend conventional training paradigms. (source: https://x.com/c_valenzuelab/status/2072032542195171788)

huggingface

8 stories
01

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Researchers introduced Agents-A1, a 35B Mixture-of-Experts agentic model designed to match trillion-parameter model performance by scaling the agent horizon. The training recipe incorporates a long-horizon knowledge-action infrastructure yielding trajectories averaging 45K tokens, followed by three-stage supervised fine-tuning, domain-level teacher training, and multi-teacher on-policy distillation. Across benchmarks, Agents-A1 achieves leading results on SEAL-0 (56.4) and IFBench (80.6), remaining highly competitive with 1T-parameter models like DeepSeek-V4-pro and Kimi-K2.6. (source: https://huggingface.co/papers/2606.30616)

02

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Researchers introduced OSWorld 2.0, a benchmark of 108 long-horizon computer-use workflows designed to evaluate autonomous agents under highly complex, real-world constraints. Everyday and professional tasks in the suite require an average of 318 tool calls with Claude Opus 4.7 using maximum thinking, compared to only 30 in OSWorld 1.0. Evaluated under a 500-step limit, Claude Opus 4.8 scored highest but achieved a final task completion rate of only 20.6%, demonstrating a substantial performance gap in current agent architectures. (source: https://huggingface.co/papers/2606.29537)

03

TACO: Tool-Augmented Credit Optimization for Agentic Tool Use

Researchers developed Tool-Augmented Credit Optimization (TACO), a GRPO reinforcement learning variant designed to optimize tool usage for multimodal code-tool agents. TACO uses two coupled advantage channels: Differential Answer-Probe Reward (DAPR), which credits tool calls based on their exact effect on final correctness, and Outcome-Gated Advantage Routing (OGAR), which distributes final outcome credit specifically to the responsible code segments. This design suppresses redundant and misleading tool calls without incurring external cost terms. (source: https://huggingface.co/papers/2606.30251)

04

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Researchers formulated and investigated Agentic Abstention, a sequential decision problem evaluating when an autonomous agent should stop tool interactions under uncertainty. Spanning web shopping, terminal environments, and QA, the authors evaluated 13 LLM systems and 2 scaffolds across 28,000 tasks, revealing that larger, capable models often struggle with timely abstention. To address this, they introduced CONVOLVE, a context engineering method that distills trajectories into stopping rules, improving Llama-3.3-70B's timely recall rate from 26.7 to 57.4 on WebShop. (source: https://huggingface.co/papers/2606.28733)

05

LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

Researchers introduced LiveEdit, a novel streaming video editing framework designed for real-time, frame-by-frame edits with content preservation. LiveEdit leverages a three-stage distillation pipeline that transfers editing capabilities from a bidirectional foundation model to an efficient unidirectional streaming editor. To support real-time operation, it uses an AR-oriented mask cache that reuses region-specific computations. Evaluations demonstrate state-of-the-art visual quality among streaming baselines, with inference speeds boosted to 12.66 FPS. (source: https://huggingface.co/papers/2606.26740)

06

DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model

Researchers released DreamForge-World 0.1 Preview, an autoregressive, controllable world model designed for real-time interactive simulation. Derived from Wan2.1-T2V-1.3B with a residual action pathway, the system achieves 14 to 15 FPS at 480p resolution on a single consumer RTX 4090 GPU. It supports live keyboard/mouse controls, mid-stream reprompting, and multi-modal initialization, offering a cost-effective route for deploying localized interactive simulators. (source: https://huggingface.co/papers/2606.30292)

07

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

Researchers proposed GUICrafter, a weakly-supervised GUI agent trained via a curriculum framework designed to reduce reliance on expensive human annotations. In the first stage, the model learns visual grounding directly from unannotated screenshots and webpages; in the second stage, it calibrates via reinforcement learning on a minimal amount of high-quality data. GUICrafter achieves competitive performance to UI-TARS while utilizing only 0.1% of its annotated dataset. (source: https://huggingface.co/papers/2606.29705)

08

AsyncOPD: How Stale Can On-Policy Distillation Be?

Researchers presented a systematic study of gradient staleness in asynchronous on-policy distillation (OPD) for LLM post-training. The study shows that forward KL-divergence remains robust to stale rollouts, whereas reverse KL is vulnerable. By resolving this via local, real-time student updates and multi-sample Monte Carlo estimators, the authors introduced AsyncOPD. AsyncOPD increases training throughput by 1.6x to 3.8x compared to synchronous training baselines. (source: https://huggingface.co/papers/2606.24143)