NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-03ENGLISH EDITION
This issue
—
All time
—

AI Blog

2 stories
01

Details Claude Fable 5 Cyber Safeguards and Jailbreak Framework

Anthropic announced the global redeployment of Claude Fable 5 along with details on its cybersecurity safeguards and a proposed AI jailbreak severity framework. Developed with Glasswing partners, this framework helps manage dual-use risks by training safety classifiers to categorize activities into four tiers: prohibited, high-risk dual-use, low-risk dual-use, and benign use. Anthropic has expanded its safety margin, requiring requests to appear clearly safe to avoid blocking, and launched a HackerOne program to crowdsource jailbreak discoveries. Feedback is being collected at cyber-safeguards@anthropic.com. (source: https://www.anthropic.com/news/fable-safeguards-jailbreak-framework)

02

Meta Accelerates Compute and Datacenter Procurement to Become Neocloud

Meta is rapidly accelerating its datacenter and compute procurement, contracting over 5GW of capacity across cloud and colocation services in the first half of the year. Analysis outlines four main use cases for this expansion: frontier model training under Meta Superintelligence Labs, scaling recommendation systems tenfold, and operating as a Neocloud provider. Meta is also reportedly in final negotiations with Anthropic to secure private instances of Claude for internal tasks and to power a token-as-a-service endpoint, mirroring high-margin compute models. (source: https://newsletter.semianalysis.com/p/meta-compute-everyone-wants-to-be)

Hacker News

8 stories
01

Introducing the Safari MCP Server for Web Developers

Apple has launched the Safari Model Context Protocol (MCP) server, designed to assist web developers using AI-powered development tools. The Model Context Protocol is an open standard that enables large language models and autonomous AI assistants to safely interact with local development environments, data sources, and system tools. By introducing this dedicated server, Safari allows developer-focused AI systems to securely access browser state, inspect DOM structures, interact with the Web Inspector, and assist in real-time debugging or page optimization. This integration represents a major step in connecting local AI coding assistants directly into the browser workflow. (source: https://webkit.org/blog/18136/introducing-the-safari-mcp-server-for-web-developers/)

02

Alibaba to ban Claude Code in workplace over alleged backdoor risks, source says

Alibaba Group is planning to ban the use of Claude Code, Anthropic's recently released AI-powered agentic coding tool, across its workplace. The decision stems from security concerns regarding potential backdoor vulnerabilities and data exfiltration risks. Because Claude Code operates as an autonomous agent with the ability to execute terminal commands and modify local files, Alibaba's information security teams fear that proprietary codebase and sensitive trade secrets could be compromised. This security policy highlights the growing tension between adopting cutting-edge developer productivity tools and managing information security perimeters in major technology enterprises. (source: https://www.reuters.com/world/china/alibaba-ban-claude-code-workplace-over-alleged-backdoor-risks-source-says-2026-07-03/)

03

Jamesob's guide to running SOTA LLMs locally

Developer jamesob has published a technical guide for running state-of-the-art large language models locally on consumer-grade hardware. The resource provides practical instructions for optimizing system resources, selecting models, and deploying quantized weights using software frameworks like llama.cpp. By configuring localized workflows, developers and researchers can bypass proprietary cloud APIs to maintain strict data privacy, reduce operational latency, and minimize computing costs. The guide features detail-oriented benchmarking advice and memory management configurations to help users achieve high-performance natural language processing on personal systems. (source: https://github.com/jamesob/local-llm)

04

Show HN: Mcpsnoop – Wireshark for MCP (transparent proxy and live TUI)

Developer kerlenton has released Mcpsnoop, an open-source debugging tool acting as a virtual Wireshark for the Model Context Protocol (MCP). Operating as a transparent proxy between MCP clients and servers, the utility intercepts live traffic and exposes message flows via a live Terminal User Interface (TUI). It allows developers to monitor request and response payloads, track schema validations, and audit communication pipelines in real-time. By providing visibility into interactions between LLMs and local tools, Mcpsnoop reduces debugging friction for complex, agentic AI architectures. (source: https://github.com/kerlenton/mcpsnoop)

05

60% Fable cost cut by converting code to images and having the model OCR it

Teamchong has introduced an unconventional cost-optimization method that reduces processing fees for multimodal large language models by 60 percent. The technique involves converting raw source code files into images and passing them to a vision-enabled model for optical character recognition (OCR) and analysis. By exploiting the pricing disparity between vision input tokens and standard text input tokens within current commercial LLM APIs, the project offers a new strategy for developers to build highly cost-effective coding tools. (source: https://github.com/teamchong/pxpipe)

06

Memorizing session transcripts isn't useful

An industry analysis argues that training large language models to memorize exact user-system interaction transcripts does not lead to generalization in agentic workflows. Instead of relying on superficial pattern matching of historical logs, developers of AI agents must build architectures that model underlying logic, system states, and dynamic transitions. This approach ensures that autonomous agents can solve novel problems and handle unexpected edge cases in live production environments, rather than failing when facing scenarios that deviate from memorized transcripts. (source: https://12gramsofcarbon.com/p/agentics-memorizing-session-transcripts)

07

Anatomy of Persistent Memory's 3 Layers: Comparing ContextNest, Mem0 and Zep

PromptOwl has published a comparative evaluation of persistent memory architectures for autonomous agents, focusing on ContextNest, Mem0, and Zep. The analysis outlines the three layers of persistent memory needed to help agents retain state, learn from continuous user interactions, and maintain coherent long-term context. It details how each software framework handles data storage, retrieval efficiency, and integration complexity. The technical comparison highlights the performance, scalability, and engineering trade-offs required to build persistent cognitive memory systems in agentic software. (source: https://promptowl.ai/resources/persistent-memory-ai-agents/)

08

Ask HN: Is anyone experimenting with different ways of using LLMs for coding?

A discussion on Hacker News addresses the limitations of current prompt-response workflows in LLM-assisted programming tools like Claude Code and Codex. Developers are criticizing the chat-based interface because its iterative loop interrupts the cognitive flow state required for software engineering. The thread explores alternative paradigms, such as non-blocking tab-completion models or streaming background processors, that could integrate code generation more smoothly into local development environments without requiring persistent manual prompting. (source: https://news.ycombinator.com/item?id=48771515)

Twitter

8 stories
01

Google Gemini Omni Flash Secures Top Spot On Video Arena Benchmark

Google DeepMind's Gemini Omni Flash model has secured the first overall position on the Video Arena benchmark. Achieving an Elo rating of 1404, the model demonstrates leading-edge performance in complex video understanding and processing capabilities. This milestone highlights Google's ongoing advancements in multimodal artificial intelligence, establishing a new standard for high-fidelity visual interpretation within large multimodal model evaluation frameworks. (source: https://x.com/ZoubinGhahrama1/status/2072922008166215981)

02

Sakana AI Introduces Multi-Agent Coordination via Sheaf-ADMM at ICML 2026

Sakana AI has published a research paper titled "Learning Multi-Agent Coordination via Sheaf-ADMM," which is scheduled for presentation at the ICML 2026 conference. The paper investigates novel frameworks for managing multi-agent systems, focusing on coordination mechanisms through the lens of Sheaf-ADMM mathematical techniques. The research aims to improve efficiency and stability in decentralized multi-agent environments, facilitating coordination in complex, large-scale operational settings. (source: https://x.com/hardmaru/status/2073038780089659642)

03

Analyzing Test-Time Compute Budgets in Frontier AI Model Evaluations

The AI Security Institute has published an investigation analyzing the impact of test-time compute budgets on frontier AI model evaluations. The study demonstrates that variable compute allocations significantly influence performance, reasoning accuracy, and overall capabilities of state-of-the-art systems during standard auditing. The findings suggest that ignoring test-time compute variability can skew performance metrics, underscoring the need for robust benchmarks that account for these computational budgets. (source: https://x.com/polynoamial/status/2072909389389021484)

04

Kling AI Introduces One Click Video Generation And Building Features

Kling AI has introduced a new video generation capability enabling users to generate high-quality architectural or cinematic visuals with a single click. By automating complex visual workflows, the platform streamlines generative video synthesis to help both professional and casual creators produce realistic, high-fidelity structural scenes. This update represents a secondary release in a series of creative potential showcases from the platform. (source: https://x.com/Kling_ai/status/2073059186167144859)

05

Google DeepMind Introduces COrigami for End-to-End Co-Design Pipeline

Google DeepMind's Discovery team has officially introduced COrigami, a specialized end-to-end pipeline designed to support co-design processes. COrigami aims to automate complex workflows and integrate systems design to enhance overall efficiency in research-intensive tasks. This release outlines Google DeepMind's ongoing efforts to integrate machine learning into collaborative development environments, potentially establishing a standardized pipeline for engineering and system-level discovery. (source: https://x.com/GoogleDeepMind/status/2073027851910050231)

06

Why Falling Token Costs Are Not Reducing Enterprise AI Expenses

Sara Hooker highlights an emerging paradox where dropping token costs for AI inference fail to lower enterprise AI expenses due to increased operational utilization. As organizations deploy complex autonomous agents and multi-step agentic workflows, the absolute consumption of compute resources offsets token pricing efficiencies. The analysis suggests enterprise success depends on improving agent comprehension and output quality rather than relying solely on model commoditization. (source: https://x.com/sarahookr/status/2073109791715811439)

07

Sam Altman's Narrative Strategy Regarding OpenAI's Ownership Stake

Yann LeCun and other industry observers are evaluating the narrative strategy of OpenAI CEO Sam Altman, specifically his proposal to allocate a five percent ownership stake in OpenAI to the United States. Observers are questioning whether this positioning is designed to influence federal regulatory sentiment, secure strategic domestic partnerships, or stabilize public perception. This development highlights the growing tension and debate around regulatory capture, AI governance, and corporate restructuring. (source: https://x.com/ylecun/status/2073038546558918735)

08

Runway Highlights Seven Years of Robust Research Infrastructure

Runway's engineering leadership has published an overview of the company's seven-year development of specialized research infrastructure and custom tooling. This backend infrastructure serves as the primary system enabling their engineering team to build, train, and deploy advanced generative video and creative models while supporting massive inference volumes. The technical review highlights the importance of highly optimized computing environments in maintaining a product pipeline for generative media. (source: https://x.com/c_valenzuelab/status/2073041661437804830)

huggingface

8 stories
01

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Researchers proposed Program-as-Weights (PAW), a novel programming paradigm designed to compile fuzzy functions from natural-language specifications into locally-executable neural artifacts. PAW uses a 4B parameter compiler trained on the newly released 10-million-example FuzzyBench dataset to output parameter-efficient adapters for a lightweight interpreter. In evaluation, a 0.6B Qwen3 interpreter executing these PAW programs matched the performance of direct prompting on a larger Qwen3-32B model. It achieved this while requiring approximately 1/50th of the inference memory and running at 30 tokens per second on a MacBook M3, successfully executing functions locally and offline. (source: https://huggingface.co/papers/2607.02512)

02

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Researchers introduced EvoPolicyGym, a standardized benchmark consisting of 16 compact interactive reinforcement learning environments to evaluate Autonomous Policy Evolution. In this setup, an AI agent iteratively refines an executable policy system under a fixed interaction budget. GPT-5.5 achieved the highest aggregate rank score and top-two performance across all 16 environments. The benchmark provides trajectory-level diagnostics that analyze how agents allocate their budget and convert environmental feedback into parametric tuning, showing that success relies on discovering task-appropriate mechanisms under bounded feedback. (source: https://huggingface.co/papers/2607.02440)

03

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Researchers developed AgenticSTS, a bounded-memory testbed and methodology designed for studying long-horizon LLM agents. Built using the closed-rule stochastic deck-building game Slay the Spire 2, the testbed requires hundreds of tactical and strategic decisions per run. Unlike conventional agents that append entire transcripts to prompts, AgenticSTS restricts decisions to fresh messages assembled by typed retrieval to stay bounded. In evaluations, a fixed-A0 ablation of the memory system showed that enabling strategic skills increased game win rates from 3/10 to 6/10. The release includes 298 completed trajectories, prompt records, and analysis scripts. (source: https://huggingface.co/papers/2607.02255)

04

AgenticDataBench: A Comprehensive Benchmark for Data Agents

To address the lack of rigorous testing environments for data science automation, researchers proposed AgenticDataBench, a benchmark designed to evaluate LLM-based data agents. The suite includes realistic tasks with fine-grained ground-truth labels across 15 vertical domains, including five real-world B2B use cases from a leading fintech company. Tasks are structured around data science skills and operational patterns extracted from Stack Overflow. The researchers introduced a systematic LLM-based task generation approach to synthesize workflows for domains lacking real-world datasets, evaluating state-of-the-art data agents to provide detailed skill-level insights. (source: https://huggingface.co/papers/2607.01647)

05

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

Researchers presented WorldDirector, a controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. The framework decouples semantic motion orchestration from visual generation, leveraging a large language model to coordinate 3D trajectories with camera movements. These trajectories act as control signals for video generation. This decoupled design maintains physical logic and visual consistency, allowing dynamic entities to preserve their exact appearance even after long periods of being out of view. Experimental results show WorldDirector supports complex, extended video synthesis with highly controllable object memory. (source: https://huggingface.co/papers/2607.02517)

06

AutoMem: Automated Learning of Memory as a Cognitive Skill

Researchers introduced AutoMem, a framework that automates the optimization of LLM memory management by treating memory as a trainable cognitive skill. AutoMem employs a two-loop optimization process: the first loop uses a teacher LLM to revise memory structures, prompts, and file schemas, while the second loop trains the agent's model directly on high-quality memory decisions. Evaluated on three procedurally generated long-horizon games (Crafter, MiniHack, and NetHack), optimizing memory alone improved the base agent's performance by 2x to 4x. This optimization allowed a 32B open-weight model to achieve performance competitive with Claude 4.5 and Gemini 3.1 Pro Thinking. (source: https://huggingface.co/papers/2607.01224)

07

PACE: A Proxy for Agentic Capability Evaluation

To reduce the high cost of agentic evaluations on benchmarks like SWE-Bench and GAIA, researchers introduced PACE, a framework that constructs cheap, compact proxy benchmarks. PACE selects a small subset of atomic, non-agentic evaluation instances and fits a regression mapping these atomic scores to predicted agentic performance. Applying this framework yielded PACE-Bench. Across 14 models, PACE-Bench predicted agentic scores with a cross-validation mean absolute error under 4%, a Spearman correlation above 0.80, and a pairwise model-ranking accuracy of approximately 85%. It achieved these results while using less than 1% of the typical agentic evaluation cost. (source: https://huggingface.co/papers/2607.02032)

08

Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Researchers proposed Task-Agnostic Pretraining (TAP), a two-stage training framework for Vision-Language-Action (VLA) models that separates physical motion execution from language alignment. TAP pretrains motor priors on cheap, unlabeled robot interaction datasets using a self-supervised Inverse Dynamics objective before aligning the priors to language instructions in a second stage. In SIMPLER benchmark evaluations, TAP matched the performance of models trained on over 1 million expert trajectories while using significantly less labeled data, yielding a 10% absolute gain over behavioral cloning. On physical WidowX robots, TAP maintained 25% success under camera perturbations. (source: https://huggingface.co/papers/2607.02466)