NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-01DEFAULT EDITION
This issue
—
All time
—

AI Blog

2 stories
01

Redeploys Claude Fable 5 and Mythos 5 After Export Controls Lifted

Anthropic announced the restoration and global redeployment of its Claude Fable 5 and Mythos 5 models across its platforms, including Claude.ai, Claude Platform, Claude Code, and Claude Cowork. The decision follows the US government's lifting of temporary export controls originally imposed on June 12, which had led Anthropic to briefly restrict access globally. Collaborative reviews with partners including Amazon, Microsoft, and Google helped resolve the restrictions, and the group is now establishing a shared industry framework called Project Glasswing to address model jailbreaks and safety safeguards systematically. (source: https://www.anthropic.com/news/redeploying-fable-5)

02

GeneBench Pro Evaluates AI Performance in Genomics and Biology Research

OpenAI introduced GeneBench-Pro, a new evaluation benchmark designed to test artificial intelligence performance in genomics, biology, and scientific research. The benchmark uses complex, real-world datasets to evaluate capabilities in sensitive scientific domains such as DNA sequence design, functional prediction, and biological reasoning. According to complementary case studies, research institutions and biotechnology partners are deploying this suite to systematically measure model limitations and establish safety boundaries, ensuring more reliable and controlled deployment of AI in laboratory environments. (source: https://openai.com/index/introducing-genebench-pro)

Hacker News

6 stories
01

Claude Fable 5 Promotional Access

Anthropic has launched a promotional access program for Claude Fable 5, its newest high-performance large language model optimized for advanced reasoning, conversational tasks, and agentic workflows. This program allows selected developers and enterprises to evaluate the model's capabilities, run benchmark tests, and perform safety evaluations. The release has sparked broad interest across the tech community (supporting discussion: https://twitter.com/claudeai/status/2072402636813607381). (source: https://support.claude.com/en/articles/15424964-claude-fable-5-promotional-access)

02

ZCode: Claude Code from the Makers of GLM

The creators of the ChatGLM (GLM) series of large language models have launched ZCode, a specialized AI coding assistant and developer tool. Built on top of the GLM architecture, the tool provides software developers with deep code intelligence, multi-file context understanding, and agentic code generation features designed to compete with Anthropic's Claude Code. ZCode is designed to streamline codebase analysis, automated debugging, and complex system refactoring tasks while reducing cognitive load. (source: https://zcode.z.ai/cn)

03

Reduce GVisor Cold Starts with GPU Snapshotting

Cerebrium has introduced a new technique utilizing gVisor and memory snapshotting to reduce GPU cold starts in serverless environments. The method captures the running state of CUDA workloads directly from memory during initialization and restores it in a fraction of a second, bypassing the slow startup phases of loading models and weights. By integrating container virtualization with NVIDIA driver state restoration, this architecture makes real-time, on-demand GPU inference scaling practical and cost-effective. (source: https://cerebrium.ai/blog/reducing-gpu-cold-starts-with-memory-snapshots-restoring-cuda-workloads-in-second)

04

Launch HN: Parsewise (YC P25) – Reason Across Documents with an API

Parsewise, a startup from the YC P25 batch, has released an API designed to convert unstructured documents into schema-compliant JSON or CSV outputs. The platform addresses common limitations of large language models—including latency, cost, volume restrictions, and validation errors—by establishing reliable data ETL pipelines. Key features include preserving data lineage across multiple documents, real-time output validation, and mechanisms to involve business experts in defining schemas. (source: https://news.ycombinator.com/item?id=48746752)

05

Show HN: AnalystAIPack – 118 runnable agent skills for malware analysis and RE

Melted in Hex has released AnalystAIPack, a security toolkit containing 118 runnable skills designed for AI agents conducting malware analysis and reverse engineering. The pack bridges large language models and security operations, allowing automated agents to decompile code, inspect binary files, and discover vulnerabilities. This modular framework provides cybersecurity teams with developer-friendly tools to construct autonomous systems capable of threat detection. (source: https://meltedinhex.com/posts/analyst-ai-pack/)

06

Monetization Gateway

Cloudflare has introduced the Monetization Gateway, a network infrastructure solution that helps web publishers manage, license, and monetize content access requests from AI bots and scrapers. The service allows publishers to establish fine-grained access policies, execute microtransactions, and build licensing agreements with artificial intelligence developers. Operating at Cloudflare's network edge, the gateway helps secure creator intellectual property in the era of generative AI. (source: https://blog.cloudflare.com/monetization-gateway/)

Twitter

8 stories
01

Anthropic Reintroduces Claude Fable 5 With Enhanced Cybersecurity Safeguards

Anthropic has announced the global re-release of its Claude Fable 5 model with updated cybersecurity safeguards, following consultations with the United States government. To address potential misuse, the model incorporates a new suite of classifiers engineered to identify and block requests associated with cyber threats. While standard coding assistance remains supported, users may encounter highly sensitive classifiers that temporarily increase false-positive flags on harmless requests. To mitigate this, Anthropic introduced guidelines and a built-in /feedback command in Claude Code to facilitate manual model calibration. Paid plan subscribers are granted temporary access to Fable 5 up to 50% of their weekly usage limit through July 7. (source: https://x.com/AnthropicAI/status/2072163884430229756)

02

GeneBench-Pro Released To Test Expert-Level Computational Biology Analysis

OpenAI President Greg Brockman announced the release of GeneBench-Pro, a novel benchmark designed to evaluate how AI models handle complex, judgment-heavy tasks in computational biology. The benchmark simulates real-world scientific workflows that typically require 20 to 40 hours of intensive effort from human domain experts. Evaluation results highlight the performance of the GPT-5.6 Sol model, demonstrating substantial progress in solving sophisticated, domain-specific challenges that demand deep analytical reasoning rather than simple information retrieval. This evaluation framework marks a significant step in assessing the practical utility of large language models for high-stakes scientific discovery. (source: https://x.com/gdb/status/2072191801122038207)

03

AdaJEPA Model Integrates Adaptive Control and Model Predictive Capabilities

Yann LeCun and his research team have introduced AdaJEPA, a machine learning architecture designed to transform the LeWorld model into an adaptive system. By incorporating model-predictive control (MPC), the framework enables sophisticated action planning and decision-making in dynamic, physical environments. This integration represents a major shift toward autonomous agents capable of predicting future states and adjusting behavior using real-time sensorimotor feedback loops. The development represents an evolutionary step for joint-embedding predictive architectures, moving beyond static data processing to bridge high-level conceptual understanding and low-level physical control. (source: https://x.com/ylecun/status/2072337476748808365)

04

Runway Partners With Bertelsmann To Integrate Generative AI Technology

Runway has announced a strategic partnership with global media conglomerate Bertelsmann to deploy its advanced generative AI tools and creative suites across Bertelsmann’s international portfolio of businesses. The integration targets core entities including RTL Group, BMG, and Bertelsmann Marketing Services. By embedding generative video and media production tools directly into these professional workflows, the collaboration aims to streamline creative processes, enhance production efficiency, and expand content distribution. This enterprise-level adoption represents a notable milestone in utilizing generative machine learning within high-end, global creative industries. (source: https://x.com/runwayml/status/2072332384855335067)

05

ARC-AGI-3 Challenge Poses Significant Hurdle for Standard AI Architectures

Francois Chollet reported that the ARC-AGI-3 benchmark has emerged as a significant barrier for modern artificial intelligence systems, resisting resolution by conventional scaling laws and standard architectures. The benchmark demands unique cognitive problem-solving capabilities that differ sharply from routine pattern recognition and retrieval tasks. Researchers are focusing on why this benchmark remains highly resistant to contemporary training techniques. Solving ARC-AGI-3 is considered a crucial milestone for the research community as they attempt to move beyond static data processing and transition toward generalized reasoning capabilities in machine intelligence. (source: https://x.com/fchollet/status/2072337007549014194)

06

Kling AI Powered Film The Last Real Man Wins Honors At Cannes Lions 2026

The short film L'Ultimo Uomo Reale (The Last Real Man), directed by Sebastian Strasser and powered by Kling AI, has won both a Silver Lion and a Bronze Lion at the 2026 Cannes Lions International Festival of Creativity. The film was honored in the competitive Film - Consumer Goods category and the newly introduced Film Craft - AI Craft category. This achievement highlights the accelerating integration of advanced generative AI models into professional, high-end cinematography and advertising workflows, demonstrating their capability to compete at the highest tier of global creative industries. (source: https://x.com/Kling_ai/status/2072296943666016418)

07

Bloome Unveils Shared Workspace For Collaborative AI Agent Teams

Francois Chollet highlighted Bloome, a new collaboration platform that enables users to integrate multiple advanced language models, including Claude, ChatGPT, and Gemini, alongside human partners in a single shared workspace. The platform focuses on the deployment and orchestration of cross-agent feedback loops to optimize iterative, complex workflows. By allowing heterogeneous artificial intelligence systems to communicate directly with one another and with human collaborators, Bloome aims to address the growing technical need for multi-agent system coordination and structured team integration within professional development environments. (source: https://x.com/fchollet/status/2072155613925437769)

08

Introducing VISReg A Novel Approach For World Models And Self-Supervised Learning

Researchers have introduced VISReg, a new methodology designed to improve self-supervised learning systems and generative world models. VISReg addresses the persistent issue of representation collapse, which often undermines the performance of latent space architectures during training. By stabilizing training dynamics and enhancing feature extraction, the framework allows intelligent agents to learn more discriminative, meaningful internal representations. This development provides a practical tool for AI researchers focused on constructing robust and highly reliable predictive world models that can interpret complex, high-dimensional real-world data without relying on extensive supervised labels. (source: https://x.com/ylecun/status/2072111978693210419)

huggingface

8 stories
01

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

Researchers have introduced Reinforcement Learning with Metacognitive Feedback (RLMF), a paradigm that improves how Large Language Models (LLMs) assess and express their own capabilities. RLMF refines completion rankings during preference optimization using the quality of a model's self-judgments. Coupled with metacognitive data selection, this two-stage, decoupled approach calibrates a model's self-reported confidence scores before mapping them to natural, context-adaptable linguistic uncertainty. In evaluations, RLMF achieved state-of-the-art faithful calibration across diverse tasks, outperforming standard reinforcement learning by up to 63% while preserving overall model accuracy. (source: https://huggingface.co/papers/2606.32032)

02

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

Researchers have introduced SWE-Interact, a new benchmark designed to evaluate software engineering agents in realistic, multi-turn, interactive developer sessions. Unlike static benchmarks that provide complete instructions upfront, SWE-Interact employs a user simulator that starts with vague requirements and progressively reveals feedback and constraints based on the agent's work. Testing reveals a significant performance gap: while top-tier models like Opus 4.8 and GPT 5.5 solve approximately 50% of single-turn baseline tasks, their success rates drop to 25% on the interactive SWE-Interact tasks, highlighting challenges in handling ambiguity and over-agentic coding mistakes. (source: https://huggingface.co/papers/2606.30573)

03

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning

To address limitations in outcome-only reinforcement learning, researchers proposed TRIAGE, a role-typed credit assignment framework for agentic reinforcement learning. In standard GRPO, uniform final outcome signals often penalize useful exploration in failures or reward redundancies in successes. TRIAGE addresses this by using a structured judge to categorize action segments into progress, exploration, infrastructure, or regression, mapping these categories to bounded segment-level process rewards. Evaluated across ALFWorld, Search-QA, and WebShop, TRIAGE successfully increased task success rates, improved policy gradients, and reduced environment-facing steps in ALFWorld and WebShop by 10.4% and 14.8% respectively. (source: https://huggingface.co/papers/2606.32017)

04

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

Researchers have proposed Evolution Fine-Tuning (EFT), a mid-training paradigm that teaches Large Language Models to iteratively evolve solutions to complex optimization tasks instead of starting search scaffolds from scratch. The authors compiled the Finch Collection, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, to train models ranging from 2B to 9B parameters. Across 22 held-out tasks, EFT models demonstrated generalized capability, outperforming base counterparts by 10.22% on average and matching state-of-the-art results on mathematical and computational problems when combined with test-time reinforcement learning. (source: https://huggingface.co/papers/2606.29082)

05

Hierarchical Experimentalist Agents

To address the limitations of static knowledge in novel domains, researchers introduced Hierarchical Experimentalist Agents (HExA), an in-context self-improvement framework that learns from active experimentation. HExA designs, simulates, and refines experiments to compile a reusable library of composable skills without requiring external supervision. Evaluated on Interphyre, a physics-based simulation benchmark, HExA improved the success rate of Claude Sonnet 4.6 from 2% to 77% on challenging tasks, and achieved a 44% transfer success rate using only skills learned from easier levels, showing clear generalization. (source: https://huggingface.co/papers/2606.29315)

06

Xiaomi-GUI-0 Technical Report

Xiaomi has released a technical report on Xiaomi-GUI-0, a native multimodal GUI agent optimized for real mobile environments. Recognizing that simulated or offline benchmarks fail to model physical device variables like payment pop-ups and risk controls, the team implemented a real-device-dominant hybrid execution infrastructure. Grounded in a training dataset that spans high-frequency tasks and error-driven recovery trajectories, the model is refined using supervised fine-tuning and step-level and agentic reinforcement learning. Xiaomi-GUI-0 achieved a 72.0% success rate on the in-house RealMobile platform and 78.9% on AndroidWorld. (source: https://huggingface.co/papers/2606.31410)

07

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

To resolve conflicts when combining specialized capabilities during post-training, researchers proposed Multi-Teacher On-Policy Distillation (MOPD). MOPD trains multiple independent, specialized reinforcement learning teachers and then distills their expertise into a single student model during the student's own on-policy rollouts, eliminating exposure bias. Evaluated on Qwen3-30B-A3B, MOPD outperformed established baselines like Mix-RL and Parameter Merging by retaining nearly all original teacher capabilities. The framework has been successfully deployed in the post-training pipeline of the frontier-scale industrial model MiMo-V2-Flash. (source: https://huggingface.co/papers/2606.30406)

08

MemLearner: Learning to Query Context memory for Video World Models

To address scene inconsistency over long-duration generations, researchers developed MemLearner, a learning-based adaptive context query method for video world models. While existing systems rely on rule-based retrieval that fails during occlusions, MemLearner uses query tokens to bridge context and predicted tokens, tapping into pre-trained visual priors without training new architectures from scratch. Using a newly collected dataset of long videos with camera pose annotations, MemLearner demonstrated superior scene consistency and memory retention under complex occlusions and dynamic object movements compared to prior video world models. (source: https://huggingface.co/papers/2606.31734)