NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-20DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

Shares Safety Frameworks for Long-Horizon Models

OpenAI has published key lessons and safety frameworks derived from deploying long-running, long-horizon AI models. The report details the unique alignment challenges and agentic failures that emerge when AI systems execute complex, multi-step tasks autonomously over extended periods. To mitigate these risks before broader deployment, OpenAI utilized iterative deployment to identify previously unobserved safety risks in real-time environments, enabling the development of improved evaluation structures and alignment guardrails to transition safely from short-prompt interfaces to autonomous agents. (source: https://openai.com/index/safety-alignment-long-horizon-models)

Hacker News

8 stories
01

Safety and alignment in an era of long-horizon models

OpenAI published an analysis detailing new safety and alignment frameworks designed specifically for long-horizon AI models. As modern AI models transition to executing multi-step, autonomous tasks over extended durations, standard alignment and reinforcement learning methods must evolve to prevent goal drift and unintended behaviors. The publication outlines evaluation methods to measure model actions across prolonged sequences, providing systematic guardrails for highly autonomous agent deployments. These developments aim to ensure control and safety in complex, real-world agentic workflows. (source: https://openai.com/index/safety-alignment-long-horizon-models/)

02

Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling

An industry analysis explores the shifting economics of frontier artificial intelligence development following the releases of Alibaba's Qwen 3.8 and Moonshot AI's Kimi K3. The report highlights how these highly capable Chinese models are challenging Western dominance by narrowing performance gaps in coding, reasoning, and multilingual processing. The analysis argues that the rapid commoditization of inference and training efficiency poses severe challenges to Anthropic's capital-intensive business model, signaling broader strategic changes across the global AI ecosystem. This development was also discussed under the context of Kimi K3's open-weights escalation. (source: https://www.emergingtrajectories.com/lh/frontier-lab-economics/)

03

Exploit brokers pay $500k for WordPress RCEs. I found one with GPT5.6 and $25

A cybersecurity researcher utilized a large language model referred to as GPT-5.6 to discover a critical Remote Code Execution vulnerability in WordPress for a total cost of twenty-five dollars. Traditionally, finding such high-value security flaws required extensive manual auditing and domain expertise, commanding payouts of up to $500,000 from exploit brokers. This successful experiment demonstrates how autonomous language models can automate static application security testing at a fraction of standard operational costs, introducing dual-use risks where both security auditors and malicious actors can systematically scan codebases for zero-day exploits. (source: https://slcyber.io/research-center/exploit-brokers-pay-500000-for-a-wordpress-rce-i-found-one-with-gpt5-6/)

04

Agent swarms and the new model economics

Cursor published an article analyzing the emergence of autonomous agent swarms and their disruptive effect on the economic landscape of large language models. The industry is transitioning from isolated, single-query interactions to collaborative, multi-agent networks that execute complex workflows. This structural shift is significantly decreasing the marginal cost of intelligence, altering token consumption patterns, and demanding high-throughput, low-latency inference architectures. The publication highlights how model providers are adapting their pricing structures and systems to optimize for continuous, high-volume agentic calculations. (source: https://cursor.com/blog/agent-swarm-model-economics)

05

Controlling Reasoning Effort in LLMs

Sebastian Raschka analyzed technical methodologies for managing latency and cost in large language models by directly controlling their reasoning effort. As models are deployed for complex multi-step tasks, balancing resource consumption with logical accuracy becomes critical. The article details optimization techniques such as structured generation limits, adaptive inference steps, and targeted system prompts to dynamically tune generation length and thinking depth. The paper demonstrates how these granular controls allow developers to match computational expenditures with specific application needs in production environments. (source: https://magazine.sebastianraschka.com/p/controlling-reasoning-effort-in-llms)

06

An Empirical Study: AI Agent Rules Need Context and Layered Enforcement

An empirical study by Eunomia explores the limitations of prompt-based safety rules and proposes a layered runtime policy enforcement framework for AI agents. Because static system prompts are highly vulnerable to evasion during complex or adversarial execution scenarios, the researchers developed a multi-tiered safety model. This approach bridges abstract semantic intent with kernel-level execution tracking by utilizing eBPF (Extended Berkeley Packet Filter) technology to enforce hard system-level constraints. Decoupling policy enforcement from the reasoning layer of autonomous agents significantly improves security and reliability. (source: https://eunomia.dev/blog/2026/07/15/ebpf-ai-agent-policy-enforcement/)

07

Inertia-1: An Open Exploration to a Unified Motion Foundation Model

Yang AI Lab announced Inertia-1, an open-source motion foundation model designed to unify movement synthesis across human, quadruped, and robotic configurations. Unlike historical pipelines that required separate architectures for different skeletal structures, Inertia-1 processes physical motion data as a unified sequence learning problem. By enabling cross-entity motion transfer and generalization, the model aims to scale physical movement intelligence in a manner similar to large language models. This research represents a step forward in creating unified, task-agnostic controllers for robotics and advanced character animation. (source: https://yang-ai-lab.github.io/Inertia-1/)

08

Nativ: Run frontier open models locally on your Mac

A developer has released Nativ, a macOS application optimized to run frontier open-source large language models locally. By targeting the neural engines and hardware architecture of Apple Silicon, the platform provides a private, offline environment to execute advanced generative AI models without cloud dependencies or subscription fees. The tool simplifies downloading, managing, and querying complex neural networks directly on consumer devices, supporting the ongoing democratization and offline execution of open-weights models. (source: https://blaizzy.github.io/nativ/)

Twitter

7 stories
01

Runway Launches Act-One For Expressive Character Performance Animation

Runway has launched Act-One, a new generative video synthesis tool that generates highly expressive character performances from only a video source and a character image. This technology maps human facial expressions, head movements, and emotions directly onto digital characters without requiring traditional motion capture hardware. Designed to preserve character consistency and emotional nuance, the browser-based tool allows creators, filmmakers, and animators to build high-fidelity animations. The launch represents a significant progression in accessible, AI-driven visual effects production. (source: https://x.com/c_valenzuelab/status/2079183351374864793)

02

Anthropic Offers Research Grants For AI-Driven Rare Disease Cures

Anthropic has introduced a scientific research initiative offering grants of up to $50,000 in Claude usage credits to researchers working to cure rare diseases. This program, under the AI for Science umbrella, aims to leverage the advanced capabilities of the Claude large language model to accelerate scientific discovery and clinical breakthroughs. By offering computational credits, Anthropic seeks to help scientific teams integrate generative artificial intelligence into complex biomedical workflows and data analysis, targeting rare conditions that are often under-resourced in traditional settings. (source: https://x.com/AnthropicAI/status/2079256626771665098)

03

Baidu Releases Unlimited-OCR Model for Long-Form Document Processing

Baidu has open-sourced Unlimited-OCR, a 3-billion-parameter optical character recognition model engineered to process documents up to 40 pages in a single inference shot. The model handles complex layouts and extended texts, addressing a common failure point for traditional OCR systems. By making the code publicly available, Baidu aims to assist developers and enterprises with high-volume document ingestion across specialized fields like finance, law, and academic research, streamlining complex workflows with long-form automated text extraction. (source: https://x.com/ylecun/status/2079082115640049716)

04

Addressing Long-Running Model Safety And Persistence Risks

An OpenAI safety lead has shared insights into research investigating the safety implications of long-running, open-ended artificial intelligence systems. These persistent, agentic models present unique security and operational risks that are frequently missed by standard, short-horizon evaluation frameworks. The ongoing study focuses on developing new testing protocols and hazard mitigation strategies tailored specifically for models that operate autonomously over extended periods in complex settings. These efforts are intended to improve industry risk management frameworks for agentic systems. (source: https://x.com/polynoamial/status/2079260550895382965)

05

Kling AI Introduces Advanced Facial Transformation Capabilities

Kling AI has released new facial transformation capabilities for its generative video platform, focusing on precise character modification and digital synthesis. The update introduces advanced features allowing creators to edit and manipulate facial expressions and structural features of digital characters with high fidelity. This release forms part of Kling AI's broader initiative to offer granular control in its generative AI video pipeline, catering to animators and film professionals seeking professional-grade characters without complex rendering software. (source: https://x.com/Kling_ai/status/2079219783690498225)

06

Claude Code Troubleshooting And Service Stability Update

The Claude development team has issued an engineering service update regarding performance and connectivity issues discovered in Claude Code. To fix backend synchronization errors and restore full functionality, developers are instructed to restart their local installations. The immediate fix is part of ongoing efforts to stabilize Anthropic's terminal-based AI agentic coding tool, ensuring local client instances align with the service's current server-side infrastructure updates. (source: https://x.com/ClaudeDevs/status/2079111020308779394)

07

Google Enhances Gemini Capabilities With New DeepMind Model Integration

Google has rolled out an update integrating research breakthroughs from its DeepMind division directly into the Gemini model ecosystem. This optimization is aimed at elevating capabilities in complex coding, logical reasoning, and multimodal understanding. By integrating architectural refinements, the release is designed to improve inference speeds and deliver more accurate outputs. The release highlights Google's ongoing strategy of deploying state-of-the-art computational model upgrades to its consumer and enterprise products. (source: https://x.com/Google/status/2079231381507645540)

huggingface

8 stories
01

Loop the Loopies!

Researchers have released Loopie, a highly capable looped Transformer series incorporating Mixture-of-Experts (MoE). The release consists of a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Loopie addresses the historical compute-efficiency limitations of looped models, outperforming standard Transformer baselines trained under identical compute budgets. A novel post-training pipeline provides the models with advanced reasoning capabilities, enabling Loopie to achieve gold-medal performance without external tools at the 2025 International Mathematical Olympiad (IMO) and International Physics Olympiad (IPhO). (source: https://huggingface.co/papers/2607.16051)

02

Understanding Reasoning from Pretraining to Post-Training

Researchers investigated the relationship between pretraining choices and reinforcement learning (RL) post-training dynamics using chess and mathematics as controlled testbeds. By pretraining models from 5M to 1B parameters, the study demonstrates that post-RL performance is highly predictable from pretraining loss, with the slope of RL reward curves improving linearly with the volume of pretraining tokens. The analysis reveals that RL does not merely sharpen existing policies; it amplifies correct moves already preferred on easy tasks while surfacing entirely new correct moves on complex tasks. (source: https://huggingface.co/papers/2607.16097)

03

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

Researchers introduced S1-Omni, a unified multimodal reasoning model designed for scientific understanding, prediction, and generation. S1-Omni maps diverse scientific data modalities, including CIF, SMILES, protein sequences, spectra, and images, into a shared representation space aligned with scientific laws. Trained on the S1-Omni-Corpus of millions of reasoning samples across 200 tasks, the model supports applications like property prediction, structure prediction, and molecular generation. S1-Omni outperforms GPT-5.5 and Gemini-3.1-Pro on most of over 60 scientific benchmarks. (source: https://huggingface.co/papers/2607.15686)

04

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Researchers developed Agon, a competitive reinforcement learning framework that implements implicit rival grading of reasoning using two competing models. Instead of relying on manual process labels or external reward models, the system rewards each model for out-solving its rival, who acts as both solver and grader in alternating rounds. This setup establishes a progressive difficulty curve. On the DeepMath benchmark with Qwen3, Agon doubled the standard GRPO pass@1 rate, demonstrating robust performance scaling across competitive programming and multiple model families. (source: https://huggingface.co/papers/2607.07690)

05

RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM

Developers released RAGU, an open-source GraphRAG engine that separates knowledge graph extraction from consolidation to improve retrieval quality. The system uses a two-stage typed extraction process, DBSCAN deduplication, and Leiden community detection. RAGU is powered by Meno-Lite-0.1, an optimized 7B model that outperforms Qwen2.5-32B on knowledge-graph construction tasks by 12.5% in relative harmonic mean. RAGU achieves an evidence recall of up to 0.84 on GraphRAG-Bench (Medical) and is installable under an MIT license. (source: https://huggingface.co/papers/2607.11683)

06

Cura 1T: Specialized Model for Agentic Healthcare

Researchers introduced Cura 1T, a healthcare-specialized LLM designed for clinical consultation, multi-modal reasoning, interactive diagnostics, and electronic health record (EHR) tool integration. Cura 1T is trained via a human-gated self-evolution loop where a training agent plans capabilities, executes targeted training, and updates data mixtures iteratively based on performance failures. Evaluated on healthcare benchmarks, Cura 1T achieved top-tier performance relative to major frontier baselines while maintaining strong generalization on out-of-domain reasoning and agentic tasks. (source: https://huggingface.co/papers/2607.15314)

07

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization

Researchers proposed Contrastive Policy Optimization (CPO), a correctness-aware advantage shaping method designed for reinforcement learning with verifiable rewards (RLVR). CPO utilizes token-level contrastive disagreement between reference-guided and vanilla generation distributions to distinguish useful uncertainty from detrimental model confusion. Theoretical and empirical evaluations demonstrate that this contrastive disagreement acts as a reliable correctness signal, addressing the zero-advantage problem and consistently outperforming entropy-based RLVR baselines on both in-domain and out-of-domain reasoning benchmarks. (source: https://huggingface.co/papers/2607.14614)

08

DSWorld: A Data Science World Model for Efficient Autonomous Agents

Researchers developed DSWorld, a framework establishing a Data Science World Model to accelerate autonomous data science agents. By predicting environment state transitions conditioned on operations, DSWorld replaces expensive trial-and-error executions with simulated predictions. It features structured state construction, cost-aware routing, and Reflective World Model Optimization. Under experimental testing, DSWorld accelerated reinforcement-learning-based agent training by 14 times and search-based inference by 3 to 6 times while outperforming top LLM baselines by 35.6% on transition prediction tasks. (source: https://huggingface.co/papers/2607.15901)