NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-30DEFAULT EDITION
This issue
—
All time
—

AI Blog

2 stories
01

Two API Settings Triple GPT-5.6 Scores on ARC-AGI-3

OpenAI announced that configuring two specific API settings tripled the performance of its GPT-5.6 model on the ARC-AGI-3 benchmark. The optimization parameters focused on retaining the model's internal reasoning capabilities and utilizing token compaction to significantly boost output efficiency. By utilizing these targeted API adjustments, the model achieved higher overall accuracy scores on complex reasoning tasks without requiring additional computational retraining. This development demonstrates that reasoning-oriented model performance can be substantially augmented through software-level parameter adjustments alone. (source: https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores)

02

Lowers GPT-5.6 Pricing for Luna and Terra Models

OpenAI announced price reductions for its GPT-5.6 model family, specifically targeting the Luna and Terra variants to make enterprise deployments more cost-effective. By optimizing model execution efficiency, OpenAI seeks to enable organizations to deploy complex conversational and analytical workflows at scale without incurring prohibitive costs. These cost reductions are designed to lower the financial barrier to entry for businesses integrating advanced capabilities into their operations, reinforcing the price-to-performance ratio of the GPT-5.6 family. (source: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6)

Hacker News

8 stories
01

Advancing the price-performance frontier with GPT‑5.6

OpenAI has announced the release of GPT-5.6, a new model focused on optimizing computational efficiency and logical reasoning performance. By enhancing the underlying transformer architecture, the model delivers faster inference speeds and improved accuracy at a significantly lower operational cost compared to previous generations. This release aims to provide developers with highly scalable natural language processing for complex reasoning tasks, code generation, and interactive workflows while reducing hosting expenses. An empirical experiment deploying this model as an autonomous business agent resulted in a $447 loss, highlighting ongoing safety and alignment challenges with fully autonomous agents. (source: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)

02

Gemini Robotics 2 brings whole body intelligence to robots

Google DeepMind has introduced Gemini Robotics 2, a physical artificial intelligence system that implements whole-body intelligence in robotic hardware. By integrating the multimodal reasoning of the Gemini foundation models, the system processes visual, tactile, and natural language inputs simultaneously instead of relying on isolated control loops. This unified approach enables smoother physical coordination and real-time decision-making in dynamic environments. The release demonstrates a methodology for bridging the gap between digital reasoning and physical robot execution using large multimodal models. (source: https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/)

03

Launch HN: Prized (YC S26) – Let non-engineer staff build secure internal tools

Prized, an enterprise software startup from the YC S26 cohort, has launched a platform enabling non-technical employees to construct secure, full-stack internal applications using natural language. The system integrates directly with corporate sign-in and database environments without requiring manual API configuration. To protect production resources from rogue agent actions, Prized isolates its execution sandbox at the network layer and manages credentials using scoped session tokens as opaque placeholders, swapping actual authentication details into headers via an egress proxy. (source: https://prized.dev)

04

Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it

Researchers have published a study demonstrating that knowledge distillation from politically censored large language models does not inherently transfer their alignment or censorship biases to student models. The study distilled DeepSeek V4 Flash into GPT-OSS-120B to evaluate performance on financial reasoning tasks, achieving a score of 83.61% on the FinanceReasoning benchmark. While the teacher model's responses to politically sensitive queries deviated significantly by 7 standard deviations from expected baselines, the distilled student model did not inherit these traits, instead aligning with its Western base model. (source: https://www.ctgt.ai/research/distillation-censorship-transfer)

05

Go LLM SDK for streaming, tool-calling AI backends (plus frontend React lib)

Grafana has released an open-source Go LLM SDK designed to facilitate the construction of streaming, tool-calling AI backends, accompanied by a React frontend library. The package simplifies integrating generative AI features into Go applications by providing structured streaming protocols and pre-built user interface elements. This release addresses common synchronization challenges during real-time token deliveries and dynamic client-side state transitions, offering developers a cohesive framework to build responsive, tool-calling agent systems within the web development ecosystem. (source: https://github.com/grafana/ai-sdk)

06

Show HN: Claude-account – switch Claude Code accounts without logging in again

Developer Hamza Rehman has launched claude-account, an open-source command-line utility that optimizes multi-tenant developer workflows for Claude Code. The tool solves authentication friction by enabling engineers to securely save credentials for separate professional and personal Anthropic accounts. Users can seamlessly toggle between saved identities using CLI commands such as add and use, which updates configurations in real time to prevent workflow disruption during agentic software development and automated codebase refactoring tasks. (source: https://github.com/hamzarehmandeveloper/claude-account)

07

Show HN: Noisegate – a differential-privacy gateway for untrusted AI agents

Developer Yash Mahajan has released Noisegate, an open-source differential-privacy gateway designed to secure sensitive data interactions with untrusted artificial intelligence agents and third-party large language models. Positioned between local enterprise data sources and external LLM services, the software applies mathematical differential-privacy algorithms to user queries. This system filters and anonymizes contextual parameters, masking individual identity-revealing patterns while preserving the semantic and analytical utility of the payload to prevent proprietary data leaks. (source: https://github.com/yashmahajan10/llm-differential-privacy-gateway)

08

You can't solve computer use by ignoring the interface

Steelman Labs has published an analysis critiquing modern methodologies for training computer-use AI agents. The piece argues that bypassing human-centric graphical user interfaces or treating application layers as black boxes limits agent adaptability. To achieve generalized computer control, AI systems must develop robust visual reasoning and understand layout designs intended for human users. The author asserts that progress relies on constructing real-time feedback loops that embrace, rather than abstract, visual complexity. (source: https://steelmanlabs.com/blog/computer-use-is-far-from-solved)

Twitter

8 stories
01

OpenAI Announces Significant Price Reductions And API Updates For GPT-5.6 Models

OpenAI has announced substantial price reductions across its GPT-5.6 model lineup to improve cost efficiency for developers. The flagship GPT-5.6 Luna model received an 80% price cut, bringing input tokens to $0.20 and output tokens to $1.20 per million. Additionally, the GPT-5.6 Terra variant has been reduced by 20% to $2 for input and $12 for output per million tokens. OpenAI also introduced a performance-oriented Fast mode for the GPT-5.6 Sol API, providing up to 2.5 times the processing speed at double the standard cost while maintaining the base model's intelligence. (source: https://x.com/sama/status/2082880720989532597)

02

Google DeepMind Unveils Gemini Robotics 2 for Physical AI Advancements

Google DeepMind has launched Gemini Robotics 2, a unified brain system designed to provide physical AI platforms with full-body intelligence and unified control. By integrating advanced vision, language, and multimodal reasoning capabilities directly into robotic architectures, the system aims to close the gap between high-level cognitive reasoning and complex real-world physical manipulation tasks. The project represents a significant step toward general-purpose physical agents that can perceive, reason, and act adaptively within dynamic corporate and industrial environments. (source: https://x.com/Google/status/2082844315420602827)

03

Google Introduces Gemini 3.6 Flash And 3.5 Flash-Lite Models

Google has expanded its generative model lineup with the release of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Designed specifically for developers and enterprise builders, these models are optimized to deliver high-level processing intelligence while balancing operational speed and scaling cost-efficiency. This model release addresses developer demand for high-throughput API workflows without sacrificing overall output quality. The rollout represents Google's strategic focus on optimizing specialized models for highly scalable generative deployments. (source: https://x.com/Google/status/2082895574294987211)

04

Google Introduces Gemini Robotics ER 2 Embodied Reasoning Model

Google has announced Gemini Robotics ER 2, an embodied reasoning model designed to integrate physical hardware platforms with advanced multimodal systems. Built on the core Gemini architecture, ER 2 enhances the spatial awareness, physical reasoning, and real-time decision-making capabilities of robotic platforms. The release highlights deep learning advances designed to assist robotic systems in executing highly complex manipulation and coordination tasks in the physical world, representing a strategic step forward in embodied artificial intelligence. (source: https://x.com/Google/status/2082850448193589435)

05

Hailuo Launches New H3 Model With Advanced Video Generation Capabilities

Hailuo has launched its new H3 video generation model, demonstrating its technical capabilities in high-fidelity generative video synthesis. Initial evaluations and user tests indicate that the H3 model provides precise prompt control, enabling detailed adherence to cinematographic instructions, specific camera movements, smooth character motion, and visual consistency. This architecture allows content creators to generate complex text and diorama-style animations from natural language prompts, marking a significant development in competitive professional-grade video synthesis tools. (source: https://x.com/Hailuo_AI/status/2082885931254821088)

06

Sam Altman Previews New Models Accelerating Scientific Discovery

OpenAI CEO Sam Altman has announced a shift in development strategy, previewing upcoming AI models tailored specifically to accelerate scientific discovery. Emphasizing a collaborative approach, Altman stated that OpenAI's priority is to empower the global scientific community with advanced models rather than solving complex physical challenges internally. This strategy aims to democratize access to advanced computational reasoning tools and accelerate breakthrough scientific research in laboratories and research institutions worldwide. (source: https://x.com/sama/status/2082628413769003269)

07

Guidelines for Standardized Harness Usage in ARC-AGI-3 Benchmarking

Francois Chollet has released clarification guidelines for testing standards in the ARC-AGI-3 benchmark, drawing a clear distinction between acceptable general-purpose API configurations and prohibited custom harness solutions that internalize test knowledge. The announcement aims to resolve measurement discrepancies by ensuring all runtime details and computational costs are transparently disclosed. This standardization effort addresses the issue of custom-engineered shortcuts, establishing a reproducible and fair baseline comparison framework for testing advanced reasoning models. (source: https://x.com/fchollet/status/2082732210436575669)

08

Kling AI Enables Expressive Performances in L’Ultimo Uomo Reale Film

Kling AI was utilized during the cinematic production of L’Ultimo Uomo Reale to generate emotional character performances. The project leverages the technical architecture of Kling AI to maintain character identity consistency and preserve fine visual details across multiple frames. By resolving common temporal artifacts and structural stabilization challenges in generative video synthesis, this implementation demonstrates the practical utility of modern AI video generators in professional filmmakers' storytelling pipelines. (source: https://x.com/Kling_ai/status/2082843670005653563)

huggingface

8 stories
01

GPT-Red: Automated Red Teaming via Self-Play at Scale

Researchers have introduced GPT-Red, an automated red-teaming agent trained to discover novel prompt injection attacks against frontier LLMs. Utilizing a scalable self-play algorithm, the model is tasked with attacking a diverse population of simultaneously-trained defender agents. GPT-Red was used to adversarially train GPT-5.6, making it their most robust model to prompt injections to date. The safety training run represents one of the largest LLM safety post-training efforts documented. GPT-Red successfully breaks prior models up to GPT-5.5, outperforming human red-teamers and demonstrating generalization to held-out environments and defender models. (source: https://huggingface.co/papers/2607.26115)

02

Can AI agents conduct open-ended AI research? Early evidence from two case studies

A study evaluated whether autonomous AI agents can conduct open-ended AI research by executing "shadow evaluations" on two unpublished NeurIPS 2026 papers. Frontier agents were given six days and thousands of dollars in compute to solve the central research questions, with the original authors grading the outputs. Although the agents successfully completed the engineering work without human assistance, they failed to make substantial progress on the core research questions. The authors identified five recurring failure modes, including poor judgment regarding publishable bars, uncreative responses to research design limits, ineffective backtracking, poor resource awareness, and instruction drift. (source: https://huggingface.co/papers/2607.27191)

03

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Researchers introduced MindForge, an automated pipeline that converts open-source command-line programs into source-free environments to train coding agents in from-scratch software engineering. Utilizing GLM-5.2 as a teacher agent, MindForge curated high-quality program synthesis trajectories to fine-tune Qwen3.6-27B. The fine-tuned model improved its average test pass rate on ProgramBench from 37.98% to 49.51%, matching the performance of much larger frontier models. It also achieved significant absolute gains on seven unseen software engineering benchmarks, including 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, and 5.04 on SWE-bench Verified. (source: https://huggingface.co/papers/2607.27146)

04

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Researchers have released OmegaUse-OfficeVal, a new benchmark designed to evaluate large language model (LLM) agents on long-horizon office tasks. The benchmark features 100 realistic tasks derived from actual practitioner requests, averaging 2.32 hours of human labor per task. A key feature of the dataset is the pairing of tasks with human labor time and price proxies to enable direct cost and efficiency comparisons with LLM inference. Code-based verifiers built on fine-grained rubrics support stable evaluations. Results show that while current frontier LLMs are faster and cheaper than humans, they do not yet match human deliverable quality. (source: https://huggingface.co/papers/2607.27155)

05

DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation

Researchers proposed DistillAlign, a framework to coordinate mode covering and mode seeking in autoregressive video distillation. Standard multi-stage pipelines decouple initialization from Distribution Matching Distillation (DMD), leading to coverage loss. DistillAlign addresses this with a joint distillation objective that combines DMD's mode-seeking capabilities with a Consistency Distillation-based mode-covering constraint. Using a new distributional evaluation protocol to measure latent precision and coverage, the authors show DistillAlign improves generation quality and diversity, with a Wan-1.3B DMD teacher model outperforming baselines distilled from a larger Wan-14B model. (source: https://huggingface.co/papers/2607.26811)

06

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Researchers introduced StealthBench, a benchmark designed to evaluate operational stealth in autonomous offensive-security agents. The framework features 11 verified operational security (OPSEC) incidents expanded into 14 dockerized scenarios across six dimensions. Evaluation is carried out via a three-model LLM judge panel with majority-vote aggregation to measure metrics including safe success rate, Stealth@Solve, and reckless solve rate. Across tested model families, none exceeded a 54% safe success rate, highlighting systematic OPSEC failures in autonomous offensive agents. (source: https://huggingface.co/papers/2607.26314)

07

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Researchers introduced SecRespond, a benchmark for evaluating LLM agents on post-compromise incident response. Built across 10 cyber ranges from compromised cloud hosts, the benchmark spans 4 entry-point types, 21 ATT&CK techniques, and 5 operating systems. Using the OpenCode agent harness, agents analyze forensic disk snapshots alongside security alerts to generate intrusion reports and remediation plans. Evaluations of 23 frontier LLMs showed that while agents identify alert-exposed issues, they struggle with proactive disk investigations and comprehensive remediation, with no model achieving full success on any single range. (source: https://huggingface.co/papers/2607.26791)

08

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Researchers proposed CoRT, a token-level credit weighting method designed for rubric-conditioned Group Relative Policy Optimization (GRPO). Instead of training an auxiliary scoring model, CoRT utilizes counterfactual replay to rescore sampled responses under both rubric-conditioned and criteria-free prompts. The resulting tokenwise log-likelihood contrasts serve as normalized weights to redistribute the signed GRPO advantage across tokens. Across instruction-tuned models, CoRT improved on response-level GRPO by an average of 4.4 percentage points, matching the performance of learned token-level baselines while avoiding additional relevance-learning stages. (source: https://huggingface.co/papers/2607.25659)