NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-28DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

Scientific Computing in the Age of Agentic AI

OpenAI published a field report detailing how researchers use AI coding agents to modernize scientific computing. These agentic AI systems automate complex data analysis pipelines, accelerate software development, and update legacy scientific codebases in disciplines like genomics. By integrating agents directly into computational workflows, laboratories are systematically overcoming traditional software bottlenecks and reducing manual programming overhead, allowing scientists to focus on primary research and discovery rather than code maintenance. (source: https://openai.com/index/scientific-computing-agentic-ai)

Hacker News

8 stories
01

Kimi Linear: An Expressive, Efficient Attention Architecture (2025)

Researchers published a paper introducing Kimi Linear, a linear attention architecture designed to replace standard quadratic softmax attention in large language models. The architecture leverages specialized gating mechanisms and hardware-aware optimization kernels to reduce computational complexity from quadratic to linear relative to sequence length. According to the technical notes published on Sebastian Raschka's blog (discussion: https://news.ycombinator.com/item?id=49085698), the architecture underpins Moonshot AI's Kimi K3 model, combining mixture-of-experts routing and reinforcement learning to manage computational bottlenecks during long-context inference (source: https://arxiv.org/abs/2510.26692).

02

A $500 RL fine-tune of a 9B open model beat frontier models on catalog review

Fermisense demonstrated that a $500 reinforcement learning fine-tune of an open-source 9B parameter model can outperform proprietary frontier models on a specialized catalog review task. By using targeted domain-specific optimization instead of relying on expensive, general-purpose commercial API models, the training run achieved high precision in structured validation pipelines while reducing long-term inference costs. This project shows that highly customized, budget-conscious reinforcement learning applied to smaller open models can achieve superior domain-specific results compared to massive commercial systems (source: https://fermisense.com/when-machines-take-the-wheel/).

03

Discovering Cryptographic Weaknesses with Claude

Anthropic released research demonstrating how its Claude model can identify complex cryptographic vulnerabilities within software systems. The study shows that Claude's natural language understanding and code analysis capabilities allow it to detect subtle implementation flaws, deprecated algorithms, and side-channel vectors that standard automated scanners miss. The paper details the dual-use implications of deploying frontier large language models for automated vulnerability discovery, explaining how they can assist software developers in code auditing while introducing risks of malicious exploitation (source: https://www.anthropic.com/research/discovering-cryptographic-weaknesses).

04

MCP 2026-07-28 Specification: transport going stateless

The Model Context Protocol community released the 2026-07-28 specification, transitioning the open protocol's transport layer to a stateless architecture. By removing session state requirements from the communication channel, the update simplifies routing mechanics, boosts resilience against transient network interruptions, and lowers the integration barrier for custom MCP endpoints. The specification defines updated request-response patterns and error-handling protocols, facilitating seamless migration for existing agentic integrations and serverless implementations (source: https://blog.modelcontextprotocol.io/posts/2026-07-28/).

05

Show HN: Formally verified 3D CSG: Trust 93 lines spec, not 1000 lines AI code

Developer schildep introduced a formally verified 3D constructive solid geometry operation implemented in the Lean 4 theorem prover. To address the reliability limits of raw AI outputs, an artificial intelligence system was used to autonomously generate over 60,000 lines of mathematical proofs to certify a complex 1,000-line mesh intersection kernel. Human developers only need to trust a concise 93-line specification, as the Lean compiler verifies the AI's proofs at compile time to guarantee complete functional correctness and geometric well-formedness (source: https://github.com/schildep/verified-3d-mesh-intersection).

06

Google's Beyond Zero: Enterprise Security for the AI Era

Google introduced Beyond Zero, a comprehensive enterprise security framework designed to address vulnerabilities introduced by deploying artificial intelligence in corporate environments. Extending traditional Zero Trust security principles, Beyond Zero secures AI data pipelines, model training environments, and inference endpoints. The architectural blueprint guides enterprises on implementing secure data lineage, establishing access controls for both human operators and automated AI agents, and defending against adversarial attacks like prompt injection and data poisoning (source: https://spawn-queue.acm.org/doi/10.1145/3819083).

07

Now Is the Time to Give LLMs Access to the ACM Digital Library

An opinion piece in the Communications of the ACM advocates for granting Large Language Models direct access to the ACM Digital Library's peer-reviewed computer science repository. The author argues that training models on structured, authoritative scientific literature will improve model accuracy, logical reasoning, and code generation. This stands in contrast to current training pipelines reliant on low-quality web data, positioning access to academic repositories as a necessary step for turning generative AI into reliable assistants for scientific research (source: https://cacm.acm.org/opinion/now-is-the-time-to-give-llms-access-to-the-acm-digital-library/).

08

"Uncensored" open LLMs are measurably more optimistic than their base models

Researchers published a study analyzing the behavioral changes in open-source large language models when safety guardrails are stripped away to make them "uncensored." The findings demonstrate a statistically significant increase in optimistic bias and positive sentiment across generated outputs compared to their aligned base versions. The study suggests that removing alignment and safety filters alters the model's underlying cognitive framing, making uncensored models systematically less risk-averse and more optimistic than aligned models (source: https://arxiv.org/abs/2607.17427).

Twitter

8 stories
01

Gemini Managed Agents API Receives New Environment Hook Features

Google has announced major updates for its Gemini Managed Agents API, introducing advanced support for model environment hooks. The update implements new block and linear functions designed to manage execution flows inside agentic workflows. These additions provide developers with deeper integration capabilities, process monitoring, and tighter coordination between the model's output and its operational environment. The tools aim to streamline complex, reactive AI agent operations in enterprise production environments. (source: https://x.com/Google/status/2082147432486342788)

02

Runway Announces Upcoming Release Of Seedance 2.5 Video Generation Model

Runway has officially teased the imminent launch of Seedance 2.5, its latest generative video synthesis model. The upcoming release targets significant enhancements in overall video quality, motion coherence, and user-directed camera control. Designed to improve high-fidelity synthesis for digital content creators, filmmakers, and artists, the model acts as a key step forward for Runway's generative suite. Specific technical details and formal release schedules remain under wraps. (source: https://x.com/runwayml/status/2082112674666529224)

03

Kling AI Celebrates X Platform Community Growth With Creative Features

Kling AI has developed new creative tools and capabilities leveraging its Kling Model Context Protocol (MCP). The update is designed to integrate generative AI technologies into user-facing social experiences, fostering closer interaction with its community. By deploying these generative features, the company aims to highlight the utility of MCP and expand interactive content creation capabilities. (source: https://x.com/Kling_ai/status/2081990170115784895)

04

Analysis of the Kimi K3 Model Architecture and Scaling Methodology

The Kimi team has released Kimi K3, a new scaled-up open-weight large language model. Architectural analysis indicates that K3 serves as a production-grade scaled version of the proprietary Kimi Linear model architecture introduced last year. By applying established scaling laws and engineering refinements, the development team has transitioned their linear framework into a highly competitive option for the open-weight model ecosystem. (source: https://x.com/rasbt/status/2082098201247600765)

05

The Evolution Of Agent SEO And Its Future Impact On Business Models

The emerging paradigm of Agent SEO is reshaping traditional search engine optimization strategies by deploying autonomous AI agents. These goal-oriented agents leverage sophisticated large language models to shift online search dynamics from keyword-focused content creation to intent-driven discovery and active interaction. This integration of agentic workflows is predicted to transform how businesses manage online visibility and discovery in an era defined by autonomous digital actors rather than passive search algorithms. (source: https://x.com/natolambert/status/2082171901787726315)

06

Viral AI Cat Sensation Daon_box Reaches Over Twenty Million Views

Kling AI has released a new episode of its Creator Off Script series featuring Korean content creator Daon_box. The creator has generated over 20 million views since the beginning of the year by crafting AI-generated cat adventures. This feature showcases the creative application of generative video synthesis in storytelling, demonstrating how modern AI video tools allow creators to build immersive narrative content and capture global audiences. (source: https://x.com/Kling_ai/status/2082071627484024995)

07

Limitations of Large Language Models in Mastering Chess Logic

Gary Marcus has highlighted critical reasoning and logical limitations in current large language models, pointing to their consistent failure to adhere to the fundamental rules of chess without external software tools. Marcus challenges the notion that massive scaling of datasets can yield true intelligence or functional world models. He argues that LLMs inherently lack rule-based reasoning, and relying on external APIs to patch these gaps fails to resolve fundamental architectural limitations in statistical models. (source: https://x.com/GaryMarcus/status/2082174654928884204)

08

Increasing Complexity in Modern Large Language Model Architectures

A review of modern large language model trends reveals a significant shift toward increased structural and architectural complexity. Rather than relying on simple, standard transformer scaling, developers and research labs are designing intricate configurations to optimize model reasoning, processing efficiency, and memory constraints. This evolutionary trajectory suggests that pushing the boundaries of intelligence demands complex architecture changes beyond simple parameter scaling. (source: https://x.com/rasbt/status/2082101241908249014)

huggingface

8 stories
01

Kimi K3: Open Frontier Intelligence

Researchers have introduced Kimi K3, a 2.8-trillion parameter Mixture-of-Experts (MoE) model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Built on Kimi Delta Attention, Attention Residuals, and Stable LatentMoE, K3 achieves a 2.5x scaling efficiency improvement over its predecessor. Post-training highlights include reinforcement learning across general, agentic, and coding domains. Evaluated across diverse benchmarks, Kimi K3 consistently outperforms other open and proprietary models in its suite, only trailing Claude Fable 5 and GPT-5.6 Sol. The model weights are publicly released to accelerate frontier intelligence research. (source: https://huggingface.co/papers/2607.24653)

02

StateAct: Program State, before Pixels, for Long-Horizon Computer-Use Agents

Researchers have introduced StateAct, a code-first, multi-agent harness built around program state for long-horizon computer-use agents. Instead of relying solely on screenshots, StateAct's main agent interacts directly with program files and APIs using code, delegating only 1.1% of steps to a GUI subagent. On the OSWorld 2.0 benchmark, StateAct increases the binary success rate of Claude Opus 4.8 from 20.6% to 26.9% and its partial success rate from 54.8% to 61.6%, while reducing task execution costs by approximately 9x compared to standard screenshot-based baselines. (source: https://huggingface.co/papers/2607.22798)

03

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation

Researchers have introduced a unified, controlled multi-turn environment to study how foundation model agents acquire, shape, and integrate multi-turn long-horizon planning. The study tracks progress across three key training stages: acquisition during pre-training, shaping via GRPO and on-policy distillation (OPD) during post-training, and integration using multi-teacher on-policy distillation (MOPD). Findings demonstrate that explicit world modeling via Chain-of-Thought yields stronger generalization, and that OPD provides more consistent update directions than GRPO under low-quality and long-horizon settings. (source: https://huggingface.co/papers/2607.24720)

04

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

Researchers have developed WorldDiT, a unified diffusion transformer architecture that integrates continuous action generation with visual world modeling. Designed for robotic control, WorldDiT operates without relying on a large pretrained vision-language model action backbone, instead generating continuous action chunks and predicting future camera frames. Evaluated across four LIBERO simulation suites, WorldDiT achieves strong performance, establishing a robust sub-billion-parameter baseline on the reported Pareto frontier for total model parameters and mean success rate. (source: https://huggingface.co/papers/2607.23909)

05

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification

Researchers have introduced Sol-Attn (Sparsifying online attention), a training-free framework designed to accelerate inference in diffusion transformers for video generation. Sol-Attn addresses the computational bottleneck of long token sequences by combining dynamic routing, sparse computation, and approximation correction in a single online-softmax pass. By utilizing on-the-fly block thresholding and reusing proxy scores, it avoids resource-heavy metric materialization. Experiments demonstrate 2.1x and 2.3x end-to-end speedups for video generation and editing, respectively, while maintaining high visual quality. (source: https://huggingface.co/papers/2607.24027)

06

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Researchers have proposed a system that pairs frozen, non-deterministic language models with a growing, persistent memory of verified solutions. This mechanism enables zero-token, bit-exact, and deterministic execution of previously solved problem families. Across 180 fresh instances spanning nine problem families, four distinct model architectures achieved perfect accuracy using zero generation tokens. The system also scales working context, supporting a 6,000,000-token sliding window on a single 46 GB GPU, bypassing standard vLLM and SGLang context constraints. (source: https://www.huggingface.co/papers/2607.23806)

07

From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search

Researchers have proposed Multi-Agent Protocol Distillation (MAPD), a joint knowledge distillation and reinforcement learning framework that bridges the distribution gap between proprietary teacher models and open-source students in agentic search. MAPD structures exploration traces into a style-normalized JSON protocol to provide a dense distillation signal to the student. Across seven QA benchmarks, MAPD improved average success rates, achieving 39.4% on Qwen3-1.7B and 44.4% on Qwen3-4B, while preventing issues like style drift and verbosity degeneration. (source: https://huggingface.co/papers/2607.24280)

08

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

Researchers have introduced OmniVAE, a jointly trained audio-video variational autoencoder (VAE) that establishes semantic alignment between audio and video latent spaces. Rather than training modalities separately, OmniVAE employs a segment-level audio-video contrastive objective alongside feature distillation from frozen modality-specific encoders. This cross-modal alignment yields stronger latent representations, translating into higher generation quality and improved synchronization for downstream text-to-audio-video generative models. (source: https://huggingface.co/papers/2607.23855)