NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-02ENGLISH EDITION
This issue
—
All time
—

Hacker News

8 stories
01

Senior SWE-Bench: Open-Source Benchmark That Assesses Agents as Senior Engineers

Snorkel AI has introduced Senior SWE-Bench, an open-source benchmarking framework designed to evaluate the capabilities of AI agents at a senior software engineering level. Unlike traditional benchmarks that focus on isolated, junior-level programming tasks, this new framework tests autonomous agents on complex software development lifecycles, including multi-file code modifications, architectural decision-making, and understanding large legacy codebases. The tool aims to set a higher performance standard for enterprise-ready AI software engineers, allowing researchers to measure their systems against realistic software industry scenarios. (source: https://senior-swe-bench.snorkel.ai/)

02

CursorBench 3.1

Cursor has released CursorBench 3.1, an updated evaluation benchmark designed to measure the performance of AI-powered code editors. The benchmark focuses on real-world coding scenarios, assessing how effectively language models can suggest, edit, and refactor code within complex developer workflows. By providing a standardized suite of tests, CursorBench 3.1 enables developers and researchers to systematically compare the efficiency, accuracy, and practical utility of various generative AI models integrated into modern IDEs. This version introduces refined metrics to better capture fine-grained editing capabilities and agentic software engineering tasks. (source: https://cursor.com/evals)

03

Launch HN: Manufact (YC S25) – MCP Cloud

Co-founders Pietro and Luigi have launched Manufact (formerly mcp-use), a YC S25 backed cloud platform designed specifically for Model Context Protocol (MCP) applications and servers. Serving as a dedicated hosting infrastructure, the platform allows developer teams to deploy, iterate, test, and monitor MCP implementations in production environments. This specializes the deployment pipeline for developers building AI agents and services that rely on the MCP standard to interface with LLMs. Alongside their cloud offering, the team continues to develop and maintain their popular open-source MCP SDKs under the mcp-use brand. (source: https://manufact.com)

04

Kimi K2.7 Code is generally available in GitHub Copilot

GitHub has officially announced the general availability of the Kimi K2.7 Code model within GitHub Copilot. Developed by Moonshot AI, this advanced model is tailored specifically for code generation, software development tasks, and complex programming workflows. By integrating Kimi K2.7 Code, GitHub Copilot expands its multi-model ecosystem, allowing developers to choose from a diverse selection of state-of-the-art artificial intelligence models. This integration aims to deliver enhanced coding efficiency, higher accuracy in logic reasoning, and localized optimization for developers globally, emphasizing the trend toward customizable AI assistance in modern integrated development environments. (source: https://github.blog/changelog/2026-07-01-kimi-k2-7-is-now-available-in-github-copilot/)

05

Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train

A research paper has investigated the structural efficiency of transformer architectures in reinforcement learning contexts, exploring whether a single transformer layer can suffice. The authors demonstrate that a single-layer transformer can achieve performance levels matching those of full-parameter reinforcement learning training. By simplifying the model's depth, the study challenges the conventional assumption that deep architectures are always necessary for complex decision-making tasks. This discovery opens up new avenues for optimizing reinforcement learning models, significantly reducing computational overhead and training times without sacrificing agent performance. (source: https://arxiv.org/abs/2607.01232)

06

Show HN: ctx – Search the coding agent history already on your machine

A Rust command-line tool named ctx has been developed to provide coding agents with a form of long-term memory. The utility ingests historical agent transcripts and log files into a local, structured SQLite database on the user's machine, allowing ranked text matches without requiring complex graph databases or external, hosted memory services. To start a new task, developers can configure their agents to run a dedicated subagent that searches this local database, retrieving historical context to recall previous error resolutions and runbook executions. (source: https://github.com/ctxrs/ctx)

07

AI can't be listed as inventor on patent applications, Japan's top court rules

Japan's Supreme Court has officially ruled that artificial intelligence systems cannot be designated as inventors on patent applications. The country's highest judicial body upheld lower court decisions, clarifying that under current patent law framework, an inventor must be a natural person. This significant ruling aligns Japan's intellectual property policies with those of other major jurisdictions, including the United States, the European Union, and the United Kingdom, which have similarly determined that AI lacks the legal personhood required for inventorship. The decision addresses a growing global debate surrounding generative AI and intellectual property rights, reinforcing traditional legal standards that attribute creative ownership solely to human beings. (source: https://japannews.yomiuri.co.jp/science-nature/technology/20260306-314930/)

08

Claude's AskUserQuestion: "No response after 60s – continued without an answer"

A GitHub issue has highlighted an operational behavior within the Claude Code CLI tool, where the AskUserQuestion prompt times out and continues execution without a response after 60 seconds. This default automated progression is raising active discussions among users and developers regarding control flow, safety, and predictability in autonomous agent developer tools. The community is currently evaluating how timeout thresholds should be managed, whether execution should pause indefinitely by default, and how to configure fallback behaviors when human-in-the-loop validation is missed, especially when agents have permissions to modify local files or execute terminal commands. (source: https://github.com/anthropics/claude-code/issues/73125)

Twitter

8 stories
01

Anthropic Integrates Claude Fable 5 Into Claude Tag Platform

Anthropic has announced the availability of Claude Fable 5 within the Claude Tag environment. The update follows internal technical discussions detailing the architectural progression of Anthropic's developer-focused tools and their expanded adoption across the organization's non-engineering departments. By integrating Claude Fable 5, Anthropic aims to provide developers with enhanced capabilities for autonomous coding assistance and complex, agentic workflow automation inside the unified Claude Tag interface. (source: https://x.com/ClaudeAI/status/2072725610061803522)

02

Sakana AI Releases Fugu Multi-Agent Orchestration Tool on OpenCode

Sakana AI has officially released Fugu, an advanced multi-agent orchestration framework, to the open-source platform OpenCode. The development team utilized the OpenCode ecosystem during the architectural design and implementation phases of the framework. By contributing Fugu to the public domain, Sakana AI aims to simplify how specialized agents communicate, coordinate, and execute distributed tasks, fostering community innovation in agentic technology. (source: https://x.com/hardmaru/status/2072495247419228162)

03

Pika Labs Introduces New MCP Integration for Enhanced Block Universe Creation

Pika Labs has launched its new Model Context Protocol (MCP) integration, designed to bridge custom digital design environments with generative video synthesis. The release enables users to connect external data environments directly to the Pika platform, facilitating greater technical control over the compositional elements of interactive block universes. This deployment makes generative video tools more interoperable with the broader developer ecosystem. (source: https://x.com/pika_labs/status/2072767397245948223)

04

Pika Labs Introduces New Voxel-It Skill For Minecraft-Style Visual Effects

Pika Labs has unveiled Voxel-It, a new creative skill designed to transform static photographs into stylized, Minecraft-inspired voxel worlds. By utilizing advanced computer vision, the tool allows users to apply a selective voxelization effect to backgrounds and objects while preserving the original human features and integrity of people in the frame. The release expands the generative capabilities of the Pika platform for digital artists. (source: https://x.com/pika_labs/status/2072767395832500655)

05

Adaption Appoints Hugo Larochelle As New Scientific Lead

Adaption has appointed Dr. Hugo Larochelle as its new Scientific Lead. Dr. Larochelle, who currently serves as the Scientific Director at Mila (the Quebec AI Institute), will execute this dual-role to guide Adaption's research in human-AI adaptation. The appointment highlights a broader industry trend of integrating academic talent to steer deep learning and machine learning strategies within specialized AI organizations. (source: https://x.com/sarahookr/status/2072711447273144328)

06

SIGReg Pretraining Outperforms CLIP and SigLIP on Large Scale Vision Tasks

Researchers have introduced SIGReg, a vision-language model that outperforms established benchmarks CLIP and SigLIP on large-scale vision tasks. By scaling vision-language joint-embedding predictive architecture (JEPA) pretraining from CC12M to the larger Datacomp-L dataset, the system demonstrates that non-contrastive, optimized pretraining strategies can capture complex visual features with high accuracy and efficiency without relying on traditional contrastive paradigms. (source: https://x.com/ylecun/status/2072719097888928052)

07

Introducing LeVLJEPA: A Fully Non-Contrastive Vision-Language Framework

Researchers have unveiled LeVLJEPA, representing the first fully non-contrastive end-to-end vision-language pretraining framework. Designed to bypass the computational limitations of conventional contrastive learning, LeVLJEPA achieves downstream task performance competitive with existing state-of-the-art multimodal vision-language models. The self-supervised architecture enables robust joint feature extraction while reducing reliance on massive paired datasets during training. (source: https://x.com/ylecun/status/2072718977244045713)

08

Anthropic Announces Built With Claude Life Sciences Virtual Hackathon

Anthropic has launched the 'Built with Claude: Life Sciences' global virtual hackathon in collaboration with the Gladstone Institutes. The week-long event gathers developers and scientists to apply the Claude large language model to specialized scientific and biomedical research workflows. Participants will leverage Claude's capabilities to solve data-analysis and laboratory challenges, establishing new methodologies for modern scientific discovery. (source: https://x.com/ClaudeDevs/status/2072704338057912695)

huggingface

8 stories
01

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Microsoft researchers introduced HealthAgentBench, a unified evaluation suite of 54 agentic healthcare tasks across seven distinct clinical categories. The benchmark replicates end-to-end clinical workflows where LLM-based agents must navigate raw medical data, EHR pipelines, and complex imaging environments. Evaluating frontier models on the suite highlights the extreme difficulty of long-horizon clinical reasoning; the top-performing model, Codex GPT-5.5, achieved a success rate of only 42%. The benchmark and environments are open-sourced to encourage development in safe, automated clinical decision-making. (source: https://huggingface.co/papers/2606.31179)

02

Autonomous Scientific Discovery via Iterative Meta-Reflection

Researchers proposed DiscoPER, an autonomous large language model framework for open-ended scientific discovery. DiscoPER dynamically writes and executes code to explore scientific datasets without pre-specified research questions, demanding that all generated hypotheses pass rigorous statistical validation. It introduces a second-order meta-reflection mechanism that treats its own accumulated discoveries as empirical data to find structural patterns and direct research toward epistemic gaps. Evaluated on the iNatDisco ecological benchmark, DiscoPER successfully recovered eight of nine known scientific patterns, demonstrating strong performance over classical causal discovery and basic LLM baselines. (source: https://huggingface.co/papers/2607.01131)

03

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Shengshu Technology introduced TurboServe, a specialized serving system designed for the unique computational demands of interactive, progressive streaming video generation. TurboServe coordinates online session placement and GPU provisioning using a migration-aware placement controller and a load-driven autoscaling controller. It implements coalesced chunk processing, GPU-CPU offloading for session suspension, and NCCL-based online migrations. Evaluated on production traces using up to 64 NVIDIA B300 GPUs, TurboServe reduced worst-case per-chunk generation latency by 37.5% and lowered total GPU operating costs by 37.2% on average. (source: https://huggingface.co/papers/2606.19271)

04

CausalMix: Data Mixture as Causal Inference for Language Model Training

Researchers proposed CausalMix, a novel framework that optimizes data mixture ratios for pretraining language models by framing the task as a causal inference problem. CausalMix treats data pool statistical features as covariates and domain mixtures as treatments to estimate the Conditional Average Treatment Effect (CATE). After training on 512 runs of Qwen2.5-0.5B, the framework successfully extrapolated and optimized mixtures for an 800K data pool, training a 7B model and generalizing to long chain-of-thought data on Qwen3-4B-Base, outperforming RegMix and other baselines. (source: https://huggingface.co/papers/2607.01104)

05

Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?

An empirical audit of repository-level performance-optimization benchmarks—specifically GSO, SWE-Perf, and SWE-fficiency—revealed significant reliability issues in measuring coding agents. By replaying official reference patches across four types of Google Cloud machines, the study found that reference patches satisfy original validity rules in all replays for only 39/102 GSO, 11/140 SWE-Perf, and 411/498 SWE-fficiency tasks. The findings demonstrate that public submission rankings depend heavily on fragile scoring rules, with SWE-Perf particularly impacted by reference patches yielding near-zero runtime differences. (source: https://huggingface.co/papers/2607.01211)

06

AutoTrainess: Teaching Language Models to Improve Language Models Autonomously

Researchers developed AutoTrainess, an autonomous language model agent designed to perform post-training processes without intensive human oversight. AutoTrainess structures post-training tasks—including iterative planning, data preparation, training execution, and evaluation—into explicit workflows and execution constraints rather than relying on a raw command-line interface. On the PostTrainBench benchmark, the framework achieved an average score of 26.94 using GPT-5.4 (Codex), outperforming CLI-only baselines and demonstrating successful generalization by improving DeepSeek-V4-Flash from 12.13 to 19.58. (source: https://huggingface.co/papers/2606.31551)

07

The State-Prediction Separation Hypothesis

Researchers proposed the state-prediction separation hypothesis, which suggests that separating the computation stream used to predict the next token from the stream used to store history results in better language modeling. To test this, they designed a novel Transformer variant featuring dual computational streams and conducted pretraining experiments across multiple scales. The experimental results show that decoupling state storage and prediction consistently improves data and compute efficiency, lowering validation loss and yielding a 2 to 3 percentage point gain on downstream tasks. (source: https://huggingface.co/papers/2607.01218)

08

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Researchers developed Graph-PRefLexOR, a family of graph-native reasoning models designed to generate traceable, scientifically valid hypotheses in materials science. Fine-tuned with Group Relative Policy Optimization (GRPO), the models organize neural generation into structured phases including mechanism exploration, graph construction, and hypothesis synthesis. Evaluated on 100 materials design questions, Graph-PRefLexOR achieved 40% to 65% improvements over base models, displaying 2 to 3 times greater semantic diversity. Tests confirmed that additional compute scales long-range conceptual recombination rather than merely expanding semantic coverage. (source: https://huggingface.co/papers/2607.00924)