NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-12DEFAULT EDITION
This issue
—
All time
—

AI Blog

3 stories
01

TCS Partners With Anthropic To Deploy Claude In Regulated Industries

Anthropic announced a strategic partnership with Tata Consultancy Services (TCS) to deploy its Claude Large Language Models across highly regulated enterprise sectors including finance, healthcare, and the public sector. Under the agreement, TCS will integrate Claude within its own workforce of 50,000 employees spanning 56 countries and join the Claude Partner Network. TCS is launching a dedicated practice to design Claude-powered systems for tasks such as insurance claims processing, banking lending advisory, and customer service operations. Additionally, developers will utilize Claude Code for software engineering. (source: https://www.anthropic.com/news/tcs-anthropic-partnership)

02

Academy Launches Three New Courses on AI Workflows

OpenAI launched three new OpenAI Academy courses designed to teach individuals practical workflows and the deployment of autonomous AI agents. The educational modules transition users from basic prompt interactions to building and deploying functional agents that can automate complex professional workflows. The curriculum focuses on establishing repeatable routines and integrating digital assistants directly into business productivity operations to streamline modern enterprise tasks. (source: https://openai-com-index-academy-courses-applying-ai-at-work)

03

Preply Integrates OpenAI to Power Personalized Language Tutoring

Preply integrated OpenAI technology to launch automated, AI-generated lesson summaries that deliver personalized language tutoring to its global users. The system synthesizes tutoring session details to automatically generate feedback, customized learning exercises, and tailored actionable insights. By combining generative artificial intelligence capabilities with live, human-led tutoring sessions, the integration aims to increase vocabulary retention and speed up feedback loops for language learners. (source: https://openai.com/index/preply)

Hacker News

7 stories
01

AI agent bankrupted their operator while trying to scan DN42

An autonomous AI agent has depleted its operator's financial resources while attempting to scan the decentralized peer-to-peer network DN42. Designed for automated network exploration, the agent initiated an uncontrolled volume of high-compute queries and API requests. Due to a lack of budget caps, rate limiting, and execution boundaries within its deployment framework, the system entered an unchecked recursive loop that quickly exhausted the operator's linked cloud resources. The incident highlights critical security and financial risks associated with deploying autonomous LLM-based systems without strict real-time cost-containment measures. (source: https://lantian.pub/en/article/fun/ai-agent-bankrupted-their-operator-scan-dn42lantian.lantian/)

02

Claude Fable is relentlessly proactive

Anthropic has developed an emerging AI model named Claude Fable, which is characterized by its highly proactive and autonomous behavior. Unlike traditional reactive assistants that depend on constant human prompting, Claude Fable can anticipate user needs, suggest developmental steps, and independently execute and refine complex workflows. The model is capable of analyzing context-rich environments and self-correcting errors during task execution. This release marks a shift toward fluid, cooperative human-AI partnerships, positioning the system to proactively manage software development, research, and data analysis tasks with minimal human oversight. (source: https://simonwillison.net/2026/Jun/11/fable-is-relentlessly-proactive/)

03

Kimi K2.7-Code: open-source coding model with better token efficiency

Moonshot AI has released Kimi K2.7-Code, an advanced open-source large language model optimized for software engineering. The model is specifically engineered to deliver superior token efficiency, reducing operational costs and latency during code generation. It excels at complex programming tasks, multi-turn interactions, code completion, and interactive debugging. By open-sourcing Kimi K2.7-Code, Moonshot AI addresses the high computational overhead typically associated with large-scale generative AI workflows, allowing developers to build more cost-effective and responsive localized programming utilities. (source: https://huggingface.co/moonshotai/Kimi-K2.7-Code)

04

Maxproof

Researchers have introduced Maxproof, a novel machine learning framework designed to improve automated theorem proving and formal verification. The system integrates advanced neural architectures with symbolic solvers to significantly optimize search space discovery. Maxproof utilizes iterative reinforcement learning to dynamically refine its strategy formulation based on historical execution feedback. In benchmarks, the framework successfully synthesized complex mathematical proofs that traditionally required expert manual intervention. This research holds major implications for verifying hardware design, ensuring software safety, and developing computational systems that require minimal human supervision. (source: https://arxiv.org/abs/2606.13473)

05

Launch HN: BitBoard (YC P25) – Analytics Workspace for Agents

Y Combinator startup BitBoard has launched an agentic analytics workspace designed to integrate autonomous AI agents into enterprise business intelligence workflows. Founded by Connor and Ambar, the platform introduces a shared visualization and database infrastructure layer. This architecture allows human users and AI coding agents to collaborate in real-time on live dashboards. By establishing persistent analytical environments, BitBoard resolves the common documentation and reporting bottlenecks associated with traditional conversational interfaces, transforming ephemeral chat analyses into structured and collaborative business intelligence reports. (source: https://bitboard.work/)

06

How to setup a local coding agent on macOS

A technical guide has been published detailing how to configure and run a fully local, private AI-powered coding agent on macOS. By utilizing open-source tools and running local large language models optimized for Apple Silicon, developers can set up secure coding assistants without relying on external cloud APIs. The tutorial walks through the installation of execution frameworks, IDE integrations, and performance tuning configurations to enable low-latency code completion and automated editing, highlighting private and offline alternatives to cloud-hosted coding systems. (source: https://ikyle.me/blog/2026/how-to-setup-a-local-coding-agent-on-macos)

07

Don't You Just Upload It to ChatGPT?

An analytical article addresses the operational challenges of relying on simple document uploads to ChatGPT for complex business processes. The author analyzes the privacy, legal, and token-limit constraints that make basic web interfaces insufficient for enterprise use cases. To resolve these limitations, the text outlines the necessity of building custom Retrieval-Augmented Generation (RAG) pipelines, implementing robust vector databases, and fine-tuning models on domain-specific datasets to achieve secure, context-aware, and precise generative AI deployments. (source: https://correresmidestino.com/dont-you-just-upload-it-to-chatgpt/)

Twitter

6 stories
01

Google DeepMind Launches New Robotics Accelerator For European Startups

Google DeepMind has officially launched its new Robotics Accelerator program to support 15 European startups focused on advancing physical AI. Selected companies will undergo an intensive three-month cohort, receiving direct access to DeepMind's specialized AI stack, including their advanced Gemini Robotics models. In addition to technical infrastructure, participants will obtain hands-on mentorship and engineering support from Google's teams to help bridge the gap between theoretical models and real-world robotic hardware deployments. (source: https://x.com/GoogleDeepMind/status/2065388989146628563)

02

Google Project Genie Access Expands to Ultra 5X Tier Subscribers Globally

Google has expanded access to Project Genie, its generative AI model capable of creating interactive 2D worlds from text prompts or images, to Google AI Ultra 5X subscribers globally. Run under Google Labs, this rollout aims to scale experimental access to its gaming and interactive simulation technologies. The initiative integrates advanced creative world-generation capabilities into Google's premium subscription tier, encouraging wider adoption of interactive, multimodal AI research tools. (source: https://x.com/GoogleDeepMind/status/2065493012956791146)

03

Runway Hosts Fourth Annual AI Festival At Lincoln Center In New York City

Runway hosted its fourth annual AI Festival at the Lincoln Center in New York City, marking its most ambitious event to date. The gathering showcased ten finalist films demonstrating generative AI in creative storytelling. Keynote speaker and director Ron Howard shared insights on the integration of technology in cinematic production. The event highlights Runway's strategic focus on establishing its multimodal tools and generative media platforms within the professional filmmaking industry. This event was also discussed by Runway co-founder Cristóbal Valenzuela. (source: https://x.com/runwayml/status/2065440244762042576)

04

Concerns Over Anthropic Data Retention Policies and Intellectual Property

Naveen Rao has raised serious concerns regarding Anthropic's mandatory prompt and history retention policies for its Claude AI platform. The policy requires storing user interaction logs, which often contain sensitive corporate assets like proprietary design files and internal documentation. This practice has been declared untenable for enterprise deployments, emphasizing a growing conflict between secure corporate data requirements and LLM training practices. This concern was also echoed and shared by researcher Yann LeCun. (source: https://x.com/NaveenGRao/status/2065282350792282119)

05

A Technical Breakdown Of Policy Gradient Derivation In Reinforcement Learning

A comprehensive mathematical derivation of policy gradient methods has been detailed, outlining a key framework in reinforcement learning. The technical documentation details how gradients of expected cumulative rewards are calculated relative to policy parameters to optimize agent behavior. This mathematical derivation provides a theoretical reference for developers building deep reinforcement learning algorithms for robotic control systems and autonomous decision-making applications in stochastic environments. (source: https://x.com/natolambert/status/2065486388641018241)

06

Hailuo AI Offers 3,000 Bonus Credits for Early Login and Platform Migration

Hailuo AI has launched a promotional incentive offering 3,000 bonus credits to users who log in to the service prior to July 1. This program is structured to facilitate migration to the integrated MiniMax Hub ecosystem. Credits earned on the Hailuo AI generative platform can be fully transferred to MiniMax Hub, letting users apply these resources seamlessly across both creative AI systems during the transition. (source: https://x.com/Hailuo_AI/status/2065277563367551101)

huggingface

8 stories
01

MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling

MiniMax researchers introduced MaxProof, a population-level test-time scaling framework designed for competition-level mathematical proofs in the MiniMax-M3 model series. MaxProof coordinates three proof-oriented capabilities: proof generation, proof verification, and critique-conditioned proof repair. Using a generative verifier engineered for low false-positive rates during reinforcement learning, the system acts as a generator, verifier, refiner, and ranker during test-time searches. Applied to mathematics benchmarks, MaxProof test-time scaling achieved a score of 35/42 on IMO 2025 and 36/42 on USAMO 2026, exceeding the gold-medal threshold on both exams. (source: https://huggingface.co/papers/2606.13473)

02

Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning

Researchers proposed SWITCH, a switchable latent reasoning framework that optimizes hidden-state recurrence in large language models using on-policy reinforcement learning. By introducing explicit discrete boundary tokens to mark the entry and exit of latent computations, the framework makes latent-state reasoning compatible with standard Group Relative Policy Optimization (GRPO) gradients. Evaluated against prior hidden-state recurrence methods, SWITCH consistently improves reasoning performance at similar scales. Mechanistic analyses confirm that the learned switching policy performs targeted, causally important computations rather than serving as an inert placeholder. (source: https://huggingface.co/papers/2606.13106)

03

MiniMax Sparse Attention

MiniMax introduced MiniMax Sparse Attention (MSA), a blockwise sparse attention mechanism built on Grouped Query Attention (GQA) to support ultra-long contexts in large language models. MSA uses a lightweight Index Branch to select key-value blocks for each GQA group, which are executed via a custom GPU path employing exp-free Top-k selection. When evaluated on a 109-billion-parameter model with a 1-million token context, MSA reduced per-token attention compute by 28.4 times. Hardware acceleration tests on an H800 GPU showed a 14.2-times speedup in prefill and a 7.6-times speedup during decoding. (source: https://huggingface.co/papers/2606.13392)

04

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

Researchers introduced Evoflux, an inference-time evolutionary search framework designed to repair executable tool workflows for compact language model agents. Unlike traditional distillation, Evoflux evolves typed workflow graphs through structured edits, adaptive execution feedback, meta-guided redesign, and diversity pruning. On the MCP-Bench benchmark, which includes live Model Context Protocol (MCP) servers and 250 tools, Evoflux raised successful tool-use execution rates from 3% to a range of 17% to 24% across small planner models. This performance outperformed alternative supervised fine-tuning and direct reinforcement learning baselines. (source: https://zju-xyc.github.io/VIA-SD-Project-Page/)

05

InterleaveThinker: Reinforcing Agentic Interleaved Generation

Researchers developed InterleaveThinker, a multi-agent pipeline designed to enable interleaved text-image sequence generation in existing image generators. The pipeline uses a planner agent to orchestrate the generation sequence and a critic agent to identify deviations and refine instructions. The framework was constructed using specialized SFT datasets and optimized using a step-wise reinforcement learning framework via GRPO. Evaluations demonstrate that InterleaveThinker improves the generation performance of base models, achieving results comparable to Nano Banana and GPT-5 on interleaved generation benchmarks while enhancing spatial and logical reasoning capabilities. (source: https://huggingface.co/papers/2606.13679)

06

The Cold-Start Safety Gap in LLM Agents

Researchers identified the "cold-start safety gap," showing that tool-calling LLM agents are highly vulnerable at the beginning of user sessions but become safer as conversations progress. To study this, they introduced the Safety Over Depth for Agents (SODA) benchmark, evaluating models across up to 20 preceding tasks. Testing seven models across four families revealed that completing regular agentic tasks first improved safety by 9% to 52%. Representation analysis verified that processing safe, standard tasks naturally shifts the model's hidden states toward its safety-aligned activation regions. (source: https://github.com/Trustworthy-ML-Lab/Agent-Cold-Start-Safety-Gap)

07

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Researchers released WeaveBench, a benchmark designed to evaluate computer-use agents on long-horizon, hybrid-interface tasks requiring coordination between GUI, CLI, and code-editing environments. The dataset contains 114 tasks spanning 8 real-world domains. Evaluation is executed in a deployed Ubuntu desktop environment and managed by a trajectory-aware judge that inspects visual and structural artifacts to detect shortcut behaviors. The evaluation of state-of-the-art models revealed a maximum pass rate of only 41.2%, exposing a major gap in the ability of current agents to execute complex cross-interface workflows. (source: https://huggingface.co/papers/2606.09426)

08

EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery

Researchers proposed EurekAgent, an environment-engineered system designed to optimize autonomous scientific discovery using LLM agents. EurekAgent focuses on engineering the agent's operating environment across four key dimensions: permissions execution boundaries, Git-based artifact collaboration, budget-aware exploration, and human-in-the-loop supervision hooks. The framework achieved state-of-the-art performance on multiple mathematical and machine learning benchmarks, finding new 26-circle packing mathematical results with under $11 in API expenditures, demonstrating that environment design is highly effective for accelerating metric-driven discovery. (source: https://huggingface.co/papers/2606.13662)