NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-10ENGLISH EDITION
This issue
—
All time
—

AI Blog

2 stories
01

Nextdoor Engineers Leverage Codex and GPT-5.5 to Streamline Development

Nextdoor engineers are utilizing OpenAI's Codex model alongside GPT-5.5 to enhance their software development lifecycle. By integrating these generative models, the engineering team is able to investigate and resolve complex, hard-to-reproduce software issues more efficiently. The deployment allows developers to minimize time spent on boilerplate code, focusing instead on delivering high-quality product outcomes and robust user experiences across multiple platforms. (source: https://openai.com/index/nextdoor)

02

Notion Integrates OpenAI Codex to Boost Engineering Productivity

Notion has integrated OpenAI Codex into its development workflow to accelerate engineering productivity. The integration allows Notion's developers to generate comprehensive technical specifications in a single shot and build features like AI Voice Input for the web platform. This setup acts as a force multiplier, enabling small engineering teams to build, iterate, and ship products at a pace typically reserved for larger organizations. (source: https://openai.com/index/notion)

Hacker News

8 stories
01

AWS Bedrock to require sharing data with Anthropic for Mythos and future models

Amazon Web Services (AWS) announced a policy shift for its Bedrock platform requiring a 30-day data retention policy for high-end Anthropic models, including Fable 5 and Mythos 5. This policy is mandated by Anthropic to identify complex patterns of misuse. Utilizing these models means customer data will exit the standard AWS security boundaries to be stored by Anthropic, with automated deletion after 30 days unless subject to an ongoing investigation. This transition highlights significant changes in cloud data sovereignty for enterprise generative AI deployments. (source: https://news.ycombinator.com/item?id=48473166)

02

German ruling declares Google liable for false answers in AI Overviews

A German court ruled that Google is legally liable for false information displayed within its AI Overviews. The court determined that because Google dynamically synthesizes and presents these AI-generated summaries under its own brand, they constitute Google's own published statements rather than mere index links. This landmark decision challenges traditional safe harbor protections and establishes that search operators must verify generative outputs or face liability for defamation, misinformation, or copyright infringement. (source: https://the-decoder.com/landmark-german-ruling-declares-googles-ai-overviews-are-googles-own-words-and-makes-it-liable-for-false-answers/)

03

DiffusionGemma: 4x Faster Text Generation

Google has announced the release of DiffusionGemma, a framework designed to accelerate text generation speeds up to four times faster than standard baseline models. By incorporating advanced diffusion-based sampling techniques directly into the generative process of the Gemma family of open models, researchers have successfully mitigated standard autoregressive bottlenecks. This approach maintains high output quality and semantic coherence while significantly reducing computational latency and resource utilization during generation tasks, establishing a new performance benchmark for scalable open-weight text generation. (source: https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/)

04

A €0.01 bank transfer could compromise a banking AI agent

A security analysis has revealed how a trivial €0.01 financial transaction can be exploited to compromise a banking AI assistant. By injecting malicious instructions into the transaction's reference field, attackers can execute indirect prompt injection attacks. When the AI agent processes the transaction history, it parses and executes the embedded instructions, leading to potential data exfiltration or unauthorized transactions. The investigation emphasizes the urgent need for strict isolation between data and instruction planes in financial AI architectures. (source: https://blue41.com/blog/how-we-helped-bunq-secure-their-financial-ai-assistant/)

05

Apache Burr: Build reliable AI agents and applications

Apache Burr has been launched as an open-source framework designed to help developers build, manage, and monitor reliable AI agents and complex stateful applications. By representing applications as state machines, Apache Burr enables developers to define clear state transitions, monitor system behavior in real-time, and debug complex autonomous or conversational agents. This structured approach directly addresses non-deterministic LLM behavior, ensuring that AI-driven workflows remain predictable, debuggable, and enterprise-ready. (source: https://burr.apache.org/)

06

Policy on the AI Exponential

Dario Amodei published an essay addressing the urgent need for a comprehensive policy framework to manage the exponential growth of artificial intelligence capabilities. Highlighting that traditional policy-making timelines are inadequate for rapid AI development, the author proposes concrete governance strategies to mitigate catastrophic risks, secure critical hardware supply chains, and establish international safety standards. The discussion emphasizes the dual-use nature of advanced AI models and the necessity of proactive collaboration between technology developers and global policymakers. (source: https://darioamodei.com/post/policy-on-the-ai-exponential)

07

Rich Sutton on AI creativity and discovery

A historical presentation by reinforcement learning pioneer Rich Sutton explores the foundational concepts of AI creativity and discovery. Sutton discusses how computational agents can bypass human-designed heuristics to discover novel strategies autonomously through trial, error, and search optimization. The talk highlights his broader advocacy for general, scalable methods like reinforcement learning over human-engineered domain knowledge, serving as a conceptual guide for researchers building autonomous agents capable of generating genuine novelty and redefining machine intelligence. (source: https://twitter.com/RichardSSutton/status/2061216087744946656)

08

Show HN: HelixDB – A graph database built on object storage

HelixDB has been released as an OLTP graph database built on object storage, integrating native vector search and full-text search capabilities. Designed to simplify data retrieval pipelines for modern AI applications, HelixDB eliminates the need for application-level query coordination. Users can natively perform joins and complex queries that span graph structures, semantic embeddings, and text searches. The project addresses the challenges of connecting disconnected database systems in modern artificial intelligence application stacks. (source: https://github.com/HelixDB/helix-db/tree/main)

Twitter

4 stories
01

Google Introduces DiffusionGemma Open Experimental Diffusion Model

Google has officially released DiffusionGemma, an experimental, open-source model that integrates text-to-image diffusion research capabilities into the Gemma framework. Designed for developers and researchers, this model aims to explore the intersection of diffusion techniques and large language model architectures. By making this model available, Google continues its commitment to transparent AI development, allowing the broader research community to experiment with its latest advancements in image generation. DiffusionGemma serves as a significant technical milestone, demonstrating how established diffusion mechanisms can be adapted to enhance the utility and creative potential of the Gemma product family. (source: https://x.com/Google/status/2064746655388209297)

02

Luma Labs Releases Ray 3.2 For Advanced Generative Video Capabilities

Luma Labs has officially released Ray 3.2, a significant update to their generative AI platform. This iteration represents the latest step in their ongoing development of high-quality video generation technology. The Ray 3.2 release focuses on performance improvements and refined visual output, allowing users to leverage enhanced capabilities for creative and professional projects. By continuing to iterate on their foundational models, Luma Labs aims to maintain its competitive edge in the rapidly evolving landscape of generative media tools. (source: https://x.com/LumaLabsAI/status/2064566309874975056)

03

Kling AI Expands Creative Capabilities With Advanced Video Generation Features

Kling AI has announced significant enhancements to its generative video platform following extensive feedback from its creator community. Since June 2024, the development team has focused on addressing common user requests regarding character realism, movement fluidity, and integrated speech capabilities. By prioritizing these specific 'if only' user feedback points, Kling AI aims to bridge the gap between creative vision and technical execution. These updates are designed to make AI-generated content feel more lifelike and expressive, providing creators with sophisticated tools to overcome previous limitations in digital synthesis. (source: https://x.com/Kling_ai/status/2064709166904594616)

04

Tiny AutoScientist Launches to Streamline Automated Research Loops

Adaption AI has officially introduced Tiny AutoScientist, a new tool engineered to optimize organizational research workflows. Building upon the foundational research loop technology developed for the original AutoScientist, this compact iteration is designed to help organizations expedite their discovery and experimentation processes. By automating complex technical workflows, the system aims to enhance efficiency, reduce manual oversight, and provide a scalable framework for integrating AI-driven insights into research pipelines. This release represents a significant step in democratizing advanced automated scientific capabilities, allowing teams to leverage sophisticated loop-based experimentation to achieve faster outcomes in their respective domains. (source: https://x.com/sarahookr/status/2064709595990532402)

huggingface

8 stories
01

Kwai Keye-VL-2.0 Technical Report

Kuaishou has released Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed for long-video understanding and agentic tasks. To address high computational costs in long contexts, Keye-VL-2.0 is the first to adapt DeepSeek Sparse Attention (DSA) to GQA-based multimodal architectures, enabling lossless processing of 256K contexts while activating only 3B parameters. The training integrates Cross-Modal Multi-Teacher On-Policy Distillation (MOPD) with Context-RL and Video-RL to mitigate catastrophic forgetting. Keye-VL-2.0 demonstrates state-of-the-art performance on benchmarks like TimeLens, Video-MME-v2, and LongVideoBench, and natively supports multi-agent collaboration with self-correction capabilities across tool, code, and search scenarios. (source: https://huggingface.co/papers/2606.10651)

02

Next Forcing: Causal World Modeling with Multi-Chunk Prediction

Researchers have introduced Next Forcing, a multi-chunk prediction (MCP) framework for causal world modeling that addresses the issues of slow training convergence and limited accuracy in autoregressive video generation. Next Forcing introduces an MCP objective that utilizes lightweight auxiliary modules to simultaneously denoise video chunks at multiple future temporal horizons. During training, the auxiliary modules provide dense multi-scale temporal supervision, speeding up convergence by 2.3 times and achieving a 93.1% relative improvement over LingBot-VA at 50 FPS. At inference, these modules enable 2x acceleration by predicting the next video chunk in parallel. The framework achieves state-of-the-art performance on the RoboTwin and PhyWorld benchmarks. (source: https://huggingface.co/papers/2606.11187)

03

EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents

Researchers proposed EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents designed to handle real-world task streams containing heterogeneous data distributions. To mitigate cross-dataset interference, EEVEE utilizes a router to partition incoming inputs into task clusters and assign them to suitable prompt configurations, optimizing this process via an interleaved router and prompt learning co-evolution strategy. Across benchmarks, the framework improves average scores by 10.38 points over Qwen3-4B-Instruct and 24.32 points over DeepSeek-V3.2, surpassing state-of-the-art methods GEPA and ACE by up to 37.2% and 48.2%, respectively, while maintaining single-benchmark efficiency. (source: https://huggingface.co/papers/2606.11182)

04

Decentralized Multi-Agent Systems with Shared Context

Researchers proposed Decentralized Language Models (DeLM), a multi-agent system (MAS) framework that decentralizes orchestration through parallel agents, a shared verified context, and a task queue. By allowing agents to asynchronously claim subtasks and write back updates to a common substrate, DeLM eliminates the communication bottlenecks of centralized controllers. On SWE-bench Verified, DeLM achieves top-tier performance across Avg.@1, Pass@2, and Pass@4, demonstrating gains of up to 10.5 percentage points over the strongest baseline while cutting costs per task by approximately 50%. It also improves average accuracy by up to 5.7 percentage points on the LongBench-v2 Multi-Doc QA benchmark across four major model families. (source: https://huggingface.co/papers/2606.10662)

05

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Researchers have introduced Workflow-GYM, a benchmark specifically designed to evaluate the performance of GUI agents on long-horizon, high-value professional workflows across diverse domains and specialized software. Unlike existing benchmarks that focus on simple applications, Workflow-GYM targets economically valuable, end-to-end tasks. Evaluation of state-of-the-art models reveals that even the strongest systems achieve success rates only slightly above 30%, indicating that long-horizon professional workflows remain highly challenging. The analysis reveals common failure modes, including workflow stage omission, error propagation, objective drift, and insufficient understanding of professional environments. (source: https://huggingface.co/papers/2606.11042)

06

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

Researchers identified that Chain-of-Thought (CoT) supervised fine-tuning systematically degrades long-context recall in hybrid linear-attention models, a phenomenon named attention amnesia. This degradation, which causes HypeNet-9B's performance on NIAH-S2@256K to drop from 67.2% to 9.4%, is attributed to CoT fine-tuning biasing attention gradients toward short-range patterns and disrupting projection weights. To address this, the authors propose QK-Restore, a training-free method that restores only the query and key projection parameters from the pre-SFT checkpoint. This intervention successfully restores long-context routing on HypeNet-5B, improving S3@256K scores from 65.4% to 76.4% while preserving reasoning performance. (source: https://huggingface.co/papers/2606.11052)

07

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

Researchers introduced ARM, a 7B discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a single next-token prediction framework. Utilizing a multi-objective semantic visual tokenizer, ARM aligns and reconstructs visual and textual sequences in a shared latent space. To boost preference-aligned behavior, the team applied reinforcement learning (RL) to optimize task-level objectives including visual quality and instruction adherence. The RL optimization significantly raised WISE scores from 0.50 to 0.56 and GEdit-Bench-EN scores from 5.75 to 6.68, while revealing unexpected cross-task synergies between text-to-image generation and instruction-guided editing. (source: https://huggingface.co/papers/2606.11188)

08

Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts

To adapt LLM agents to complex tasks without human-labeled datasets, researchers have introduced Retrospective Harness Optimization (RHO), a self-supervised method designed to optimize agent tools and workflows using past trajectories. RHO extracts a coreset of difficult historical tasks, re-runs them in parallel, and applies self-validation and self-consistency to select the best candidate harness updates via pairwise self-preference. Evaluated across three domains, a single round of RHO optimization successfully increased the pass rate on SWE-Bench Pro from 59% to 78% without any external grading, showing that the method alters behavior patterns and targets prior failure modes. (source: https://huggingface.co/papers/2606.05922)