NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-22ENGLISH EDITION
This issue
—
All time
—

AI Blog

5 stories
01

Introduces Daybreak Tools for Organization-Wide Security

OpenAI has introduced Daybreak, a new suite of security tools designed to automate organization-wide vulnerability patching and threat mitigation. The product line includes Codex Security and the GPT-5.5-Cyber model, which leverage advanced artificial intelligence to automate the identification, validation, and remediation of security bugs in digital systems. The suite aims to scale enterprise continuous vulnerability management capabilities, allowing organizations to maintain robust cyber defenses. (source: https://openai.com/index/daybreak-securing-the-world)

02

Samsung Electronics deploys ChatGPT Enterprise and Codex globally

Samsung Electronics has launched a global, corporate-wide deployment of OpenAI's ChatGPT Enterprise and Codex model to its workforce. This rollout represents one of OpenAI's largest enterprise partnerships to date and integrates conversational AI capabilities alongside intelligent code-generation features directly into Samsung's operational systems. The deployment aims to support software engineering and general productivity workflows across the company's global business units. (source: https://openai.com/index/samsung-electronics-chatgpt-codex-deployment)

03

Design Patterns for Evaluating Cybersecurity Capabilities in AI Agents

Eugene Yan released an analysis detailing the standard architectural design patterns used to evaluate cybersecurity capabilities within autonomous AI agents. The framework highlights four primary elements common to existing benchmarks like Cybench and CVE-Bench: a sandboxed target system inside Docker containers, calibrated challenge difficulty inputs, specialized development tools, and a deterministic grading system. To capture nuanced agent behaviors, the evaluation methodology employs a multi-level attack-chain pyramid to assign partial credits for discovery and exploitation tasks. (source: https://eugeneyan.com//writing/cybersecurity-evals/)

04

Patch the Planet Initiative Launched to Support Open Source Maintainers

OpenAI has announced the launch of Patch the Planet, a new Daybreak initiative aimed at helping open-source software maintainers detect and resolve security vulnerabilities. The program combines automated AI-driven analysis with professional human developer review to systematically identify, validate, and patch vulnerabilities in critical digital infrastructure. By streamlining the vulnerability lifecycle, the initiative seeks to alleviate security management burdens from global open-source projects. (source: https://openai.com/index/patch-the-planet)

05

Codex Maxxing For Long Running Work

OpenAI detailed how developer Jason Liu leverages the Codex model to carry context across extended workflows and manage multi-step engineering projects. The report details practical design patterns for context management and prompt strategies, allowing developers to maintain consistent performance and preserve task state when using large language models on complex, long-running coding operations that span beyond a single query. (source: https://openai.com/index/codex-maxxing-long-running-work)

Hacker News

6 stories
01

The text in Claude Code’s “Extended Thinking” output is not authentic

An analysis of Claude Code's "Extended Thinking" output has revealed that the displayed reasoning steps are not raw, internal monologues from the large language model. Instead, the visible outputs are post-processed and reconstructed narratives optimized for readability and safety alignment. This finding indicates that the chain-of-thought representations are curated justifications rather than direct, unmediated traces of neural computations, raising transparency concerns for AI interpretability. (source: https://patrickmccanna.net/the-text-in-claude-codes-extended-thinking-output-is-not-authentic/)

02

Sakana Fugu

Sakana AI has introduced Fugu, a novel approach designed to optimize and evolve foundational artificial intelligence models. By utilizing evolutionary computation and biologically inspired algorithms, the method targets neural network adaptation to discover optimal configurations without the computational overhead of massive retraining. This development seeks to bypass standard scaling bottlenecks in deep learning architectures. (source: https://sakana.ai/fugu/)

03

GLM 5.2 vs. Opus

A technical comparison has evaluated the performance and architectural trade-offs of the GLM 5.2 model against Claude 3 Opus. While Claude 3 Opus displays advanced qualitative reasoning and safety-aligned outputs, GLM 5.2 demonstrates strong competition in mathematical formulation, logical reasoning, and localized linguistic capabilities. The analysis outlines key differences in inference speeds, API reliability, and cost efficiency across academic datasets. (source: https://techstackups.com/comparisons/glm-5.2-vs-opus/)

04

Show HN: Oak – Git replacement designed for agents

An innovative version control system named Oak has been designed as a modern Git replacement optimized specifically for autonomous AI agents. Developed to minimize latency and context-management overhead, Oak utilizes virtual mounts that enable agents to work on code tasks without executing a full repository clone. This structure facilitates parallel task execution and increases operational speed, as shown during several months of self-bootstrapped development. (source: https://oak.space/oak/oak)

05

Prompt Injection as Role Confusion

A research paper has introduced a novel security framework that defines prompt injection vulnerabilities in large language models as cases of role confusion. The authors demonstrate that models fail to distinguish between their system-defined operational persona and external user inputs representing alternative personas. This study analyzes the boundaries of model identity and suggests new mitigation paths to construct robust barriers against adversarial manipulation. (source: https://role-confusion.github.io)

06

Moebius: 0.2B image inpainting model with 10B-level performance

Researchers from HUSTVL have developed Moebius, an efficient image inpainting model containing 0.2 billion parameters that delivers performance comparable to 10-billion parameter models. The model employs optimized architectural designs to reduce computational demands while achieving state-of-the-art results in context-aware generation, object removal, and pixel-level refinement. This release represents a significant advancement in bringing highly capable, resource-efficient generative vision models to edge devices. (source: https://hustvl.github.io/Moebius/)

Twitter

8 stories
01

Introducing GPT-5.5-Cyber for Advanced Cybersecurity Operations

OpenAI has officially launched GPT-5.5-Cyber, a specialized model designed to enhance digital defense systems. Engineered to actively remediate software vulnerabilities rather than just identifying them, the model has demonstrated state-of-the-art performance benchmarks on the CyberGym platform. Developed in collaboration with the United States government and the broader security community, the launch is supported by secondary initiatives: 'Patch The Planet' and the 'Codex Security' plugin, which automates threat modeling and patch generation. (source: https://x.com/sama/status/2069121360744550796)

02

Sakana AI Introduces Fugu Multi-Agent Orchestration System

Sakana AI has announced the launch of Fugu, a service designed to orchestrate multi-agent systems via a single unified API. The platform executes complex coding and engineering tasks by deploying a swarm of diverse, specialized AI agent instances and coordinating their resources. The system debuts alongside 'Fugu Ultra,' a model optimized for autonomous agent interactions, and features integration with Composer 2.5 for architectural scoping. It has also been deployed directly onto the Vercel AI Gateway for simplified cloud integration. (source: https://x.com/hardmaru/status/2068861824205025759)

03

GLM-5.2 Marks A New Era For Open Source Agentic Capabilities

The open-source community has released GLM-5.2, bringing advanced agentic capabilities to a non-proprietary model. The release delivers high-level complex reasoning and task execution, matching frontier-level intelligence parameters that were previously restricted to commercial architectures. Due to its potential disruptive impact on the open-source landscape, proponents and researchers are calling for proactive educational campaigns targeted at regulators and policymakers to establish balanced safety, governance, and deployment guidelines for high-capacity open models. (source: https://x.com/natolambert/status/2069073545632813193)

04

Google DeepMind Enters Strategic Research Partnership With A24

Google DeepMind has entered into a strategic research partnership with film production company A24 to bridge the gap between artificial intelligence and filmmaking. This collaboration aims to bring storytellers and filmmakers directly into the design process of generative AI video tools. By incorporating professional creative feedback, DeepMind intends to optimize models for real-world entertainment workflows while addressing artistic standards and ethical considerations in synthetic media production. (source: https://x.com/GoogleDeepMind/status/2069066675895337405)

05

Runway Updates Aleph 2.0 With Advanced Video Aspect Ratio Expansion

Runway has introduced an updated video scene expansion feature within its Aleph 2.0 platform. The utility leverages generative models to automatically adjust video aspect ratios, extending physical scene boundaries to fit new formats while maintaining composition and visual coherence. This release is targeted at content creators repurposing high-fidelity video assets across variable digital and social environments. Detailed training tutorials for implementing the tool have been released on the Runway Academy platform. (source: https://x.com/runwayml/status/2069147959896240320)

06

Kling 3.0 Turbo Integration Enhances Cinematic Video Generation on Picsart

Picsart has integrated the Kling 3.0 Turbo model to bring high-speed, generative video capabilities directly to its platform users. This integration facilitates the generation of cinematic videos featuring high-fidelity motion and synchronized native audio, while lowering the processing time required for rendering. The release also coincides with new Kling AI features focusing on user protagonist video generation, alongside strategic case studies highlighting digital creators scaling their reach using the generative engine. (source: https://x.com/Kling_ai/status/2068961124796563676)

07

Adaption AI Partners With NCS To Accelerate Sovereign AI In Asia-Pacific

Adaption AI has entered into a strategic partnership with technology services firm NCS to scale sovereign AI systems in the Asia-Pacific region. This initiative focuses on building and deploying customized AI infrastructures and adaptable intelligence systems aligned with regional data privacy standards and local regulatory laws. By combining NCS’s local market presence with Adaption AI's research capabilities, the partnership intends to help regional governments and enterprises safely utilize secure, generative architectures. (source: https://x.com/sarahookr/status/2069055188615446805)

08

Gary Marcus Questions SpaceX's AGI Progress Amid GPU Capacity Leasing

Industry critic Gary Marcus has questioned SpaceX's progress toward achieving Artificial General Intelligence following reports of the company securing large-scale GPU contracts to lease compute capacity. Marcus argues that prioritizing hardware leasing suggests a business model focused on infrastructure revenue rather than actual model development. This debate highlights ongoing friction and competition between specialized GPU cloud networks like CoreWeave and industrial giants entering the hardware resource layer. (source: https://x.com/GaryMarcus/status/2069093373877874992)

huggingface

8 stories
01

GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents

Researchers have introduced GateMem, a novel benchmark designed to evaluate memory governance in multi-principal, shared-memory AI agents. Spanning medical, educational, household, and office domains, GateMem measures long-horizon utility alongside access control across authorization boundaries and agent-facing active forgetting. Evaluation of multiple backbone models and retrieval methods reveals that none simultaneously achieve robust access control, reliable forgetting, and high utility. While long-context prompting achieved the highest governance scores, it incurred significant token costs, and retrieval-based alternatives frequently leaked unauthorized or deleted information (source: https://huggingface.co/papers/2606.18829).

02

MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision

Researchers proposed MemSlides, a hierarchical memory framework designed for personalized presentation generation agents. MemSlides divides memory into user profile memory for zero-round personalization, session-level working memory, and tool memory to preserve execution experience. Coupled with a slide-local revision system, the agent performs targeted edits on specific regions rather than regenerating entire decks. Controlled experiments demonstrate that separating persistent user profiles from working memory and tool-execution records significantly improves persona-alignment, localized editing, and preference preservation throughout multi-turn revision rounds (source: https://huggingface.co/papers/2606.17162).

03

Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models

Researchers have proposed Reflective Masking (RM), a post-training method that enables multi-turn masking and denoising for Mask Diffusion Models (MDMs). Unlike autoregressive models that rely on sequential generation for refinement, RM allows MDMs to execute local, iterative corrections on previous outputs. To leverage reasoning history, the authors introduce History Reference, a parameter-free mechanism utilizing intermediate denoising states. Applied across diverse modalities like text generation, Sudoku, and image editing, RM requires no structural changes and consistently outperforms traditional masking baselines (source: https://huggingface.co/papers/2606.16700).

04

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

Researchers introduced WorldLines, a project-driven benchmark designed to evaluate long-horizon embodied agents in household settings. Generating temporally extended traces of dialogues, actions, feedback, and environmental changes, the benchmark translates interactions into structured evaluations for Memory QA and Embodied Task Planning. Additionally, the authors developed ObsMem, an observer-grounded memory framework that tracks visibility-aware memories and action-native state histories. Initial baseline evaluations reveal persistent challenges in managing partial observability and translating long-term memory into cohesive embodied planning decisions (source: https://huggingface.co/papers/2606.18847).

05

SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG

Researchers developed SproutRAG, an attention-guided hierarchical retrieval-augmented generation (RAG) framework designed for long documents. SproutRAG structures sentences into semantically cohesive, progressively larger blocks using learned inter-sentence attention to build a binary chunking tree. The framework avoids lossy summarization and external LLM API calls during indexing or retrieval by learning which internal attention layers capture document structure. At retrieval time, a hierarchical beam search extracts candidates at multiple granularities, achieving a 6.1% average information efficiency improvement over baseline systems across four benchmarks (source: https://huggingface.co/papers/2606.18381).

06

MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval

Researchers introduced MCompassRAG, a metadata-guided retrieval framework that improves the speed and precision of paragraph-level search in RAG systems. To avoid semantic noise from mixing topics in dense chunk embeddings, MCompassRAG embeds topic metadata alongside text segments and trains a lightweight retriever via LLM-teacher distillation. During inference, it runs topic-aware retrieval without additional LLM calls. Across six complex retrieval benchmarks, MCompassRAG improved information efficiency by an average of 8.24% while executing with over five times lower latency than baseline efficient RAG approaches (source: https://huggingface.co/papers/2606.18508).

07

Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations

Researchers introduced the Call Playbook dataset, featuring five classification tasks derived from real-world B2B sales conversations, and proposed a novel knowledge extraction method for low-resource domains. Rather than concatenating verbose examples that increase context length, this approach distills in-context learning examples into compact, structured classification criteria and explicit task descriptions. The method achieved up to a 7% improvement in macro-averaged AUC and a 99% reduction in token usage compared to traditional in-context learning, while remaining robust against the performance degradation seen in alternative token compression baselines (source: https://huggingface.co/papers/2606.15641).

08

StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs

Researchers developed StylisticBias, a controlled benchmark for evaluating social bias in multimodal large language models (MLLMs). It features 25,000 generated images derived from 500 base faces, using 50 single-attribute variations per face to isolate specific visual cues while keeping identities fixed. Evaluating six MLLMs across 25 social judgment scenarios, the study found that age and body type drive identity-level effects, while fashion style drives major attribute-level shifts. Approximately 15 attributes account for nearly 80% of total variation, with model sensitivity strongest in socioeconomic and style-related judgments (source: https://huggingface.co/papers/2606.20527).