NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-18DEFAULT EDITION
This issue
—
All time
—

AI Blog

3 stories
01

Improving Health and Wellness Responses in ChatGPT

OpenAI has updated ChatGPT's health and wellness capabilities using the GPT-5.5 Instant model to deliver more reliable and structured health intelligence. This deployment enhances the conversational assistant's health queries through stronger logical reasoning, clearer communication, and deeper contextual understanding of user inputs. To ensure safety and accuracy, OpenAI evaluated these capabilities against standards informed by medical physicians. The enhancement aims to deliver dependable wellness guidance directly within the chat interface. (source: https://openai.com/index/improving-health-intelligence-in-chatgpt)

02

AI Reasoning Model Helps Diagnose Rare Genetic Diseases in Children

Researchers utilized an OpenAI reasoning model to assist in diagnosing rare genetic diseases, successfully identifying 18 new diagnoses in previously unsolved pediatric cases. The study demonstrates the clinical utility of advanced reasoning models in analyzing complex medical data to support physicians. By processing genomic information alongside patient symptoms, the AI model helped uncover critical diagnostic insights that had previously eluded medical teams, highlighting the potential of advanced reasoning systems to accelerate the identification of rare conditions. (source: https://openai.com/index/diagnose-rare-childhood-diseases)

03

Introducing LifeSciBench for Evaluating AI in Life Science Research

OpenAI announced the release of LifeSciBench, an expert-authored and expert-reviewed benchmark designed to evaluate how artificial intelligence systems handle real-world life science research tasks and decisions. This framework assesses the capabilities of AI models in executing complex scientific workflows and making critical research choices. By leveraging expert curation and review, LifeSciBench aims to set a rigorous standard for measuring AI performance in highly specialized biological and medical research domains. (source: https://openai.com/index/introducing-life-sci-bench)

Hacker News

8 stories
01

DeepSeek Introduces Vision

DeepSeek has officially integrated vision capabilities into its chat platform, allowing its model to perform multimodal tasks. Users can now upload images, diagrams, and document screenshots directly into the chat interface for visual reasoning, document parsing, and code generation from UI mockups. This release expands DeepSeek's utility beyond natural language processing to compete directly in the multimodal AI space. The integration has been widely discussed for its potential to bridge the gap between computer vision and language generation (source: https://chat.deepseek.com/).

02

Noam Shazeer Joins OpenAI

Prominent AI researcher Noam Shazeer has joined OpenAI, marking a significant talent transition in the generative artificial intelligence industry. Shazeer is widely recognized for co-authoring the seminal 2017 Transformer paper and previously co-leading Google's Gemini project. This shift comes after his brief return to Google via its acquisition of Character.ai. Shazeer's recruitment is expected to bolster OpenAI's core research capabilities in scaling neural network architectures and developing next-generation frontier models (source: https://twitter.com/NoamShazeer/status/2067400851438932297).

03

Midjourney Medical

Midjourney has announced its expansion into the healthcare sector with the launch of Midjourney Medical. Known primarily for its artistic image synthesis, the company is adapting its proprietary diffusion models and visual generation technologies for clinical and medical imaging demands. The initiative focuses on high-fidelity anatomical visualization, medical illustration, and diagnostic assistance frameworks. This move signals a strategic shift of generative AI technology into specialized domains where precise spatial comprehension is critical (source: https://www.midjourney.com/medical/blogpost).

04

We built a persistent agent memory layer on Elasticsearch with 0.89 recall

Elasticsearch has built and demonstrated a persistent agent memory layer that achieves a 0.89 recall rate. This memory system allows autonomous AI agents to maintain long-term context and execute workflows by indexing past interactions and knowledge. The architecture combines semantic vector search with keyword search to address common LLM challenges like memory decay and context window limitations. This development offers a scalable approach to building stateful production-grade agents (source: https://www.elastic.co/search-labs/blog/agent-memory-elasticsearch).

05

Launch HN: TesterArmy (YC P26) – Agents that test web and mobile apps

TesterArmy has launched an agentic quality assurance platform designed to automate end-to-end checks for web and mobile applications. The system utilizes AI agents to interpret test scenarios written in natural language, executing them programmatically without requiring static testing scripts. It also integrates with external coding agents, enabling automated pipelines to manage and run test suites. This launch aims to streamline software development workflows and reduce the manual overhead of QA maintenance (source: https://tester.army).

06

The Token Compression Illusion: Why I'm Skeptical of RTK

This technical critique analyzes Recurrent Token Compression (RTK) techniques, challenging the assumption that they offer a viable path to large language model efficiency. The author argues that the computational overhead of compression and reconstruction can offset theoretical latency benefits. Furthermore, the loss of critical semantic details during compression degrades performance in high-precision tasks like code generation. The article suggests alternative architectures like structured attention or state space models for long-context tasks (source: https://mroczek.dev/articles/the-token-compression-illusion-why-im-skeptical-of-rtk/).

07

Local Qwen isn't a worse Opus, it's a different tool

This technical analysis compares the performance of local open-weights AI models, such as Qwen, to proprietary APIs like Claude 3 Opus. The author argues that local language models provide unique operational advantages, including guaranteed data privacy, lower latency, and zero-marginal-cost execution for specialized tasks. Rather than replacing massive frontier models, local models serve as specialized tools optimized for local environments. The article advocates for hybrid architectures that utilize both local and cloud-based models (source: https://blog.alexellis.io/local-ai-is-not-opus/).

08

Migrate from OpenClaw

Nous Research has released a comprehensive guide outlining the migration path from the OpenClaw framework to its Hermes Agent ecosystem. The documentation provides instructions on transitioning LLM-based agent applications, custom workflows, and prompts over to the Hermes platform. By migrating, developers can take advantage of improved system scalability, enhanced resource management, and deeper integrations with modern Large Language Models and open-source agent tooling (source: https://hermes-agent.nousresearch.com/docs/guides/migrate-from-openclaw).

Twitter

8 stories
01

MCP Introduces Enterprise-Managed Authentication For Centralized Connector Access

The Model Context Protocol (MCP) team has officially launched support for the Enterprise-Managed Authentication extension to streamline centralized connector access across organizations. This update enables administrators to centrally authorize and manage MCP connectors, ensuring that tools, Identity Providers (IdP), and data sources are seamlessly integrated upon initial user login. This protocol addition standardizes cross-platform communication and simplifies identity management, resolving redundant decentralized per-application OAuth setups to reduce security risks and administrative overhead (source: https://x.com/ClaudeDevs/status/2067655887662272723).

02

Claude Integrates Ecosystem Connectors for Enhanced Workflow Automation

Anthropic has launched a new beta integration suite for Claude that enables direct connectivity with major tools including Asana, Atlassian, Canva, Figma, Granola, Linear, and Supabase. This update ensures access is synchronized across Claude Chat, Claude Code, and Cowork environments to streamline collaborative, agentic workflows. Additionally, Anthropic introduced a new feature allowing team and enterprise users to convert ongoing Claude Code sessions into shareable, real-time updated web page Artifacts (source: https://x.com/ClaudeDevs/status/2067655890166247690).

03

Noam Shazeer Joins OpenAI Team to Advance Artificial Intelligence Research

OpenAI has officially added highly respected researcher Noam Shazeer to its research team to advance its development of high-performance large language models and general intelligence. Shazeer is a key pioneer behind foundational deep learning architectures, notably the Transformer model, during his tenure at Google. This critical hire strengthens OpenAI's roster of top-tier talent as it builds next-generation frontier systems and optimizes complex reasoning capabilities (source: https://x.com/polynoamial/status/2067402274662744465).

04

Sam Altman Announces Collaboration With Noam Brown At OpenAI

Sam Altman, CEO of OpenAI, announced he is officially collaborating with researcher Noam Brown, a long-standing recruitment goal since the initial formation of the company. The partnership, marked on the tenth anniversary of their professional aspirations, aims to expand OpenAI's research in complex reasoning capabilities and advanced reinforcement learning strategies for large scale models (source: https://x.com/sama/status/2067427421083652131).

05

GLM-5.2 Released Featuring Multi-Head Latent and Sparse Attention

The new open-weight GLM-5.2 model has been released, demonstrating strong natural language processing capabilities. Architecturally, the model builds upon GLM-5 and GLM-5.1 foundations by incorporating Multi-head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA) to optimize processing speed and resource efficiency. Evaluated on PostTrainBench, GLM-5.2 showcased advanced post-training alignment. Additionally, shape-rotator benchmark results reveal GLM-5.2 outperforming Gemini in specialized spatial reasoning tests (source: https://x.com/rasbt/status/2067612153020838055).

06

Unanticipated Drivers Behind The Second Wave Of PPO In The LLM Era

John Schulman published observations on the resurgence of Proximal Policy Optimization (PPO) in Large Language Model (LLM) alignment, identifying key factors not anticipated in original research. PPO's importance-ratio objective successfully mitigates numeric errors, forward-pass noise, and asynchronous training bias. New findings on the clipping objective reveal mechanisms that impact entropy, as detailed in the DAPO paper, explaining why PPO remains a fundamental technique for instruction following, safety, and stable optimization (source: https://x.com/johnschulman2/status/2067410492008841643).

07

Kling 3.0 Turbo Launches With Enhanced Performance And Visual Fidelity

Kling AI has launched Kling 3.0 Turbo, a model designed to accelerate generative video inference speeds while improving visual quality, including lip-sync accuracy and sharpness. Optimized for production-ready efficiency, the model provides improved prompt adherence and multi-shot consistency. The Kling 3.0 Turbo engine has officially integrated with third-party generative platforms, including Fal AI, SeaArt, Clipfly, Fotor, Glam AI, Runware, and Morphic, delivering lower rendering costs across ecosystems (source: https://x.com/Kling_ai/status/2067565582141309059).

08

Luma Labs Introduces Skills For Repeatable Creative Workflows

Luma Labs has announced the launch of Luma Skills, a workflow capability designed to simplify creative processes within its platform. This update allows users to transform static assets and brand identity into reusable, scalable workflows that generate hundreds of product-accurate concepts. By automating the design phase while preserving consistent styling across projects, the system allows creative teams to establish structured generative pipelines and execute complex operations without reconfiguring parameters from scratch (source: https://x.com/LumaLabsAI/status/2067655789465292977).

huggingface

8 stories
01

Native Active Perception as Reasoning for Omni-Modal Understanding

Researchers introduced OmniAgent, a native omni-modal agent that frames video understanding as a partially observable Markov decision process. Instead of processing entire videos uniformly, OmniAgent dynamically executes actions to distill audio-visual cues into a persistent textual memory. To support this framework, the authors introduced Agentic Supervised Fine-Tuning and Agentic Reinforcement Learning with Turn-aware Adaptive Uncertainty Rescaled Advantage (TAURA). The resulting 7B OmniAgent exhibits test-time scaling and achieves state-of-the-art results across ten benchmarks, outperforming the much larger Qwen2.5-VL-72B on LVBench by a margin of 50.5% to 47.3% (source: https://huggingface.co/papers/2606.19341).

02

Sumi: Open Uniform Diffusion Language Model from Scratch

Researchers released Sumi, a 7B parameter uniform diffusion language model (UDLM) trained from scratch on 1.5 trillion tokens. While masked diffusion and autoregressive language models have been scaled extensively, uniform diffusion language models have lacked open-source, large-scale reference points. Sumi matches autoregressive baselines trained on similar token budgets in coding, reasoning, and knowledge benchmarks, though it lags on commonsense metrics due to its education-centric data mixture. The authors have open-sourced the model weights, checkpoints, and complete training configurations to facilitate further study of uniform diffusion dynamics (source: https://huggingface.co/papers/2606.19005).

03

CEO-Bench: Can Agents Play the Long Game?

Researchers introduced CEO-Bench, a benchmark designed to evaluate how well language model agents handle long-horizon, multi-step planning tasks under uncertainty. The benchmark simulates running a startup company for 500 days via a programmable Python interface, forcing agents to balance interconnected variables like pricing, marketing, and budgeting. Agents are evaluated on their ability to analyze noisy databases and code custom customer simulations. Most state-of-the-art models failed to maintain their starting budget, with only Claude Opus 4.8 and GPT-5.5 finishing above the initial $1M balance, though neither consistently generated a profit (source: https://huggingface.co/papers/2606.18543).

04

STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability

Researchers proposed STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability) to address the issue of policy entropy collapse in reinforcement learning with verifiable rewards algorithms like GRPO. Based on a gradient analysis of token-level entropy dynamics, STARE identifies critical token subsets using batch-internal surprisal quantiles and reweights their advantages during training. Across model scales from 1.5B to 32B and multiple reasoning tasks, STARE stabilizes policy entropy. It outperformed DAPO and other baselines on AIME24 and AIME25 by 4% to 8% in average accuracy while maintaining balanced exploration (source: https://huggingface.co/papers/2606.19236).

05

The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL

Researchers proposed Discriminator-Guided RL (DRL), a reinforcement learning framework that corrects the structural mismatch found in traditional score- and flow-matching models. DRL trains a discriminator in a pretrained representation space to distinguish training data from base-model samples, utilizing its logit as a reward in KL-regularized RL. Applied across models like SiT and JiT, DRL improved sample visual realism and structure without human preference data. For instance, DRL reduced guidance-free FID on SiT from 9.38 to 2.62 and enhanced the Pareto frontier between human preference alignment and overall image fidelity (source: https://huggingface.co/papers/2606.19162).

06

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Researchers introduced MaineCoon, a 22B parameter real-time audio-visual autoregressive model optimized as a social world model. MaineCoon achieves sub-second human interactions with streaming generations at up to 47.5 frames per second on a single GPU. The model uses self-resampling, cross-modal representation alignment, domain-aware preference optimization, and reinforced online-policy distillation. To sustain extended generation sequences without drift, the authors developed an agentic streaming inference framework featuring cache management and prompt planning, establishing a new baseline for low-latency, long-horizon multimedia generation (source: https://huggingface.co/papers/2606.17800).

07

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

Researchers introduced RODS (Reward-driven Online Data Synthesis), an online training framework designed to prevent the depletion of informative training samples in multi-turn tool-use RL. RODS leverages progress reward variance during GRPO rollouts as a zero-cost boundary detector to identify complex samples near the agent's capability limits. It then synthesizes new multi-turn task variants via a skill-aligned resampling pipeline. Starting with 400 human-derived seeds and maintaining an active training pool of ~800 samples, RODS matched the performance of a static 17K offline dataset while requiring 20 times fewer trajectory samples (source: https://huggingface.co/papers/2606.19047).

08

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Researchers revealed a vulnerability in sparse autoencoder (SAE) latent-space defenses, demonstrating that clamping target features fails to completely eliminate suppressed behaviors. The authors formulated this phenomenon as a post-intervention recovery problem and designed encoder-orthogonal updates to isolate recovery paths. Across experiments including refusal steering and unlearning, the stress test showed models could recover forbidden behavior. In refusal steering settings, models achieved a 95.8% recovery rate while keeping targeted feature drift to 0.131. Attribution analysis traced the recovered behavior back to the SAE reconstruction residual (source: https://huggingface.co/papers/2606.18322).