NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-24DEFAULT EDITION
This issue
—
All time
—

AI Blog

2 stories
01

and Broadcom Introduce Jalapeño Inference Chip

OpenAI has partnered with Broadcom to introduce Jalapeño, a custom AI silicon chip designed specifically for large language model inference. The new processor focuses on optimizing the deployment of advanced generative models by delivering improvements in operational performance, energy efficiency, and scalability across enterprise AI systems. This joint effort is a strategic development in managing the hardware demands of next-generation artificial intelligence technologies, specifically tailored for inference workloads to help scale model deployment. (source: https://openai.com/index/openai-broadcom-jalapeno-inference-chip)

02

Building Shared Standards for Advanced AI

OpenAI announced a partnership with the Appia Foundation to help build shared standards, robust evaluation frameworks, and safety practices for advanced AI. The initiative aims to establish unified benchmarks and cooperative guidelines across the artificial intelligence sector to address emergent risks. By collaborating on global safety standards, the partnership seeks to ensure safe development and alignment as AI capabilities continue to progress. (source: https://openai.com/index/helping-build-shared-standards-for-advanced-ai)

Hacker News

7 stories
01

Computer use in Gemini 3.5 Flash

Google has introduced a new computer use capability for its Gemini 3.5 Flash model, enabling active desktop automation. This feature allows the model to interact directly with digital interfaces by interpreting visual on-screen elements, executing keystrokes, and controlling cursor movements and clicks. Designed as a lightweight, low-latency model, Gemini 3.5 Flash can automate multi-step workflows across various software applications in real time. This release marks a shift from passive conversational assistants to active agentic operators capable of performing complex tasks directly within standard operating systems. (source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-computer-use-gemini-3-5-flash/)

02

Qwen-AgentWorld: Language World Models for General Agents

Researchers have introduced Qwen-AgentWorld, a framework that utilizes Language World Models to power general-purpose AI agents. By employing large language models to construct interactive, natural language simulation environments, the framework acts as a digital sandbox where agents can plan, reason, and anticipate outcomes before execution. This approach mitigates the risks and high costs of real-world trials by allowing agents to simulate complex decision consequences. The framework demonstrates notable performance improvements in multi-step planning, task success rates, and generalized decision-making across several evaluation domains. (source: https://arxiv.org/abs/2606.24597)

03

NSA lost access to Mythos amid Anthropic dispute

The National Security Agency has lost access to Mythos, a specialized artificial intelligence tool developed by Anthropic, following an active dispute between the agency and the AI startup. Mythos was previously used by the agency to assist in national security analysis and intelligence gathering. The service disruption underscores growing tension between commercial AI developers and government bodies regarding data privacy, model safety, and regulatory compliance. The dispute reflects a broader trend of commercial frontier model vendors exercising strict control over how government agencies deploy proprietary technologies. (source: https://www.nytimes.com/2026/06/23/us/politics/nsa-lost-access-anthropic-tool.html)

04

For Most of the World, Open-Source AI Is the Only Way Forward

An industry report argues that open-source artificial intelligence is the primary mechanism for equitable global technology adoption, serving as an alternative to proprietary models. While commercial models managed by a few technology corporations dominate the current landscape, they introduce significant cost and accessibility barriers. Open-source models allow developers worldwide to customize systems for local languages, regional cultures, and specific regulatory standards. By democratizing access to high-performance base models, open-source initiatives foster localized innovation and prevent market monopolies, ensuring technology benefits are distributed globally. (source: https://techstrong.ai/articles/for-most-of-the-world-open-source-ai-is-the-only-way-forward/)

05

Haystack: Open-Source AI Framework for Production Ready Agents, RAG

Deepset has developed Haystack, an open-source AI orchestrator designed to build production-ready applications powered by large language models, retrieval-augmented generation (RAG), and autonomous agents. The framework provides modular pipelines for complex tasks such as question answering, semantic search, and document ingestion. By offering structured components to connect vector databases, transformer models, and external APIs, Haystack simplifies the deployment of high-performance cognitive search engines. It is tailored for enterprise scalability, making it easier for developers to build robust, agentic AI solutions. (source: https://haystack.deepset.ai/)

06

RubyLLM: A Ruby framework for all major AI providers

RubyLLM has launched an open-source, unified framework designed to streamline large language model API integrations for Ruby developers. The framework provides a standardized interface that abstracts the distinct characteristics of different AI provider endpoints. By reducing boilerplate code, RubyLLM simplifies the process of switching between underlying models and providers with minimal configuration changes. The project aims to drive broader adoption of advanced generative AI capabilities within the Ruby and Ruby on Rails communities, offering a viable alternative to Python-dominated development ecosystems. (source: https://rubyllm.com/)

07

Big AI labs are hiring philosophers

Major artificial intelligence laboratories are hiring academic philosophers to address critical ethics, value alignment, and conceptual dilemmas in frontier AI development. As large language models and autonomous systems grow more capable, laboratories face complex issues regarding artificial agency, decision-making constraints, and existential risk. Philosophers work alongside engineering teams to translate abstract ethical values into concrete algorithmic guardrails. Additionally, these interdisciplinary experts assist in structuring advanced benchmarks to evaluate the moral and logical reasoning capacities of generative systems, ensuring they remain beneficial and aligned with human values. (source: https://www.economist.com/science-and-technology/2026/06/24/why-big-ai-labs-are-hiring-so-many-philosophers)

Twitter

8 stories
01

OpenAI Introduces Significant Performance Enhancements To GPT-5.5 Instant

OpenAI announced a series of performance improvements to its GPT-5.5 Instant model. This update focuses on refining the model's conversational experience, making interactions more engaging and nuance-focused for daily use cases. The fine-tuning process targets natural responsiveness and high-speed processing for interactive artificial intelligence applications. Users are encouraged to integrate the updated model immediately to experience the refined conversational style and performance characteristics in their active development workflows. (source: https://x.com/gdb/status/2069845493199597944)

02

Open-Source Model Achieves Record-Breaking Performance on ARC-AGI Benchmark

A newly released open-source model has achieved the strongest performance to date on the Abstraction and Reasoning Corpus (ARC-AGI) benchmark. Created by Francois Chollet, ARC-AGI is designed to measure general intelligence by testing a model's ability to solve unseen logic puzzles rather than relying on memorized data. This milestone indicates that open-source optimization techniques and architectural improvements are narrowing the capabilities gap with large-scale proprietary systems, marking a notable shift in non-proprietary reasoning performance. (source: https://x.com/fchollet/status/2069858556552298519)

03

Introducing Jalapeño: A Custom-Built Accelerator for LLM Inference

Developers have introduced Jalapeño, a custom hardware accelerator designed specifically to optimize Large Language Model (LLM) inference. Developed over a nine-month period, the project leveraged internal AI models during its design phase to accelerate engineering timelines. Early performance metrics demonstrate highly efficient power consumption, delivering superior performance per watt compared to traditional architectures. This release highlights a broader industry shift toward vertical hardware integration to address computational bottlenecks in LLM production environments. (source: https://x.com/gdb/status/2069809298612621629)

04

Sakana AI Launches Fugu Ultra Model on OpenRouter Platform

Sakana AI has officially released its Fugu Ultra model on the OpenRouter platform, expanding access to its specialized model architecture. This deployment allows developers to utilize Fugu Ultra through OpenRouter's modular interface, promoting multi-model systems and decentralized AI infrastructure. The launch has been integrated into OpenRouter's ecosystem, facilitating broader testing, deployment, and interoperability. This represents a strategic effort to challenge centralized AI access patterns by providing flexible, high-performance alternatives. (source: https://x.com/hardmaru/status/2069818125265293684)

05

Bitrobot Releases Massive Humanoids-in-the-Wild Teleop Dataset

Bitrobot has released HIW-500, an extensive humanoid teleoperation dataset featuring 500 hours of real-world home interaction data. Recorded in uncontrolled domestic environments, this release represents the largest collection of human-controlled humanoid robotic movement data to date. The dataset serves as a resource to accelerate the development of imitation and behavioral learning algorithms, enabling humanoid robots to navigate complex, human-centric spaces with improved physical control systems. (source: https://x.com/Thom_Wolf/status/2069817782280094073)

06

HalluHard Leaderboard Integrates GLM-5.2 With Adaptive Reasoning Capabilities

The HalluHard benchmark has integrated GLM-5.2 into its official evaluation leaderboard to assess the model's adaptive thinking and reasoning processes. Notable for dynamically allocating computational effort during complex logic tasks, GLM-5.2 represents a growing industry trend toward evaluating reasoning efficiency and logical precision rather than simple output accuracy. This benchmark update aims to provide researchers with critical metrics on model hallucination rates, logical inference, and the efficacy of adaptive architectures. (source: https://x.com/GaryMarcus/status/2069633163786367361)

07

Runway Launches Automated Ad Localization Feature for Global Marketing

Runway has launched an automated ad localization feature designed to streamline global marketing operations. The tool takes a single advertisement image and automatically generates localized variants in multiple languages. By automating transcreation and translation, this tool reduces the manual effort required for multi-market campaigns, allowing creative teams to deploy assets globally. This update highlights Runway's strategic focus on integrating generative artificial intelligence tools into production-ready enterprise workflows. (source: https://x.com/runwayml/status/2069796562805440964)

08

Yann LeCun Criticizes Current AI Scaling Laws and Large Language Models

Meta's Chief AI Scientist Yann LeCun has challenged prevailing AI scaling practices, arguing that brute-force scaling of current autoregressive Large Language Models is insufficient for human-level intelligence. LeCun highlighted that generative models fail to grasp physical reality, reasoning, or long-term planning due to the absence of a genuine world model. He advocates shifting research priorities toward World Model architectures and Joint Embedding Predictive Architectures (JEPA) to bypass current structural limits. (source: https://x.com/ylecun/status/2069765820121485385)

huggingface

8 stories
01

OpenThoughts-Agent: Data Recipes for Agentic Models

Researchers have released OpenThoughts-Agent, a fully open data curation pipeline for training agentic models. To address the challenge of generalizing across diverse agentic tasks, the project ran over 100 ablation experiments, culminating in a training dataset of 100,000 examples. Fine-tuning a Qwen3-32B model on this dataset produced a 44.8% average accuracy across seven agentic benchmarks, representing a 3.9 percentage point improvement over the previous open-data baseline, Nemotron-Terminal-32B. The training sets, pipeline, and models have been made publicly available. (source: https://huggingface.co/papers/2606.24855)

02

Qwen-AgentWorld: Language World Models for General Agents

Researchers introduced Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, which are language world models capable of simulating agentic environments across 7 domains using long chain-of-thought reasoning. The models were trained on over 10 million real-world environment interaction trajectories through a three-stage pipeline: continued pre-training, supervised fine-tuning, and reinforcement learning with hybrid rewards. Evaluated on the new AgentWorldBench benchmark, Qwen-AgentWorld outperformed existing frontier models. This system supports scalable environment simulation for agentic reinforcement learning and serves as a downstream performance warm-up. (source: https://huggingface.co/papers/2606.24597)

03

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

Researchers introduced NatureBench, a cross-discipline benchmark containing 90 scientific coding tasks derived from Nature-family publications to evaluate AI agents. It utilizes NatureGym to construct standardized, containerized environments for each task. Under strict testing without web-search capabilities, the strongest evaluated frontier agent achieved the state-of-the-art criterion on only 17.8% of the tasks. The analysis revealed that successes were primarily due to translating tasks into standard supervised learning formats rather than original scientific discovery, with failures mostly caused by bad methodological choices and compute constraints. (source: https://huggingface.co/papers/2606.24530)

04

MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization

Researchers presented MobileForge, an annotation-free adaptation system for mobile GUI agents designed to reduce the cost of manually labeling tasks or rewards. It combines MobileGym, which grounds task generation and rollout evaluation in real-world mobile app interaction, with Hierarchical Feedback-Guided Policy Optimization (HiFPO) to convert multi-level feedback into step-level GRPO updates. Using only automatically generated data, MobileForge adapted Qwen3-VL-8B to achieve 67.2% Pass@3 on AndroidWorld, while the adapted ForgeOwl-8B model established a state-of-the-art open-data benchmark of 77.6% Pass@3. (source: https://huggingface.co/papers/2606.19930)

05

MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management

To address ReAct-style prompt dilution in long-horizon mobile agent tasks, researchers developed MemGUI-Agent, an end-to-end mobile GUI agent utilizing Context-as-Action (ConAct) for proactive context management. ConAct handles context modification as standard actions, maintaining structured and compressed representations of folded histories and UI states. The authors introduced MemGUI-3K, a dataset of 2,956 trajectories containing full annotations for training. The resulting MemGUI-8B-SFT model achieved leading open-data 8B performance on the MemGUI-Bench evaluation suite and generalized successfully to the out-of-distribution MobileWorld benchmark. (source: https://huggingface.co/papers/2606.19926)

06

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction

Researchers introduced AOHP (Android Open Harness Project), an open-source, OS-level agent harness built on the Android Open Source Project (AOSP) to natively support AI agents as first-class OS actors. AOHP implements specialized system mechanisms for personalized service composition, optimized agent interfaces, and secure information flows while maintaining the legacy Android ecosystem. In preliminary experiments, AOHP improved agent task completion rates by 21.12%, reduced execution costs by 51.55% in token consumption, and demonstrated strong compliance with system security policies. (source: https://huggingface.co/papers/2606.23449)

07

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning

To prevent LLM agents from storing erroneous trajectories as successful experiences, researchers introduced EDV (Execute-Distill-Verify), a collaborative experience learning framework. In the Execute stage, multiple heterogeneous agents generate diverse candidate trajectories. In the Distill stage, a third-party agent analyzes these inputs to minimize execution bias. Finally, in the Verify stage, the execution group validates the candidate experiences using a consensus mechanism before storing them. EDV achieved consistent improvements over baseline models when evaluated on the tau2-bench, Mind2Web, and MMTB benchmarks. (source: https://huggingface.co/papers/2606.24428)

08

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Researchers proposed the Holistic Data Scheduler (HDS), an online data mixing framework that formulates data scheduling during LLM pre-training as a reinforcement learning problem. Running in a continuous control space, HDS utilizes the Soft Actor-Critic algorithm combined with a multi-objective reward function incorporating data quality, inter-domain influence, and model weight norms. On The Pile benchmark, HDS reached equivalent final validation perplexity while using 44% fewer training iterations than the next best method, and achieved a 7.2% improvement on 0-shot MMLU. (source: https://huggingface.co/papers/2606.24133)