NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-15DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

GPT-Red Unlocks Automated Red Teaming and Self-Improvement

OpenAI has introduced GPT-Red, an automated red teaming system designed to enhance the safety, alignment, and robustness of large language models against prompt injection attacks. By utilizing self-play methodologies, the system enables AI models to iteratively probe their own systems, uncover hidden vulnerabilities, and execute defensive hardening. This framework automates the detection of edge cases and malicious inputs to streamline safety alignment before models are publicly deployed. (source: https://openai.com/index/unlocking-self-improvement-gpt-red)

Hacker News

8 stories
01

Inkling: Our Open-Weights Model

Thinking Machines has released Inkling, a brand-new open-weights model designed to advance accessibility in the artificial intelligence development landscape. By providing open access to the model's weights, this initiative allows developers and researchers to customize, fine-tune, and deploy the system for specialized tasks. This release aims to provide a robust alternative for enterprises looking to maintain control over proprietary data and run local deployments on their own architectures. (source: https://thinkingmachines.ai/news/introducing-inkling/)

02

Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU

A technical demonstration by NeoMind Labs showcases Google's Gemma 4 26B model running at a practical speed of 5 tokens per second on a 13-year-old Intel Xeon processor without GPU acceleration. The implementation utilizes advanced CPU execution, model quantization, and memory management optimizations. This performance demonstrates that large-scale open-weights models can be deployed on low-cost, legacy enterprise server hardware instead of expensive modern GPU clusters, democratizing local model inference. (source: https://www.neomindlabs.com/2026/06/08/running-gemma-4-26b-at-5-tokens-sec-on-a-13-year-old-xeon-with-no-gpu/)

03

Launch HN: Coasty (YC S26) – An API for computer-use agents

Nitish and Prateek have launched Coasty, a Y Combinator S26-backed startup that provides an API and platform designed for computer-use agents. The service automates workflows within legacy desktop and web applications that lack traditional APIs. Developers trigger actions via natural-language instructions or direct API calls. Coasty runs these tasks inside custom virtual machines or secure browser configurations, navigating interfaces by analyzing screenshots and simulating keyboard or mouse inputs. (source: https://coasty.ai/docs)

04

OpenAI loses trademark dispute at EU court

OpenAI has lost a critical trademark dispute at the European Union Intellectual Property Office or European court system, representing a major regulatory setback. The ruling prevents OpenAI from securing exclusive proprietary naming rights for its signature AI technologies in the European market. This decision sets a legal precedent that allows competitors to use similar descriptive terms in their marketing and challenges OpenAI's brand monopolization across EU member states. (source: https://dpa-international.com/economics/urn:newsml:dpa.com:20090101:260715-930-389143/)

05

Open-source memory for coding agents, synced over SSH

An open-source project called Deja-Vu introduces a specialized memory layer for autonomous coding agents. The system synchronizes and persists agent memory across environments using secure SSH connections, resolving context loss in disconnected sessions. By maintaining state and previous code modifications, Deja-Vu improves agent performance without relying on proprietary third-party cloud services, giving developers full privacy and control over their codebases. (source: https://github.com/vshulcz/deja-vu/)

06

A General Goal-Conditioned Minecraft Model

Researchers have introduced a general goal-conditioned model designed specifically for the open-ended virtual environment of Minecraft. The model utilizes hierarchical decision-making and pre-trained multi-modal representations to translate abstract player instructions into sequential in-game actions. By moving beyond traditional single-task reinforcement learning, this framework tackles key challenges in long-horizon planning, spatial reasoning, and object manipulation, demonstrating robust behaviors for autonomous virtual agents. (source: https://pantograph.com/journal/pan-1)

07

DSLs Enable Reliable Use of LLMs

Martin Fowler explores how Domain-Specific Languages (DSLs) can be used to constrain the output space of Large Language Models (LLMs) to make them more reliable in software engineering. By enforcing a predefined DSL grammar, developers can prevent issues such as model hallucinations, formatting errors, and syntactic invalidity. This integration bridges natural language reasoning with deterministic execution systems, enabling stable automated agentic workflows. (source: https://martinfowler.com/articles/llm-and-dsls.html)

08

J-space comparisons across open models

A new interactive visualization tool has been developed to present detailed J-space comparisons across various open-source large language models. The platform maps high-dimensional model representations and performance metrics into an intuitive spatial representation. By studying these plots, researchers can evaluate how distinct training methodologies, fine-tuning datasets, and architectural tweaks alter a model's latent representation space and relative performance clustering. (source: https://eliebak.com/viz/jspace-open-v2)

Twitter

8 stories
01

Inkling Model Released With Open Weights And Enhanced Agentic Capabilities

John Schulman has announced the launch of Inkling, a performant open weights foundation model optimized for agentic workflows, coding, and reasoning tasks. The model's training process began with initial pretraining during the winter, followed by concentrated alignment and agentic optimization phases starting in mid-January. Inkling is now available to developers on the Tinker platform, where it operates as a modular, adaptable baseline for building specialized downstream applications. The release marks a strategic effort by Thinky Machines to provide a unified, customizable post-training service layer across its integration stack. This launch was also discussed by other industry practitioners on social media, emphasizing its utility as an accessible baseline. (source: https://x.com/johnschulman2/status/2077460227327467982)

02

Addressing The Validation Bottleneck In AI-Driven Scientific Discovery

Google DeepMind has released an analysis detailing the role of autonomous AI agents in scientific discovery and the bottleneck that occurs during physical validation. While deep-learning agents excel at generating novel hypotheses and designing complex experimental pipelines digitally, transitioning these concepts to physical laboratory tests remains a major rate-limiting step. To resolve this friction, DeepMind outlines four strategic priorities for policymakers and researchers. These recommendations focus on building reliable validation infrastructures and standardized evaluation frameworks to safely integrate automated research agents into real-world laboratory workflows. The findings reflect a broader shift toward structured, agent-led experimentation in material sciences and drug discovery. (source: https://x.com/GoogleDeepMind/status/2077372568143642972)

03

Gemini Spark Expands Availability To More Countries And Languages

Google has expanded the global availability of Gemini Spark, its personal AI assistant feature, to Google AI Ultra subscribers across additional international regions and languages. Designed for deep integration within the Google product ecosystem, the tool assists premium subscribers with complex personal productivity tasks through localized generative AI interactions. By rolling out the feature to a wider international user base, Google aims to scale its premium consumer-facing assistant technology and strengthen its market position against competing generative AI service providers. This rollout marks an iterative step in localized deployment of Google's flagship multimodal assistant capabilities. (source: https://x.com/Google/status/2077465898609197344)

04

GPT-Red Enhances Security Against Prompt Injection Vulnerabilities

GPT-Red has been introduced as an automated red teaming tool designed to identify and mitigate prompt injection vulnerabilities in large language models. The software enables developers to systematically stress-test production models against sophisticated adversarial inputs, offering a scalable defense framework for enterprise application deployment. By automating the threat detection process, the tool helps security engineers harden applications and maintain model alignment standards without requiring manual heuristic interventions. This launch represents a notable technical advancement in automated vulnerability scanning, addressing one of the primary infrastructure security concerns for production-grade large language models. (source: https://x.com/gdb/status/2077464463251554327)

05

Analysis of Recursive Self Improvement in Modern AI Models

A new research paper evaluating recursive self-improvement in artificial intelligence has challenged the feasibility of rapid takeoff scenarios. By analyzing model optimization loops, the study finds that machine intelligence gains scale at a measured pace defined by the progression rate of [Intelligence]^0.07. This empirical performance data suggests that self-optimizing feedback loops produce gradual, predictable advancement rather than the runaway exponential growth often predicted in long-term safety debates. The findings offer a critical counterpoint to speculative singularity timelines, providing AI safety researchers with empirical constraints to model future system capabilities and scaling laws. (source: https://x.com/GaryMarcus/status/2077282297540145192)

06

Straightening Latent Trajectories to Enhance World Model Planning

A new technical methodology has been introduced that optimizes generative world model planning by straightening latent trajectories. By locally adjusting these paths within the latent space, researchers have reduced the optimization complexity inherent in long-term sequence predictions, resulting in more coherent and stable simulated outcomes. This advancement directly addresses trajectory drift, a primary challenge in using world models for autonomous navigation and robotic decision-making. The technique simplifies latent pathing, making predictive world models more reliable for planning and testing physical actions in virtual environments. (source: https://x.com/ylecun/status/2077234041711972374)

07

Kling AI Announces Winners of the 2026 Campus AIGC Creation Competition

Kling AI has announced the winners of its NEXTGEN 2026 Campus AIGC Creation Competition, awarding the top Apex Award to the short film 'Kite'. Created by student filmmakers Shihua Lin, Mingwei Chen, and Xuanwei Liu, the project was honored for its narrative and visual exploration of freedom and identity. The film was rendered and generated using Kling AI's video synthesis platform, highlighting the growing role of generative video tools in academic and creative fields. This competition serves as a benchmark for the adoption of AIGC software in cinematic storytelling, demonstrating practical creative applications of the platform's video generation models. (source: https://x.com/Kling_ai/status/2077181348146479481)

08

French Startup Develops Advanced Humanoid Robot In Record Time

Rémi Cadene, a former Tesla AI and Hugging Face researcher, has developed a functional French-designed humanoid robot within a nine-month development cycle. By leveraging rapid engineering methodologies and advanced machine learning integration, Cadene's team successfully navigated from initial concept to physical hardware implementation. This rapid turnaround highlights the accelerating intersection of top-tier AI engineering talent with physical robotic design. The project signals a growing trend of high-performance robotic hardware development emerging outside of traditional Silicon Valley tech hubs, demonstrating how agile software pipelines can optimize physical robot manufacturing. (source: https://x.com/ylecun/status/2077235013553139895)

huggingface

6 stories
01

Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation

Researchers introduced SpectraReward, a training-free reward function that turns pretrained multimodal large language models (MLLMs) into zero-shot reward models for text-to-image reinforcement learning. Instead of utilizing preference labels or model fine-tuning, SpectraReward measures how well an original prompt can be recovered from a generated image using an image-conditioned, teacher-forced forward pass. The team also developed Self-SpectraReward, where a unified model's understanding branch provides rewards for its own generation branch. Across five out-of-distribution benchmarks and nine MLLM backbones ranging from 4B to 235B parameters, the method consistently improved generation quality, demonstrating that reward-policy alignment is crucial for image-generation RL. (source: https://huggingface.co/papers/2607.11886)

02

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Researchers introduced SearchGen-20K and SearchGen-Bench, a dataset and benchmark containing 20,839 prompts designed to evaluate the world-knowledge limitations of open visual generators. On this benchmark, frontier generators scored only 21 to 28 out of 100 due to their static training corpora failing to cover long-tailed and post-cutoff visual concepts. To address this, the authors propose a teach-then-search co-training framework that allows agents to discover and navigate their changing knowledge boundaries. This system leverages a 1-million-sample search corpus to dynamically retrieve context only when the generator's internal knowledge is insufficient, mitigating noise injection. (source: https://huggingface.co/papers/2607.05382)

03

Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution

A new QA-driven framework called ACQUIRE was introduced to improve how LLM-based coding agents resolve repository-level software issues. Rather than immediately generating code patches, the framework decouples repository understanding from the fix phase. It operates in two stages: first, a Questioner and Answerer agent pair collaboratively identify and bridge the model's repository knowledge gaps through autonomous exploration; second, a Resolver generates a patch informed by this structured QA knowledge. Evaluated on SWE-bench Verified, ACQUIRE raised the Pass@1 metric by up to 4.4 percentage points over existing pre-repair methods with modest time and cost overhead. (source: https://huggingface.co/papers/2607.11111)

04

Towards Autonomous and Auditable Medical Imaging Model Development

Researchers developed AMID, an autonomous multi-agent framework designed to automate medical imaging model development. Since medical imaging tasks require modality-specific configurations and strict validation, AMID uses Data-Conditioned Method Planning to refine complex tasks into parallel executable lanes grounded in data analysis. It then employs Verification-Guided Two-Stage Optimization to selectively exploit high-performing lanes while strictly verifying metrics and predictions. Evaluated across 20 medical imaging challenge tasks, AMID outperformed existing general-purpose machine learning engineering systems and approached the performance of solutions designed by human experts. (source: https://huggingface.co/papers/2607.10522)

05

What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness

An empirical study analyzed Eternis-Forecaster 8B, GLM-4.7-Flash, and GLM-4.5-Air to determine if internal representations offer better calibration and faithfulness than chain-of-thought (CoT) reasoning. By training representation-pooling probes on intermediate activations, researchers achieved superior calibration. They discovered that CoT traces often remain unchanged even when influential prompt evidence is removed, whereas internal probes successfully detect these shifts and act as lie detectors. Furthermore, forced answering tests showed that forecasting decisions are largely made before the reasoning tokens are generated, enabling token savings of 30% to 47% via routing. (source: https://huggingface.co/papers/2607.08046)

06

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

A new research paper provides a principled analysis of the evaluation and design paradigms in deep reinforcement learning (RL). The authors introduce the theoretical foundations of scaling laws in RL, demonstrating that asymptotic performance rankings and data-regimes do not maintain a monotone relationship. Through large-scale experiments, the study reveals that traditional design and evaluation paradigms have led to incorrect conclusions in RL literature. The findings offer a core analytical framework addressing scaling, model capacity, and algorithmic complexity to guide future deep reinforcement learning research. (source: https://huggingface.co/papers/2607.07769)