NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-07-16DEFAULT EDITION
This issue
—
All time
—

AI Blog

1 story
01

Cars24 Integrates OpenAI to Scale Voice and Chat Automation

Cars24 has integrated OpenAI-powered voice and chat agents to automate customer interactions and streamline internal operations. By implementing these AI-driven systems, the automotive e-commerce platform successfully manages over one million monthly conversation minutes. The deployment has enabled Cars24 to recover twelve percent of previously lost leads and distribute agentic workflows to various internal departments. These integrations allow the company to scale its customer support operations efficiently while building conversational features faster than with previous methodologies. (source: https://openai.com/index/cars24)

Hacker News

8 stories
01

Kimi K3 is now live

Kimi has officially launched its latest model, Kimi K3, representing a significant progression in conversational artificial intelligence. The new conversational large language model is designed to deliver superior comprehension capabilities, more coherent contextual awareness, and refined response generation compared to its predecessors. It focuses on handling complex queries and long-context processing with lower latency while maintaining high precision. Widely discussed on Hacker News, this release underscores the accelerating competition in the generative AI landscape as developers optimize performance and expand processing capabilities. (source: https://www.kimi.com/en)

02

German AI consortium releases Soofi S, an open 30B model that tops benchmarks

A German AI consortium has officially released Soofi S, an open-source 30-billion-parameter language model. Designed to provide localized linguistic optimization, the model successfully tops several established benchmarks in both English and German, positioning itself as a highly competitive alternative to proprietary models in the 30B parameter class. Developed through collaborative research, the model combines robust architectural engineering with optimized multi-lingual training corpora to achieve superior efficiency, reasoning, and context understanding without requiring massive proprietary licensing fees. (source: https://the-decoder.com/german-ai-consortium-releases-soofi-s-an-open-30b-model-that-tops-benchmarks-in-both-english-and-german/)

03

NotebookLM is now Gemini Notebook

Google has officially rebranded its AI-powered research and note-taking assistant, NotebookLM, as Gemini Notebook. This transition integrates the tool more deeply into Google's unified Gemini ecosystem, highlighting its reliance on the advanced capabilities of the Gemini model family. Gemini Notebook continues to provide users with a workspace for synthesizing documents, generating summaries, and creating audio overviews under a unified brand. The rebranding signals a strategic consolidation of Google's consumer and enterprise generative AI features, focusing on providing advanced contextual understanding, document analysis, and collaborative note-taking solutions. (source: https://blog.google/innovation-and-ai/products/gemini-notebook/notebooklm-gemini-notebook/)

04

Schema Harness Achieves ~99% on Arc‑AGI‑3 Public

The Schema Harness project has achieved an accuracy rate of approximately 99% on the Arc-AGI-3 Public benchmark dataset. The Abstraction and Reasoning Corpus (ARC), designed by François Chollet, measures artificial general intelligence capabilities by testing system adaptability and reasoning on grid-based visual tasks. Schema Harness uses a structured and domain-specific approach to improve performance on complex abstraction tasks. This milestone demonstrates a scalable framework that bridges deep learning heuristics and strict algorithmic reasoning, pointing toward more dependable and adaptive machine intelligence frameworks. (source: https://schema-harness.github.io/)

05

Launch HN: Traceforce (YC S26) – Company-wide security monitoring for AI apps

Traceforce, co-founded by Xia and Varun, has launched its enterprise security monitoring solution designed specifically for generative AI applications. The platform offers organizations deep visibility and control over generative AI tools like ChatGPT and Claude across corporate devices. Traceforce discovers active AI applications and monitors how they interface with internal and external databases through Model Context Protocols (MCP). To support developers, Traceforce released MCP-Xray, an open-source dynamic penetration testing utility designed to detect vulnerabilities in these Model Context Protocol connections. (source: https://news.ycombinator.com/item?id=48937020)

06

Agent-talk: Enabling coding agents to work together

Developer xhluca has introduced Agent-talk, an open-source framework designed to enable multiple autonomous coding agents to collaborate seamlessly on complex programming tasks. The system establishes structured communication protocols that allow individual agents to share context, negotiate task distribution, and collaboratively debug code. This multi-agent approach addresses the limitations of single-agent systems, which often struggle with large-scale codebases and complex logic. The repository provides the foundational tools, protocols, and environments necessary to orchestrate these specialized agents. (source: https://github.com/xhluca/agent-talk)

07

How to Train a Gen AI Kick Drum Model on Your Old Linux Desktop with 6GB VRAM

A developer has published a technical guide demonstrating how to train a generative AI diffusion model for kick drum synthesis on a local Linux desktop equipped with 6GB of VRAM. The tutorial outlines optimization strategies for running audio diffusion models under constrained memory conditions using PyTorch. By detailing data preparation, model architecture selection, and training pipeline configurations, the guide demonstrates that specialized audio generation models can be trained locally on consumer-grade hardware without expensive cloud GPUs. (source: https://www.zhinit.dev/blog/training-a-kick-drum-diffusion-model)

08

Show HN: QBasic Gorillas (Repeeled)

A developer has recreated the classic QBasic Gorillas game in vanilla JavaScript as a hands-on side project to test AI-assisted development methodologies and language models. Utilizing the Fable compiler and the Opus framework, the web implementation preserves the original gameplay mechanics while adding modern visual effects. The creator highlighted this project as a practical demonstration of how recreating vintage applications serves as an effective mechanism for exploring the capabilities of AI-driven coding assistants and generative AI workflows. (source: https://gameswithtony.com/gorillas/)

Twitter

8 stories
01

Gemini Omni Flash Tops Artificial Analysis Leaderboards for Multimodal Tasks

Google announced that its Gemini Omni Flash model has secured the number one position on the Artificial Analysis Leaderboards across both the Text to Video and Image to Video categories. This benchmark placement highlights the model's performance in multimodal generation and processing. Gemini Omni Flash demonstrated high efficiency in handling complex visual inputs and synthesizing video outputs from various prompts, marking a competitive milestone in Google's generative AI offerings. (source: https://x.com/Google/status/2077802568420339975)

02

Runway Launches Agent 2.0 To Advance Agentic Video Generation

Runway has unveiled Agent 2.0, a significant advancement in generative video technology designed to establish a new category of agentic video workflows. This release focuses on improving narrative coherence, cinematic language, and production quality in automated video generation. By shifting to an autonomous, agentic system model, Runway aims to streamline complex visual storytelling and close the gap between human creative intent and automated execution. (source: https://x.com/c_valenzuelab/status/2077793488272257135)

03

Gemini Omni Integration Enhances Video Editing In Google Vids

Google has integrated its Gemini Omni multimodal AI model into Google Vids to support natural language-based video editing. This update allows users to generate and edit professional-quality video content using simple textual prompts, minimizing the need for traditional manual video editing skills. The integration aims to reduce overall video production times and represents Google's broader effort to implement generative AI capabilities across its Workspace suite. (source: https://x.com/Google/status/2077789853828162031)

04

New Benchmarking Harness Achieves Exceptional Performance on ARC-A

Researchers announced a new evaluation harness designed to test model reasoning capabilities on the ARC-A benchmark, showing exceptional performance metrics. When utilized with the Opus 4.8 and Fable 5 models, the harness achieved a 99% Relative Human Average Effort (RHAE). Additionally, testing with the GPT-5.6 Sol model yielded a 95.35% success rate, offering a more precise evaluation methodology for advanced reasoning and problem-solving systems. (source: https://x.com/Michael_J_Black/status/2077775037302522146)

05

Rapid Evolution Of Mathematical Reasoning Capabilities In Large Language Models

This update outlines the historical progression of mathematical reasoning in Large Language Models, which scaled from elementary proficiency in 2023 to high school level in 2024, and reached International Mathematical Olympiad (IMO) gold medal standards by 2025. This progression is also reflected in the newly discussed GPT-5.6 Sol Pro, which targets open statistics questions and complex quantitative research issues. (source: https://x.com/polynoamial/status/2077762676932165996) (related discussion: https://x.com/gdb/status/2077622035984105848)

06

Google DeepMind Partners With Isomorphic Labs To Bolster Bioresilience

Google DeepMind has formed a strategic collaboration with Isomorphic Labs to develop proactive defenses against health outbreaks and biological risks using frontier AI models. The initiative applies computational biology and machine learning to construct predictive infrastructure and improve biosecurity metrics. This partnership targets scalable and precise defensive strategies to safeguard public health against emerging biological threats. (source: https://x.com/GoogleDeepMind/status/2077721122116640969)

07

Kling AI Showcases Advanced Text To Video Generation Capabilities

Kling AI has released a technical demonstration highlighting its advanced text-to-video generation model, showing progress in temporal consistency and realistic rendering. The release emphasizes the model's capacity to synthesize coherent cinematic motion dynamics from text prompts. Additionally, Kling AI recognized Kuan Cheng with a Silver Award in its 4K Short Film Creative Contest, demonstrating practical application of the video generation tools. (source: https://x.com/Kling_ai/status/2077770231129235615) (related context: https://x.com/Kling_ai/status/2077747579794952408)

08

Rate Limit Restrictions Lifted for All Claude Users

The Claude development team announced a complete reset of all weekly and five-hour rate limits across user accounts to optimize platform access. This infrastructure adjustment aims to deliver uninterrupted service and accommodate growing usage volume for complex, professional workflows. The limit removal indicates successful scaling of underlying infrastructure to support high-demand generative AI interactions without compromising overall system stability. (source: https://x.com/ClaudeDevs/status/2077603834453770467)

huggingface

8 stories
01

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning

Researchers introduced Ring-Zero, a training pipeline designed to scale zero-data reinforcement learning (zero RL) to a 1-trillion parameter model named Ring-2.5-1T-Zero. Utilizing algorithmic and system optimizations including clipped importance sampling, training-inference ratio correction, and mixed-precision control, the project explores training dynamics and emergent capabilities at scale. The 1T parameter model demonstrates improved sample efficiency and spontaneously develops advanced cognitive behaviors such as self-verification and parallel reasoning. Evaluated on seven mathematical benchmarks, the model achieves competitive reasoning performance while producing structured and concise reasoning traces. (source: https://huggingface.co/papers/2607.12395)

02

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

Researchers proposed KnowAct-GUIClaw, a novel Know-Route-Act-Reflect framework built to address cross-platform and self-improvement limitations in the OpenClaw agent system. The system employs a host agent to decompose long-horizon tasks, coupled with a pluggable GUI subagent featuring an experience-attributable memory system and a self-evolving skill library. Tested across Android, iOS, HarmonyOS, and Windows, GUIClaw paired with open-source Kimi-2.6 models achieved a 64.1% success rate on the long-horizon MobileWorld benchmark, outperforming several baseline frameworks and closed-source models such as Seed-2.0-Pro and GPT-5.5. (source: https://huggingface.co/papers/2607.12625)

03

Self-Improvements in Modern Agentic Systems: A Survey

A comprehensive survey on self-improving autonomous agents was published, framing modern agents as adaptive systems that convert experience into capability gains without human intervention. The authors present a system-level framework that defines an agent as a foundation model coupled with an operational scaffold of memory, tools, prompts, and control logic. Within this framework, self-improvement is conceptualized as a self-induced update operator targeting model parameters or scaffold configurations. The survey organizes existing literature by update target, change signals, applications, and evaluation metrics, while maintaining a repository of technical updates. (source: https://huggingface.co/papers/2607.13104)

04

Length Penalties Make Chain-of-Thought Less Monitorable

A study on reinforcement learning with length penalties revealed that compressing chain-of-thought (CoT) reasoning makes the underlying factors driving a model's answers harder to detect. Researchers trained Qwen3-4B and Qwen3-14B variants under different target lengths and evaluated them using biasing-hint interventions. While compression preserved multiple-choice accuracy and reduced token usage, it dramatically lowered model monitorability. At the strongest target compression, the rate at which a monitor caught hint utilization dropped from 69% to 49% for Qwen3-14B, demonstrating a clear compression-monitorability frontier where cheaper reasoning compromises transparency. (source: https://huggingface.co/papers/2607.09786)

05

PalmClaw: A Native On-Device Agent Framework for Mobile Phones

Developers released PalmClaw, an open-source agent framework designed to run natively on mobile phones and directly manage sessions, memory, skills, and tools on-device. Unlike conventional mobile agents that rely exclusively on sequential GUI actions like tapping and swiping, PalmClaw exposes raw device capabilities as structured tools with explicit arguments and defined execution boundaries. Experimental evaluations of the framework demonstrate an 11.5% relative improvement in task success and a 94.9% reduction in completion time compared to the strongest baseline mobile agent setups. (source: https://huggingface.co/papers/2607.13027)

06

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

Researchers introduced Vinci2, an on-device proactive assistant designed to analyze continuous egocentric videos and decide when to intervene without explicit user prompting. Alongside the system, the team released EgoServe, a benchmark for proactive assistance comprising over 3,000 service instances spanning four temporal memory horizons and ten categories. To address these tasks, they developed EgoMemo, a training-free, memory-augmented agent utilizing multi-scale temporal summaries, semantic knowledge graphs, and visual embeddings to perform retrieval-augmented reasoning. EgoMemo demonstrated strong baseline performance on the EgoServe benchmark. (source: https://huggingface.co/papers/2607.11523)

07

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

Researchers proposed ShortOPD, a short-to-long On-Policy Distillation (OPD) training schedule designed to recover the generation quality of structured pruned LLMs. Standard long-rollout OPD spends excessive optimization budget on low-information repetitive suffixes. ShortOPD mitigates this by identifying teacher-confirmed repetitive suffixes, treating the surviving prefix as the effective rollout length, and adjusting the budget dynamically. Across math, coding, and open-ended generation benchmarks, ShortOPD matched a fixed 8192-token rollout horizon within two points while requiring only 25% of the training time (8.5 hours versus 35.9 hours) and 71% fewer tokens. (source: https://huggingface.co/papers/2607.13124)

08

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

Developers introduced AgentCompass, an open-source, lightweight, and extensible infrastructure designed to evaluate LLM-based autonomous agents. To resolve fragmented evaluation pipelines, AgentCompass separates the evaluation process into three independent modules: Benchmark, Harness, and Environment. The system features a fault-tolerant asynchronous runtime and diagnostic tools to analyze execution trajectories for failure modes like reward-hacking. Natively supporting over 20 benchmarks across five capability dimensions, AgentCompass provides a scalable and reproducible environment for agent research. (source: https://huggingface.co/papers/2607.13705)