NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-10-02DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Gemini 3.0 Pro – early tests

Initial reports have emerged regarding the early testing phase of Gemini 3.0 Pro, Google's next-generation artificial intelligence model. These preliminary evaluations are crucial for assessing the model's enhanced capabilities across various benchmarks, including its performance in complex reasoning, advanced multimodal understanding, and sophisticated language generation tasks. The 'Pro' designation suggests a strategic focus on advanced applications, potentially targeting enterprise solutions, high-demand computational scenarios, and developers requiring state-of-the-art performance. Industry observers are keenly anticipating further details on its architectural improvements, scalability, and efficiency, especially in comparison to previous iterations of the Gemini family and competing models. The early test phase is vital for identifying potential areas for optimization, validating the model's stability and robustness, and ensuring its readiness for deployment across a diverse range of real-world applications before a wider public release. This period helps to refine its capabilities and solidify its competitive position in the rapidly evolving landscape of large language and multimodal models, setting expectations for its impact on AI development.

02

OpenAI's H1 2025: $4.3B in income, $13.5B in loss

OpenAI, a prominent artificial intelligence research and deployment company, reportedly achieved a significant revenue of $4.3 billion during the first half of 2025. This financial performance underscores the rapid commercialization success and expansion of its AI services, including its widely adopted large language models and other generative AI technologies. Despite this substantial income, the company is also reported to have incurred considerable losses, totaling $13.5 billion, within the same period. This substantial difference between revenue and loss highlights OpenAI's aggressive and ongoing investment strategy in cutting-edge AI research and development. It reflects massive expenditures on advanced computing infrastructure, talent acquisition, and foundational model training, crucial for sustaining its competitive edge in the fast-evolving AI landscape and pushing the boundaries of machine learning innovation.

03

Daniel Stenberg on 22 curl bugs found by AI and fixed

Daniel Stenberg, the lead developer of the widely-used curl project, has confirmed the successful identification and subsequent resolution of 22 bugs discovered through an artificial intelligence system. This significant development highlights the increasing utility of AI in enhancing software quality assurance and security. The AI system's ability to pinpoint these vulnerabilities underscores its potential as a powerful tool in automated bug detection, complementing traditional testing and manual review processes. The findings not only improve the robustness and reliability of curl, a critical component in countless applications and systems, but also offer valuable insights into the practical application of AI in identifying complex code flaws. This event serves as a notable case study demonstrating how AI technologies can contribute directly to the stability and security of foundational open-source software. The collaboration between AI-driven analysis and human expertise in patching these issues represents an important step forward in modern software development practices.

04

Two Amazon delivery drones crash into crane in commercial area of Tolleson, AZ

Two Amazon Prime MK30 delivery drones recently crashed into a crane located in a commercial area of Tolleson, Arizona. This incident, extensively reported by news outlets like The Verge and ABC15, brings to the forefront significant operational challenges and safety considerations inherent in advanced autonomous drone delivery systems. The collision underscores the intricate complexities involved in deploying Unmanned Aerial Vehicles (UAVs) within urban or semi-urban environments, particularly concerning the efficacy of their obstacle detection, avoidance, and overall navigation capabilities. Ongoing investigations are anticipated to thoroughly examine the drones' integrated navigation systems, sensor array performance, and underlying flight control algorithms to pinpoint the precise root cause of this dual collision. Such occurrences are instrumental in catalyzing the refinement of critical safety protocols, bolstering regulatory frameworks, and advancing the foundational artificial intelligence and robotics technologies that govern safe and reliable autonomous aerial operations, ultimately aiming to foster public trust and widespread adoption of future drone delivery services. The incident has reportedly prompted Amazon to temporarily suspend its drone delivery operations in the affected region.

05

NL Judge: Meta must respect user's choice of recommendation system

A Dutch court has ruled that Meta, the parent company of Facebook and Instagram, must provide users with the option to choose their preferred recommendation system. This landmark decision stems from a lawsuit initiated by Bits of Freedom, a digital rights organization. The ruling mandates Meta to move beyond purely personalized, data-driven algorithmic feeds, empowering users to opt for alternative or non-personalized content delivery mechanisms. This judgment has significant implications for how large social media platforms design and implement their recommendation algorithms, emphasizing user autonomy and control over their digital experience. It signals a growing legal push to regulate the pervasive influence of artificial intelligence in content curation, potentially paving the way for similar requirements in other jurisdictions. Compliance will necessitate considerable adjustments to Meta's underlying machine learning infrastructure, highlighting a crucial intersection between user rights, technological design, and platform governance.

06

Meta will listen into AI conversations to personalize ads

Meta is reportedly planning to leverage user interactions with its AI systems to personalize advertisements, a development that signals a significant shift in data collection and targeted advertising strategies. This approach would involve analyzing conversations users have with Meta's AI assistants to infer preferences, interests, and needs, subsequently tailoring ad content. While Meta's stated goal is to enhance user experience through more relevant ads, the initiative immediately raises substantial privacy concerns regarding the extent of data collection from private conversations and its potential implications for user trust and data security. Critics are likely to question the ethical boundaries of such a practice, especially concerning the processing of sensitive conversational data. This move underscores the ongoing tension between technological advancement in AI, commercial interests in advertising, and individual privacy rights, potentially prompting increased regulatory scrutiny on how AI-driven platforms handle user information. The strategy could also reshape the competitive landscape for personalized advertising, pushing other tech giants to explore similar data exploitation methods.

GitHub

3 stories
01

Tunix: A JAX-native LLM Post-Training Library

Tunix (Tune-in-JAX) is an early-stage, JAX-native library designed to streamline the post-training of Large Language Models, leveraging JAX for accelerated computation and seamless integration with Flax NNX. It offers comprehensive support across key methodologies including Supervised Fine-Tuning (SFT), Reinforcement Learning (RL), and Knowledge Distillation. For SFT, Tunix facilitates both full weights fine-tuning and Parameter-Efficient Fine-Tuning (PEFT) via LoRA/Q-LoRA layers. Its RL capabilities encompass Proximal Policy Optimization (PPO), Group Relative Policy Optimization (GRPO), and Token-level Group Sequence Policy Optimization (GSPO-token), alongside Direct Preference Optimization (DPO) for preference alignment. The library also provides diverse Knowledge Distillation strategies, such as logit matching, attention transfer, and feature pooling. Tunix emphasizes modularity, enabling easy customization, and efficiency, with native support for distributed training strategies like DP, FSDP, and TP on accelerators like TPUs. Currently under active development, Tunix aims to expand its advanced algorithms, scalability, and agentic RL training features, and is notably collaborating with the GRL framework for scalable LLM RL experiments on TPUs.

02

Pathway Live Data Framework

Pathway is a Python ETL framework designed for stream processing, real-time analytics, LLM pipelines, and Retrieval-Augmented Generation (RAG). It provides an intuitive Python API that integrates seamlessly with existing Python ML libraries, making it suitable for both development and production environments, handling both batch and streaming data with the same codebase. Underpinning Pathway is a high-performance, scalable Rust engine that leverages Differential Dataflow for incremental computation, enabling multithreading, multiprocessing, and distributed computations. Key features include a wide range of connectors (Kafka, GDrive, PostgreSQL, Airbyte), support for stateful transformations, persistence for computation state, and built-in consistency mechanisms. The framework also offers dedicated LLM tooling with wrappers, parsers, embedders, and an in-memory real-time Vector Index, facilitating the rapid development and deployment of live LLM and RAG applications. It can be easily deployed with Docker and Kubernetes, offering superior performance compared to traditional streaming technologies like Flink and Spark, and is available under a BSL 1.1 License.

03

Claude Agent SDK for Python

The Claude Agent SDK for Python provides a robust toolkit for interacting with Claude Code, enabling the development of sophisticated AI agent applications. It offers an asynchronous `query()` function for basic conversational interactions and a `ClaudeSDKClient` for more advanced, bidirectional conversations. Key features include the ability to define custom in-process tools as Python functions, which streamlines development, enhances performance by eliminating IPC overhead, and simplifies deployment compared to external MCP servers. The SDK also supports hooks, allowing developers to inject deterministic processing and automated feedback into the Claude agent loop at specific points, such as pre-tool use validation. Prerequisites include Python 3.10+, Node.js, and Claude Code. The SDK provides comprehensive error handling and type definitions for robust application building, making it a powerful solution for integrating Claude's agentic capabilities into Python projects.

huggingface

6 stories
01

DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search

Although RLVR has become an essential component for developing advanced reasoning skills in LLMs, contemporary studies have documented training plateaus that emerge following thousands of optimization steps, demonstrating notable decreases in performance gains despite increased computational investment. This limitation stems from the sparse exploration patterns inherent in current RLVR practices, where models rely on limited rollouts that often miss critical reasoning paths and fail to provide systematic coverage of the solution space. We present DeepSearch, a framework that integrates Monte Carlo Tree Search directly into RLVR training. In contrast to existing methods that rely on tree search only at inference, DeepSearch embeds structured search into the training loop, enabling systematic exploration and fine-grained credit assignment across reasoning steps. Through training-time exploration, DeepSearch addresses the fundamental bottleneck of insufficient exploration, which leads to diminishing performance improvements over prolonged training steps. Our contributions include: (1) a global frontier selection strategy that prioritizes promising nodes across the search tree, (2) selection with entropy-based guidance that identifies confident paths for supervision, and (3) adaptive replay buffer training with solution caching for efficiency. Experiments on mathematical reasoning benchmarks show that DeepSearch achieves 62.95% average accuracy and establishes a new state-of-the-art for 1.5B reasoning models - using 5.7x fewer GPU hours than extended training approaches. These results highlight the importance of strategic exploration over brute-force scaling and demonstrate the promise of algorithmic innovation for advancing RLVR methodologies. DeepSearch establishes a new direction for scaling reasoning capabilities through systematic search rather than prolonged computation.

02

GEM: A Gym for Agentic LLMs

The training paradigm for large language models (LLMs) is moving from static datasets to experience-based learning, where agents acquire skills via interacting with complex environments. To facilitate this transition we introduce GEM (General Experience Maker), an open-source environment simulator designed for the age of LLMs. Analogous to OpenAI-Gym for traditional reinforcement learning (RL), GEM provides a standardized framework for the environment-agent interface, including asynchronous vectorized execution for high throughput, and flexible wrappers for easy extensibility. GEM also features a diverse suite of environments, robust integrated tools, and single-file example scripts demonstrating using GEM with five popular RL training frameworks. Along with this, we also provide a set of baselines across 24 environments using REINFORCE with Return Batch Normalization (ReBN), which -- unlike GRPO -- is compatible with the full RL setting of dense per-turn rewards and offers better credit assignment. We further conduct apple-to-apple benchmarking of PPO, GRPO and REINFORCE in both single- and multi-turn settings using GEM to shed light on the algorithmic designs. Lastly, GEM also functions as a convenient evaluation toolkit besides a training environment. We hope this framework can help accelerate future agentic LLM research.

03

SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights

Post-training quantization has emerged as the most widely used strategy for deploying large language models at low precision. Still, current methods show perplexity degradation at bit-widths less than or equal to 4, partly because representing outliers causes precision issues in parameters that share the same scales as these outliers. This problem is especially pronounced for calibration-free, uniform quantization methods. We introduce SINQ to augment existing post-training quantizers with an additional second-axis scale factor and a fast Sinkhorn-Knopp-style algorithm that finds scales to normalize per-row and per-column variances, thereby minimizing a novel per-matrix proxy target for quantization: the matrix imbalance. Our method has no interactions between layers and can be trivially applied to new architectures to quantize any linear layers. We evaluate our method on the Qwen3 model family and DeepSeek-V2.5. SINQ improves WikiText2 and C4 perplexity significantly against uncalibrated uniform quantization baselines and can be further enhanced by combining it with calibration and non-uniform quantization levels. Code to reproduce the results of this work and to easily quantize models using SINQ is available at https://github.com/huawei-csl/SINQ.

04

Code2Video: A Code-centric Paradigm for Educational Video Generation

While recent generative models advance pixel-space video synthesis, they remain limited in producing professional educational videos, which demand disciplinary knowledge, precise visual structures, and coherent transitions, limiting their applicability in educational scenarios. Intuitively, such requirements are better addressed through the manipulation of a renderable environment, which can be explicitly controlled via logical commands (e.g., code). In this work, we propose Code2Video, a code-centric agent framework for generating educational videos via executable Python code. The framework comprises three collaborative agents: (i) Planner, which structures lecture content into temporally coherent flows and prepares corresponding visual assets; (ii) Coder, which converts structured instructions into executable Python codes while incorporating scope-guided auto-fix to enhance efficiency; and (iii) Critic, which leverages vision-language models (VLM) with visual anchor prompts to refine spatial layout and ensure clarity. To support systematic evaluation, we build MMMC, a benchmark of professionally produced, discipline-specific educational videos. We evaluate MMMC across diverse dimensions, including VLM-as-a-Judge aesthetic scores, code efficiency, and particularly, TeachQuiz, a novel end-to-end metric that quantifies how well a VLM, after unlearning, can recover knowledge by watching the generated videos. Our results demonstrate the potential of Code2Video as a scalable, interpretable, and controllable approach, achieving 40% improvement over direct code generation and producing videos comparable to human-crafted tutorials. The code and datasets are available at https://github.com/showlab/Code2Video.

05

ACON: Optimizing Context Compression for Long-horizon LLM Agents

Large language models (LLMs) are increasingly deployed as agents in dynamic, real-world environments, where success requires both reasoning and effective tool use. A central challenge for agentic tasks is the growing context length, as agents must accumulate long histories of actions and observations. This expansion raises costs and reduces efficiency in long-horizon tasks, yet prior work on context compression has mostly focused on single-step tasks or narrow applications. We introduce Agent Context Optimization (ACON), a unified framework that optimally compresses both environment observations and interaction histories into concise yet informative condensations. ACON leverages compression guideline optimization in natural language space: given paired trajectories where full context succeeds but compressed context fails, capable LLMs analyze the causes of failure, and the compression guideline is updated accordingly. Furthermore, we propose distilling the optimized LLM compressor into smaller models to reduce the overhead of the additional module. Experiments on AppWorld, OfficeBench, and Multi-objective QA show that ACON reduces memory usage by 26-54% (peak tokens) while largely preserving task performance, preserves over 95% of accuracy when distilled into smaller compressors, and enhances smaller LMs as long-horizon agents with up to 46% performance improvement.

06

BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in subject-consistent video generation due to an inherent difficulty in parsing prompts that specify complex spatial relationships, temporal logic, and interactions among multiple subjects. To address this issue, we propose BindWeave, a unified framework that handles a broad range of subject-to-video scenarios from single-subject cases to complex multi-subject scenes with heterogeneous entities. To bind complex prompt semantics to concrete visual subjects, we introduce an MLLM-DiT framework in which a pretrained multimodal large language model performs deep cross-modal reasoning to ground entities and disentangle roles, attributes, and interactions, yielding subject-aware hidden states that condition the diffusion transformer for high-fidelity subject-consistent video generation. Experiments on the OpenS2V benchmark demonstrate that our method achieves superior performance across subject consistency, naturalness, and text relevance in generated videos, outperforming existing open-source and commercial models.