NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-09-22DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Qwen3-Omni: Native Omni AI model for text, image and video

Alibaba Cloud's Qwen team has unveiled Qwen3-Omni, representing a significant evolution in its large AI model series. This new model is engineered with a native 'Omni' capability, allowing it to seamlessly process and understand information across text, images, and video modalities from the ground up. This integrated approach distinguishes it from models that combine separate modules for different data types, aiming for a more unified and coherent multimodal understanding. Qwen3-Omni is designed to serve as a foundational framework for diverse applications, promising enhanced contextual awareness and sophisticated interactions across various media formats. This release underscores the ongoing trend in AI research towards developing general-purpose systems that can interpret and generate content across different data types efficiently and effectively, potentially setting new standards in multimodal AI performance.

02

OpenAI and Nvidia announce partnership to deploy 10GW of Nvidia systems

OpenAI and Nvidia have forged a strategic partnership to deploy a monumental 10-gigawatt capacity of Nvidia's advanced computing systems. This collaboration signals a significant acceleration in the development of artificial intelligence capabilities, particularly in the realm of large-scale model training and inference. The deployment of such an immense computational infrastructure underscores the escalating demand for processing power required to push the boundaries of current AI models, including the next generation of large language models and other sophisticated AI systems. Industry observers suggest that a 10GW deployment could translate into unprecedented supercomputing capabilities, enabling OpenAI to tackle highly complex AI research challenges, foster innovation in areas like generative AI, and potentially lead to breakthroughs that redefine the landscape of artificial intelligence. This partnership highlights the critical role of specialized hardware, like Nvidia's GPUs, in realizing the full potential of AI, marking a substantial investment in the future of AI development and demonstrating a shared commitment to advancing the field at an industrial scale.

03

California issues historic fine over lawyer's ChatGPT fabrications

California has reportedly imposed a "historic fine" on a lawyer for using OpenAI's ChatGPT to generate fabricated content, marking a significant development in the legal and ethical oversight of artificial intelligence tools in professional practice. This incident highlights the growing challenges and responsibilities associated with the adoption of large language models in fields requiring accuracy and factual integrity, such as law. The fine underscores a clear stance from regulatory bodies regarding the accountability of professionals when integrating AI, particularly when the technology leads to misrepresentation or the creation of false information. This case is likely to set a precedent for future legal and ethical guidelines surrounding AI-generated content, prompting a re-evaluation of current practices and emphasizing the critical need for human oversight and verification when leveraging advanced AI systems. It also brings into focus the broader debate on AI's reliability, the potential for misuse, and the necessary regulatory frameworks to ensure ethical deployment across various industries. Regulators and legal professionals are now facing the imperative to develop robust guidelines that address AI's role in critical applications, balancing innovation with the imperative for professional integrity and public trust. This landmark decision could influence AI policy and professional standards nationwide.

04

We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

This research emphasizes the critical need for Large Language Models (LLMs) to develop cultural communication competence, exemplified by the Persian art of Taarof. The paper argues that current LLMs, despite their linguistic capabilities, often fall short in navigating complex social protocols and indirect communication inherent in many cultures. By failing to grasp nuances like Taarof —a sophisticated system of politeness, deference, and indirectness —LLMs risk generating culturally insensitive or inappropriate responses. The study advocates for integrating socio-pragmatic understanding into LLM training, moving beyond purely semantic and syntactic processing. This approach aims to enhance AI systems' social intelligence, enabling them to engage in more contextually appropriate, respectful, and effective cross-cultural interactions, ultimately broadening their applicability and acceptance in diverse global contexts.

05

SWE-Bench Pro

SWE-Bench Pro, an enhanced and professional iteration of the widely recognized SWE-Bench benchmark, has been introduced by Scale AI to rigorously evaluate the software engineering capabilities of advanced AI agents and large language models (LLMs). This benchmark, available on GitHub, is specifically designed to test an AI's proficiency in addressing complex, real-world software development challenges, ranging from debugging and code refactoring to implementing new features across diverse open-source projects. Unlike its predecessor, the 'Pro' version likely incorporates a larger and more varied dataset, alongside potentially more sophisticated evaluation metrics, to provide a more comprehensive and demanding assessment. The initiative aims to push the frontiers of AI in software development, enabling researchers and developers to accurately gauge the practical applicability, autonomous reasoning, and robust problem-solving skills of AI models in intricate coding environments. SWE-Bench Pro is poised to become an essential tool for advancing the development of highly capable AI systems that can significantly contribute to the software engineering lifecycle.

06

AI-Generated "Workslop" Is Destroying Productivity

The proliferation of AI-generated content, characterized as "workslop," is increasingly being identified as a significant impediment to organizational productivity rather than a driver of efficiency. While artificial intelligence holds immense promise for automating tasks and streamlining workflows, the current output often lacks the necessary quality, precision, or contextual understanding to be directly usable. This frequently necessitates extensive human intervention for editing, refinement, and verification, thereby negating the intended time-saving benefits. The article highlights that this suboptimal AI output, if not properly managed and integrated, can lead to increased workload as employees spend valuable time correcting or completely redoing tasks initially handled by AI. The core issue lies in the gap between AI's generative capabilities and the specific, high-standard requirements of professional work. To counter this trend, organizations must focus on developing more sophisticated prompting techniques, implementing robust AI quality control mechanisms, and fostering better human-AI collaboration strategies. Ultimately, the unchecked deployment of AI without adequate oversight and refinement processes risks turning a potential productivity booster into a considerable drain on resources and efficiency.

Twitter

6 stories
01

gdb_Nvidia GPU Partnership

This announcement details a significant strategic partnership between the user gdb and NVIDIA, focusing on a substantial acquisition of GPU compute power. The deal involves acquiring "millions of GPUs," representing a compute capacity equivalent to NVIDIA's entire projected shipments for 2025. Furthermore, an investment of up to $100 billion is committed as these GPUs are deployed. This collaboration signals a major step forward in high-performance computing infrastructure, likely aimed at accelerating advanced AI research, model training, or large-scale inference workloads. The scale of the GPU acquisition and the financial commitment underscore the immense demand for cutting-edge hardware in the rapidly evolving AI landscape. The partnership is poised to drive significant advancements and capabilities in areas requiring massive computational resources.

02

GaryMarcus_AI Risks Letter

A significant letter, signed by 200 prominent leaders including Nobel laureates, AI experts, and former heads of state, has been delivered to the United Nations. The letter urgently calls for the establishment of international 'red lines' to prevent unacceptable risks posed by artificial intelligence. Gary Marcus, a signatory, expressed his belief that too much has been overlooked, and insufficient action has been taken to confront these potential dangers. The tweet also highlights a video contribution from Nobel laureate Maria Ressa, underscoring the gravity and broad consensus behind this call for AI governance and safety measures. This initiative reflects a growing concern among global thought leaders regarding the unchecked advancement of AI technologies and the imperative need for proactive regulatory frameworks to ensure responsible development and deployment.

03

natolambert_Reasoning Models Evolve

The evolution of reasoning models is shifting focus beyond just "thinking." With OpenAI's o1-preview release marking a year, the core primitives have expanded to include searching and executing tools. These additions compensate for the limitations of probabilistic models that rely on outdated parameter information. The synergy of thinking, searching, and acting is poised to form the bedrock of future systems. The engineering aspects of these combined functionalities are as crucial as achieving optimal model weights, highlighting a comprehensive approach to AI system development. This integrated approach signifies a move towards more robust and capable AI systems, ready to tackle complex tasks by leveraging diverse functionalities beyond core inference.

04

GoogleDeepMind_AI Safety Framework

Google DeepMind is emphasizing its commitment to responsible AI development as they build increasingly powerful artificial intelligence models. To this end, they are implementing their latest Frontier Safety Framework, which is described as their most comprehensive approach to date for identifying and proactively addressing emerging risks associated with advanced AI. This framework is designed to ensure that the development of powerful AI technologies is guided by robust safety protocols and risk mitigation strategies. The company aims to stay ahead of potential challenges and ensure that AI advancements are pursued with a strong focus on safety and ethical considerations. Further details about this initiative are available through the provided link, which offers a deeper dive into the methodologies and objectives of the Frontier Safety Framework.

05

ylecun_LLM JEPAs Release

The tweet from @ylecun highlights the release of the first iteration of JEPAs (Joint Embedding Predictive Architectures) specifically designed for Large Language Models (LLMs). This signifies a new development in the architecture and training methodologies for LLMs, potentially improving their predictive capabilities and understanding of sequential data. The announcement also emphasizes the availability of the code to the public, indicating a commitment to open-source development and collaboration within the AI research community. This release is likely to foster further research and experimentation with JEPAs in the context of LLMs, potentially leading to advancements in areas such as text generation, comprehension, and reasoning. The open-source nature of the project encourages wider adoption and contributions from other researchers and developers, accelerating progress in the field.

06

gdb_ChatGPT Usage Study Released

A large-scale study has been released detailing the current usage patterns of ChatGPT among consumers. The findings indicate a significant broadening of adoption beyond initial user groups, suggesting a mainstream integration of the technology. The research highlights the creation of substantial economic value stemming from both personal and professional applications of ChatGPT. This broad adoption points to the growing impact of AI tools in everyday life and professional environments, underscoring the maturation and widespread acceptance of large language models in various sectors. The study provides valuable insights into how individuals and businesses are leveraging AI for productivity and innovation, reflecting a dynamic shift in technological engagement.

GitHub

4 stories
01

Tongyi DeepResearch

Tongyi DeepResearch, developed by Tongyi Lab, is an innovative agentic large language model boasting 30.5 billion total parameters with a highly efficient 3.3 billion activated per token. Specifically engineered for long-horizon, deep information-seeking tasks, it has achieved state-of-the-art performance across a diverse set of agentic search benchmarks, including Humanity's Last Exam, BrowserComp, and WebWalkerQA. The model's robust capabilities stem from a sophisticated, fully automated synthetic data generation pipeline, enabling agentic pre-training, supervised fine-tuning, and reinforcement learning. It benefits from large-scale continual pre-training on high-quality agentic interaction data to enhance reasoning and maintain freshness. Furthermore, Tongyi DeepResearch employs an end-to-end reinforcement learning strategy, utilizing a customized Group Relative Policy Optimization framework with token-level policy gradients for stable training. At inference, it supports both the rigorous ReAct paradigm and an IterResearch-based 'Heavy' mode for optimized performance, making it a versatile solution for advanced web agent applications.

02

Elasticsearch

Elasticsearch is a powerful, distributed search and analytics engine, scalable data store, and vector database, meticulously optimized for speed and relevance in production environments. It serves as the cornerstone of Elastic's open Stack platform, enabling near real-time search over massive datasets, advanced vector searches, and seamless integration with generative AI applications. Key use cases span Retrieval Augmented Generation (RAG), traditional full-text search, and specialized applications such as logs, metrics, application performance monitoring (APM), and security logs. Users can easily get started via Elastic Cloud's managed service or through self-managed installations, including a quick local Docker setup using the `start-local` script. Elasticsearch supports interaction through robust REST APIs, language clients like Python, curl, and Kibana's Dev Tools Console for indexing and querying structured or unstructured data efficiently. It facilitates the storage and indexing of diverse data types, making them immediately available for search and exploration through Kibana.

03

像老乡鸡那样做饭

This GitHub repository, "CookLikeHOC," functions as a non-official, community-driven culinary resource dedicated to replicating the popular dishes served by the Chinese fast-food chain "Laoxiangji" (Home Original Chicken). Conceived and largely completed in 2024, the project diligently aggregates and structures recipes originally documented in the "《老乡鸡菜品溯源报告》" (Laoxiangji Dish Traceability Report). A significant recent update includes the launch of a dedicated web interface, available at cooklikehoc.soilzhu.su, enhancing accessibility for users. Furthermore, the platform has begun integrating AI-generated imagery for certain stew recipes, aiming to provide visual guidance, while actively soliciting authentic user-contributed photographs. The repository explicitly states its non-affiliation with the official Laoxiangji brand, emphasizing its role as a consumer-driven initiative. It serves as a comprehensive and interactive database for culinary enthusiasts seeking to explore and prepare these specific dishes.

04

tldraw

tldraw is an open-source React library designed for building highly interactive infinite canvas experiences, serving as the core technology behind the popular digital whiteboard tldraw.com. This monorepo provides a comprehensive SDK that allows developers to seamlessly integrate drawing, collaboration, and various visual interaction functionalities into their web applications. It offers straightforward installation and usage instructions, complemented by detailed guides for local development and contributions. A unique feature is the inclusion of CONTEXT.md files and an associated script, specifically designed to help AI agents rapidly gain context and understand the codebase, thereby enhancing automated development workflows. The tldraw SDK is available under a flexible license permitting both commercial and non-commercial projects, with options to remove the "Made with tldraw" watermark through a business license. The project fosters an active community through Discord and offers transparent guidelines regarding trademarks and contributions, positioning itself as a leading solution for sophisticated frontend graphical applications.

huggingface

6 stories
01

RPG: A Repository Planning Graph for Unified and Scalable Codebase

Large language models excel at function- and file-level code generation, yet generating complete repositories from scratch remains a fundamental challenge. This process demands coherent and reliable planning across proposal- and implementation-level stages, while natural language, due to its ambiguity and verbosity, is ill-suited for faithfully representing complex software structures. To address this, we introduce the Repository Planning Graph (RPG), a persistent representation that unifies proposal- and implementation-level planning by encoding capabilities, file structures, data flows, and functions in one graph. RPG replaces ambiguous natural language with an explicit blueprint, enabling long-horizon planning and scalable repository generation. Building on RPG, we develop ZeroRepo, a graph-driven framework for repository generation from scratch. It operates in three stages: proposal-level planning and implementation-level refinement to construct the graph, followed by graph-guided code generation with test validation. To evaluate this setting, we construct RepoCraft, a benchmark of six real-world projects with 1,052 tasks. On RepoCraft, ZeroRepo produces repositories averaging nearly 36K LOC, roughly 3.9times the strongest baseline (Claude Code) and about 64times other baselines. It attains 81.5% functional coverage and a 69.7% pass rate, exceeding Claude Code by 27.3 and 35.8 percentage points, respectively. Further analysis shows that RPG models complex dependencies, enables progressively more sophisticated planning through near-linear scaling, and enhances LLM understanding of repositories, thereby accelerating agent localization.

02

MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer

Unified multimodal Large Language Models (LLMs) that can both understand and generate visual content hold immense potential. However, existing open-source models often suffer from a performance trade-off between these capabilities. We present Manzano, a simple and scalable unified framework that substantially reduces this tension by coupling a hybrid image tokenizer with a well-curated training recipe. A single shared vision encoder feeds two lightweight adapters that produce continuous embeddings for image-to-text understanding and discrete tokens for text-to-image generation within a common semantic space. A unified autoregressive LLM predicts high-level semantics in the form of text and image tokens, with an auxiliary diffusion decoder subsequently translating the image tokens into pixels. The architecture, together with a unified training recipe over understanding and generation data, enables scalable joint learning of both capabilities. Manzano achieves state-of-the-art results among unified models, and is competitive with specialist models, particularly on text-rich evaluation. Our studies show minimal task conflicts and consistent gains from scaling model size, validating our design choice of a hybrid tokenizer.

03

Lynx: Towards High-Fidelity Personalized Video Generation

We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure identity fidelity. The ID-adapter employs a Perceiver Resampler to convert ArcFace-derived facial embeddings into compact identity tokens for conditioning, while the Ref-adapter integrates dense VAE features from a frozen reference pathway, injecting fine-grained details across all transformer layers through cross-attention. These modules collectively enable robust identity preservation while maintaining temporal coherence and visual realism. Through evaluation on a curated benchmark of 40 subjects and 20 unbiased prompts, which yielded 800 test cases, Lynx has demonstrated superior face resemblance, competitive prompt following, and strong video quality, thereby advancing the state of personalized video generation.

04

BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent

In the field of AI-driven human-GUI interaction automation, while rapid advances in multimodal large language models and reinforcement fine-tuning techniques have yielded remarkable progress, a fundamental challenge persists: their interaction logic significantly deviates from natural human-GUI communication patterns. To fill this gap, we propose "Blink-Think-Link" (BTL), a brain-inspired framework for human-GUI interaction that mimics the human cognitive process between users and graphical interfaces. The system decomposes interactions into three biologically plausible phases: (1) Blink - rapid detection and attention to relevant screen areas, analogous to saccadic eye movements; (2) Think - higher-level reasoning and decision-making, mirroring cognitive planning; and (3) Link - generation of executable commands for precise motor control, emulating human action selection mechanisms. Additionally, we introduce two key technical innovations for the BTL framework: (1) Blink Data Generation - an automated annotation pipeline specifically optimized for blink data, and (2) BTL Reward -- the first rule-based reward mechanism that enables reinforcement learning driven by both process and outcome. Building upon this framework, we develop a GUI agent model named BTL-UI, which demonstrates consistent state-of-the-art performance across both static GUI understanding and dynamic interaction tasks in comprehensive benchmarks. These results provide conclusive empirical validation of the framework's efficacy in developing advanced GUI Agents.

05

BaseReward: A Strong Baseline for Multimodal Reward Model

The rapid advancement of Multimodal Large Language Models (MLLMs) has made aligning them with human preferences a critical challenge. Reward Models (RMs) are a core technology for achieving this goal, but a systematic guide for building state-of-the-art Multimodal Reward Models (MRMs) is currently lacking in both academia and industry. Through exhaustive experimental analysis, this paper aims to provide a clear "recipe" for constructing high-performance MRMs. We systematically investigate every crucial component in the MRM development pipeline, including reward modeling paradigms (e.g., Naive-RM, Critic-based RM, and Generative RM), reward head architecture, training strategies, data curation (covering over ten multimodal and text-only preference datasets), backbone model and model scale, and ensemble methods. Based on these experimental insights, we introduce BaseReward, a powerful and efficient baseline for multimodal reward modeling. BaseReward adopts a simple yet effective architecture, built upon a {Qwen2.5-VL} backbone, featuring an optimized two-layer reward head, and is trained on a carefully curated mixture of high-quality multimodal and text-only preference data. Our results show that BaseReward establishes a new SOTA on major benchmarks such as MM-RLHF-Reward Bench, VL-Reward Bench, and Multimodal Reward Bench, outperforming previous models. Furthermore, to validate its practical utility beyond static benchmarks, we integrate BaseReward into a real-world reinforcement learning pipeline, successfully enhancing an MLLM's performance across various perception, reasoning, and conversational tasks. This work not only delivers a top-tier MRM but, more importantly, provides the community with a clear, empirically-backed guide for developing robust reward models for the next generation of MLLMs.

06

Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue

The ultimate goal of embodied agents is to create collaborators that can interact with humans, not mere executors that passively follow instructions. This requires agents to communicate, coordinate, and adapt their actions based on human feedback. Recently, advances in VLAs have offered a path toward this goal. However, most current VLA-based embodied agents operate in a one-way mode: they receive an instruction and execute it without feedback. This approach fails in real-world scenarios where instructions are often ambiguous. In this paper, we address this problem with the Ask-to-Clarify framework. Our framework first resolves ambiguous instructions by asking questions in a multi-turn dialogue. Then it generates low-level actions end-to-end. Specifically, the Ask-to-Clarify framework consists of two components, one VLM for collaboration and one diffusion for action. We also introduce a connection module that generates conditions for the diffusion based on the output of the VLM. This module adjusts the observation by instructions to create reliable conditions. We train our framework with a two-stage knowledge-insulation strategy. First, we fine-tune the collaboration component using ambiguity-solving dialogue data to handle ambiguity. Then, we integrate the action component while freezing the collaboration one. This preserves the interaction abilities while fine-tuning the diffusion to generate actions. During inference, a signal detector functions as a router that helps our framework switch between asking questions and taking actions. We evaluate the Ask-to-Clarify framework in 8 real-world tasks, where it outperforms existing state-of-the-art VLAs. The results suggest that our proposed framework, along with the training strategy, provides a path toward collaborative embodied agents.