NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-12-01DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

DeepSeek-v3.2: Pushing the frontier of open large language models

DeepSeek-v3.2 marks a notable stride in the development of open large language models, aiming to expand the frontiers of accessible, high-performance AI. Released by DeepSeek AI, this version is designed to elevate the capabilities of open-source LLMs, likely through sophisticated architectural innovations, optimized training regimes, and a focus on enhanced reasoning and generation. The availability of DeepSeek-v3.2 on platforms like Hugging Face signifies a commitment to the open-source community, enabling researchers and developers to leverage, experiment with, and build upon advanced AI technology. This release positions itself as a strong contender in the landscape of LLMs, striving to offer competitive performance compared to closed-source alternatives across a spectrum of tasks, including complex problem-solving and nuanced conversational interactions. By fostering greater transparency and collaborative engagement, DeepSeek-v3.2 contributes to democratizing advanced AI, providing a powerful resource for innovation and diverse applications in the broader artificial intelligence domain.

02

Search tool that only returns content created before ChatGPT's public release

A novel search tool, dubbed "Slop Evader," has been introduced with the specific objective of filtering web content to display only results published prior to the public release of ChatGPT. This initiative directly addresses growing concerns among users and researchers regarding the proliferation of AI-generated content on the internet and its potential impact on information quality and authenticity. By restricting search results to the pre-ChatGPT era, the tool aims to offer a pristine digital environment, free from the influence of modern large language models. This allows users to access information, articles, and discussions that are unequivocally human-authored, providing a valuable resource for those seeking original content or conducting research uninfluenced by contemporary generative AI paradigms. The tool underscores a nascent demand for mechanisms to navigate and curate information in an increasingly AI-saturated online landscape.

03

Cara Hunter on the deepfake video that nearly ended her political career

Cara Hunter, a notable political figure, experienced a severe threat to her public career following the circulation of a highly explicit deepfake video. This digitally fabricated content, designed to falsely portray Hunter, exemplifies the increasing weaponization of artificial intelligence in political spheres. The incident illuminates the escalating challenges posed by synthetic media, particularly its capacity to disseminate misinformation and inflict substantial reputational harm on individuals. This specific case underscores the critical need for multifaceted approaches to counter deepfakes, encompassing advanced technological detection mechanisms, widespread public education initiatives, and the implementation of stringent legal and regulatory frameworks aimed at safeguarding individuals from such malicious digital assaults. The experience of Cara Hunter serves as a compelling and cautionary tale, illustrating how sophisticated AI technologies, when maliciously deployed, can yield devastating personal and professional ramifications, thereby intensifying ongoing debates concerning digital ethics, online safety, and the integrity of public discourse in the digital age.

04

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning

DeepSeek AI has unveiled DeepSeekMath-V2, an advanced large language model specifically engineered to significantly enhance mathematical reasoning capabilities through a novel self-verifiable approach. This new iteration aims to address the common challenges faced by existing large language models concerning accuracy and reliability when tackling complex mathematical problem-solving. By integrating sophisticated mechanisms for self-verification, DeepSeekMath-V2 is designed not only to generate potential solutions but also to critically evaluate and validate them, potentially through iterative refinement, formal proof generation, or consistency checks. This development represents a crucial step towards building more robust and trustworthy artificial intelligence systems for demanding scientific and engineering domains. The model's emphasis on verifiable reasoning is anticipated to substantially improve performance in fields requiring high precision and certainty, such as formal mathematics, theoretical physics, and intricate computational tasks, offering a new paradigm for AI-driven mathematical discovery and problem-solving.

05

Why I'm Betting Against the AGI Hype

This article explores the growing skepticism surrounding the immediate realization of Artificial General Intelligence (AGI), challenging the prevailing narrative of rapid progress often amplified by media and tech evangelists. The author argues that despite significant advancements in specific AI domains, particularly with large language models, these systems still lack fundamental aspects of human-like understanding, common sense, and true generalization capabilities. The piece delves into the inherent limitations of current AI architectures, suggesting that incremental improvements in narrow AI do not necessarily pave the way for AGI as quickly as some predict. It emphasizes the complex philosophical and technical hurdles that remain, such as achieving genuine consciousness, emotional intelligence, and real-world embodiment, which are often overlooked in the hype cycle. The author concludes by advocating for a more pragmatic and long-term view of AI development, urging a focus on robust and beneficial narrow AI applications rather than premature expectations for AGI.

06

Google, Nvidia, and OpenAI

This article, titled 'Google, Nvidia, and OpenAI', is expected to analyze the complex and competitive landscape involving these three pivotal entities shaping the future of artificial intelligence. It likely delves into Google's expansive AI ecosystem, encompassing cutting-edge research, powerful cloud computing services, and proprietary hardware like TPUs, positioning it as a key competitor to OpenAI. The piece would highlight Nvidia's indispensable role as the dominant provider of high-performance GPUs, which are foundational for training and deploying advanced AI models across the industry, including for both Google and OpenAI. Furthermore, the analysis would scrutinize OpenAI's rapid advancements in generative AI and large language models, examining its market impact, strategic alliances, and ongoing rivalry with established tech giants. The objective is to unravel the strategic implications of their collaborations, market competitions, and technological breakthroughs, offering insights into the evolving AI industry and its potential future directions.

GitHub

3 stories
01

TrendRadar

TrendRadar is an open-source GitHub project designed to aggregate, filter, and analyze real-time hot news from over 11 mainstream platforms, deployable in as little as 30 seconds. It offers intelligent push strategies (daily, current, incremental), precise content filtering using customizable keywords, and real-time trend analysis to understand topic evolution. The platform supports multi-channel notifications (WeChat, Feishu, Telegram, Email, Slack, etc.) and multi-terminal adaptation, including GitHub Pages and Docker deployment for data persistence. A key feature is its AI smart analysis, built on the Model Context Protocol (MCP), enabling natural language queries, in-depth trend analysis, sentiment analysis, and cross-platform data insights, making it an efficient tool for investors, content creators, and businesses to cut through information overload.

02

Agent Development Kit (ADK) for Go

The Agent Development Kit (ADK) for Go is an open-source, code-first toolkit designed to streamline the building, evaluating, and deploying of sophisticated AI agents. Developed by Google, ADK applies software development principles to AI agent creation, offering a flexible and modular framework for orchestrating agent workflows from simple tasks to complex multi-agent systems. While optimized for Google's Gemini, it maintains model and deployment agnosticism, ensuring compatibility across various AI models and frameworks. This Go version leverages the language's strengths in concurrency and performance, making it ideal for developers creating cloud-native agent applications, particularly in environments like Google Cloud Run. Key features include idiomatic Go design, a rich tool ecosystem for diverse agent capabilities, and a code-first approach that enhances flexibility, testability, and versioning. ADK-Go facilitates the creation of scalable applications through modular multi-agent system design, enabling easy containerization and deployment.

03

➤ Cursor Free VIP

Cursor Free VIP is a cross-platform utility tool designed for educational and research purposes, aiming to help users manage and reset configurations of the Cursor AI IDE. It supports Windows, macOS, and Linux operating systems across various architectures, offering automated installation scripts for ease of use. Key features include the ability to reset Cursor's settings and multi-language support, enhancing accessibility for a global user base. The project emphasizes its role as a learning aid, explicitly stating it does not generate fake email accounts or facilitate unauthorized OAuth access, and encourages users to support the original Cursor project. It provides detailed configuration options for various system paths and timing parameters, and includes a comprehensive changelog. Users are advised to run the script with administrator privileges and ensure the Cursor IDE is closed for optimal performance. This tool presents itself as a valuable resource for developers and researchers exploring the Cursor AI environment.

huggingface

6 stories
01

Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models

This work explores the challenge of building "Machines that Can Remember", framing long-term memory as the problem of efficient ultra-long context modeling. We argue that this requires three key properties: sparsity, random-access flexibility, and length generalization. To address ultra-long-context modeling, we leverage Hierarchical Sparse Attention (HSA), a novel attention mechanism that satisfies all three properties. We integrate HSA into Transformers to build HSA-UltraLong, which is an 8B-parameter MoE model trained on over 8 trillion tokens and is rigorously evaluated on different tasks with in-domain and out-of-domain context lengths to demonstrate its capability in handling ultra-long contexts. Results show that our model performs comparably to full-attention baselines on in-domain lengths while achieving over 90% accuracy on most in-context retrieval tasks with contexts up to 16M. This report outlines our experimental insights and open problems, contributing a foundation for future research in ultra-long context modeling.

02

AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement

Recently, multi-person video generation has started to gain prominence. While a few preliminary works have explored audio-driven multi-person talking video generation, they often face challenges due to the high costs of diverse multi-person data collection and the difficulty of driving multiple identities with coherent interactivity. To address these challenges, we propose AnyTalker, a multi-person generation framework that features an extensible multi-stream processing architecture. Specifically, we extend Diffusion Transformer's attention block with a novel identity-aware attention mechanism that iteratively processes identity-audio pairs, allowing arbitrary scaling of drivable identities. Besides, training multi-person generative models demands massive multi-person data. Our proposed training pipeline depends solely on single-person videos to learn multi-person speaking patterns and refines interactivity with only a few real multi-person clips. Furthermore, we contribute a targeted metric and dataset designed to evaluate the naturalness and interactivity of the generated multi-person videos. Extensive experiments demonstrate that AnyTalker achieves remarkable lip synchronization, visual quality, and natural interactivity, striking a favorable balance between data costs and identity scalability.

03

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

The landscape of high-performance image generation models is currently dominated by proprietary systems, such as Nano Banana Pro and Seedream 4.0. Leading open-source alternatives, including Qwen-Image, Hunyuan-Image-3.0 and FLUX.2, are characterized by massive parameter counts (20B to 80B), making them impractical for inference, and fine-tuning on consumer-grade hardware. To address this gap, we propose Z-Image, an efficient 6B-parameter foundation generative model built upon a Scalable Single-Stream Diffusion Transformer (S3-DiT) architecture that challenges the "scale-at-all-costs" paradigm. By systematically optimizing the entire model lifecycle -- from a curated data infrastructure to a streamlined training curriculum -- we complete the full training workflow in just 314K H800 GPU hours (approx. $630K). Our few-step distillation scheme with reward post-training further yields Z-Image-Turbo, offering both sub-second inference latency on an enterprise-grade H800 GPU and compatibility with consumer-grade hardware (<16GB VRAM). Additionally, our omni-pre-training paradigm also enables efficient training of Z-Image-Edit, an editing model with impressive instruction-following capabilities. Both qualitative and quantitative experiments demonstrate that our model achieves performance comparable to or surpassing that of leading competitors across various dimensions. Most notably, Z-Image exhibits exceptional capabilities in photorealistic image generation and bilingual text rendering, delivering results that rival top-tier commercial models, thereby demonstrating that state-of-the-art results are achievable with significantly reduced computational overhead. We publicly release our code, weights, and online demo to foster the development of accessible, budget-friendly, yet state-of-the-art generative models.

04

REASONEDIT: Towards Reasoning-Enhanced Image Editing Models

Recent advances in image editing models have shown remarkable progress. A common architectural design couples a multimodal large language model (MLLM) encoder with a diffusion decoder, as seen in systems such as Step1X-Edit and Qwen-Image-Edit, where the MLLM encodes both the reference image and the instruction but remains frozen during training. In this work, we demonstrate that unlocking the reasoning capabilities of MLLM can further push the boundaries of editing models. Specifically, we explore two reasoning mechanisms, thinking and reflection, which enhance instruction understanding and editing accuracy. Based on that, our proposed framework enables image editing in a thinking-editing-reflection loop: the thinking mechanism leverages the world knowledge of MLLM to interpret abstract instructions, while the reflection reviews editing results, automatically corrects unintended manipulations, and identifies the stopping round. Extensive experiments demonstrate that our reasoning approach achieves significant performance gains, with improvements of ImgEdit (+4.3%), GEdit (+4.7%), and Kris (+8.2%) when initializing our DiT from the Step1X-Edit (ReasonEdit-S), and also outperforms previous open-source methods on both GEdit and Kris when integrated with Qwen-Image-Edit (ReasonEdit-Q).

05

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning

Large language models have made significant progress in mathematical reasoning, which serves as an important testbed for AI and could impact scientific research if further advanced. By scaling reasoning with reinforcement learning that rewards correct final answers, LLMs have improved from poor performance to saturating quantitative reasoning competitions like AIME and HMMT in one year. However, this approach faces fundamental limitations. Pursuing higher final answer accuracy doesn't address a key issue: correct answers don't guarantee correct reasoning. Moreover, many mathematical tasks like theorem proving require rigorous step-by-step derivation rather than numerical answers, making final answer rewards inapplicable. To push the limits of deep reasoning, we believe it is necessary to verify the comprehensiveness and rigor of mathematical reasoning. Self-verification is particularly important for scaling test-time compute, especially for open problems without known solutions. Towards self-verifiable mathematical reasoning, we investigate how to train an accurate and faithful LLM-based verifier for theorem proving. We then train a proof generator using the verifier as the reward model, and incentivize the generator to identify and resolve as many issues as possible in their own proofs before finalizing them. To maintain the generation-verification gap as the generator becomes stronger, we propose to scale verification compute to automatically label new hard-to-verify proofs, creating training data to further improve the verifier. Our resulting model, DeepSeekMath-V2, demonstrates strong theorem-proving capabilities, achieving gold-level scores on IMO 2025 and CMO 2024 and a near-perfect 118/120 on Putnam 2024 with scaled test-time compute.

06

DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action

To build a generalizable Vision-Language-Action (VLA) model with strong reasoning ability, a common strategy is to first train a specialist VLA on robot demonstrations to acquire reliable manipulation skills, and then incorporate mixed annotated robot data together with multimodal data to restore broader reasoning capabilities. However, we observe that the resulting reasoning VLA often suffers from degraded action performance compared to the specialist model before fine-tuning, a phenomenon we refer to as action degeneration. To address this issue, we propose DualVLA, which enhances action performance through carefully designed post-training while still preserving reasoning capability. We first introduce a dual-layer data pruning method that removes redundant embodied reasoning, preventing it from adversely influencing action learning. To further strengthen action generation, we design a dual-teacher adaptive distillation strategy that assigns different supervision signals to different data domains while maintaining reasoning ability. To fill the evaluation gap for generalist VLAs, we also propose VLA Score, which decouples VLA capability into reasoning, intention, action, and alignment dimensions for a more fine-grained assessment. Experiments show that DualVLA achieves an average success rate of 61.0 in SimplerEnv and an average score of 65.4 across eight competitive multimodal benchmarks, demonstrating a stronger balance between precise action execution and multimodal understanding. Project Website: https://costaliya.github.io/DualVLA/.