NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-12-05DEFAULT EDITION
This issue
—
All time
—

Hacker News

4 stories
01

Show HN: SerpApi MCP Server

The Hacker News story introduces the 'SerpApi MCP Server,' a new project from SerpApi, a company recognized for its robust real-time search engine results APIs. The accompanying GitHub repository confirms 'MCP' stands for 'Multi-Cloud Proxy,' indicating this server is a critical infrastructure component. SerpApi's core operations involve complex web scraping and large-scale data extraction, demanding sophisticated proxy management, distributed architectures, and efficient request handling across diverse cloud environments. The MCP Server likely plays a pivotal role in optimizing these aspects, providing a robust solution for managing proxies across multiple cloud providers to ensure reliable and scalable data acquisition. This 'Show HN' initiative suggests an open-source contribution, offering valuable insights into the technical solutions employed by leading data extraction services. The project is particularly relevant for developers, engineers, and organizations focused on building scalable web scraping tools, distributed systems, or exploring the backend mechanics of large-scale data collection. It underscores SerpApi's commitment to technological advancement and community sharing in the domain of web data acquisition, indirectly supporting the data demands of AI and Machine Learning applications.

02

Jony Ive's OpenAI Device Barred from Using 'Io' Name

A collaborative project between renowned designer Jony Ive and artificial intelligence powerhouse OpenAI is reportedly facing a branding challenge, as their forthcoming device has been legally barred from utilizing the name 'Io'. This development suggests a potential conflict over intellectual property, trademark infringement, or a similar naming dispute, which could impact the product's market entry and branding strategy. Jony Ive, celebrated for his iconic industrial designs during his tenure at Apple, is venturing into new hardware territory with OpenAI, a leader in advanced AI research and development. This partnership is anticipated to produce an innovative AI-powered hardware device, aiming to blend cutting-edge artificial intelligence with sophisticated design and user experience. The prohibition from using the 'Io' designation underscores the intricate legal and commercial hurdles involved in launching high-profile technology products. This situation may necessitate a rebranding effort, potentially causing delays in the device's official unveiling or market availability, and highlights the broader complexities surrounding intellectual property in the rapidly evolving tech landscape.

03

Gemini 3 Pro: the frontier of vision AI

Google's latest release, Gemini 3 Pro, is positioned as a significant advancement at the forefront of vision AI, pushing the boundaries of what artificial intelligence can achieve in understanding and processing visual information. This iteration of the Gemini model family is expected to feature enhanced capabilities in areas such as object recognition, image analysis, video comprehension, and multimodal reasoning, integrating seamlessly with other data types. The announcement likely details architectural improvements and new techniques that enable more robust and nuanced interpretations of complex visual data, aiming to set new industry standards. Developers are anticipated to gain access to powerful tools for building next-generation applications in diverse fields, from augmented reality to automated inspection systems, leveraging Gemini 3 Pro's advanced visual intelligence to unlock novel functionalities and drive innovation across the AI landscape. The release underscores Google's ongoing commitment to leading AI research and development.

04

NeurIPS 2025 Best Paper Awards

The NeurIPS 2025 organizing committee has officially announced the recipients of its prestigious Best Paper Awards, recognizing groundbreaking research presented at the conference. These annual accolades celebrate exceptional contributions that significantly push the boundaries of machine learning and artificial intelligence, encompassing a broad spectrum of topics from fundamental theoretical advancements to innovative practical applications. The selected papers are meticulously chosen based on stringent criteria, including their originality, technical depth, empirical significance, clarity of presentation, and potential transformative impact on future research directions within the AI community. This year's honorees are anticipated to feature leading-edge work across diverse domains, such as novel deep learning architectures, advanced reinforcement learning algorithms, critical ethical AI considerations, and the development of more robust, efficient, and interpretable intelligent systems. This highly anticipated announcement underscores the vibrant and rapidly evolving research landscape within the global AI community, acknowledging the profound efforts of researchers who are consistently shaping the future of intelligent systems and driving technological progress in the field. These awards serve as a crucial benchmark for the most impactful and influential studies emerging from one of the premier conferences dedicated to neural information processing systems.

GitHub

1 story
01

Next AI Draw.io

Next AI Draw.io is a cutting-edge Next.js web application that seamlessly integrates advanced AI capabilities with draw.io diagrams, enabling users to create, modify, and enhance complex visualizations through natural language commands. Leveraging Large Language Models (LLMs), the platform supports direct diagram manipulation, image-based replication, and real-time refinement via an interactive chat interface. Key features include comprehensive diagram version history, specialized support for generating AWS, GCP, and Azure architecture diagrams with appropriate icons, and dynamic animated connectors for improved visualization. The application is built with Next.js, Vercel AI SDK, and react-drawio, representing diagrams as modifiable XML. It boasts multi-provider AI support, including AWS Bedrock, OpenAI, Anthropic, Google AI, and DeepSeek, offering flexibility for various LLM backends. Next AI Draw.io streamlines the diagramming process, making it accessible and efficient for technical and non-technical users alike, particularly valuable for documenting system architectures and general visual communication.

huggingface

6 stories
01

SIMA 2: A Generalist Embodied Agent for Virtual Worlds

We introduce SIMA 2, a generalist embodied agent that understands and acts in a wide variety of 3D virtual worlds. Built upon a Gemini foundation model, SIMA 2 represents a significant step toward active, goal-directed interaction within an embodied environment. Unlike prior work (e.g., SIMA 1) limited to simple language commands, SIMA 2 acts as an interactive partner, capable of reasoning about high-level goals, conversing with the user, and handling complex instructions given through language and images. Across a diverse portfolio of games, SIMA 2 substantially closes the gap with human performance and demonstrates robust generalization to previously unseen environments, all while retaining the base model's core reasoning capabilities. Furthermore, we demonstrate a capacity for open-ended self-improvement: by leveraging Gemini to generate tasks and provide rewards, SIMA 2 can autonomously learn new skills from scratch in a new environment. This work validates a path toward creating versatile and continuously learning agents for both virtual and, eventually, physical worlds.

02

Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction

The evolution of Large Language Models (LLMs) from passive responders to autonomous agents necessitates a fundamental shift in learning paradigms -- from static imitation to incentive-driven decision making. However, this transition is significantly impeded by the lack of scalable infrastructure capable of constructing high-quality interaction signals for effective policy learning. To address this, we introduce a comprehensive method designed to systematically scale the diversity and complexity of interactive environments. Our method realizes this scaling by addressing three orthogonal dimensions: (1) Complexity: NexAU, a flexible agent framework that supports building complex agent hierarchies via simple configurations; (2) Diversity: NexA4A automatically generates diverse agent hierarchies from natural language to cover infinite domains; and (3) Fidelity: NexGAP bridges the simulation-reality gap by integrating dynamic real-world environment for grounded trajectories synthesis. We train Nex-N1 upon the diverse and complex interactive environments established by our infrastructure. Empirical results on benchmarks such as SWE-bench and tau2 demonstrate that Nex-N1 consistently outperforms SOTA open-source models and achieves competitive performance against frontier proprietary models on complex agentic tasks. We open-source the Nex ecosystem and model weights to facilitate further research.

03

TV2TV: A Unified Framework for Interleaved Language and Video Generation

Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new class of omni video-text models that integrate ideas from recent LM reasoning advances to address this challenge. More specifically, we present TV2TV, a unified generative modeling framework which decomposes video generation into an interleaved text and video generation process. TV2TV jointly learns language modeling (next-token prediction) and video flow matching (next-frame prediction) using a Mixture-of-Transformers (MoT) architecture. At inference time, TV2TV decides when to alternate between generating text and video frames, allowing the model to "think in words" about subsequent content before "acting in pixels" to produce frames. This design offloads much of the responsibility for deciding what should happen next to the language modeling tower, enabling improved visual quality and prompt alignment of generated videos. It also enables fine-grained controllability, allowing users to modify the video generation trajectory through text interventions at any point in the process. In controlled experiments on video game data, TV2TV demonstrates substantial improvements in both visual quality and controllability. TV2TV also scales to natural videos, as we show by augmenting sports videos with interleaved natural language action descriptions using vision-language models (VLMs). Training TV2TV on this corpus yields strong visual quality and prompt alignment, showcasing the model's ability to reason about and generate complex real-world action sequences. Together, these results highlight TV2TV as a promising step toward video generation with open-ended textual reasoning and control.

04

ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning

Reward models are critical for aligning vision-language systems with human preferences, yet current approaches suffer from hallucination, weak visual grounding, and an inability to use tools for verification, limiting their reliability on complex multimodal reasoning tasks. We present ARM-Thinker, an Agentic multimodal Reward Model that autonomously invokes external tools (e.g., image cropping, doc page retrieval) to ground judgments in verifiable evidence, replacing static, non-interactive reward scoring. This enables the model to verify fine-grained visual details, cross-reference multi-page evidence, and validate reasoning claims, which are capabilities absent in existing reward models. We train ARM-Thinker with multi-stage reinforcement learning, jointly optimizing tool-calling decisions and judgment accuracy. To evaluate agentic reward modeling, we introduce ARMBench-VL, comprising three benchmarks that assess fine-grained visual grounding (image-level tools), multi-page document understanding (retrieval tools), and instruction following (text-level verification). ARM-Thinker achieves +16.2% average improvement on reward modeling benchmarks, +9.6% on tool-use tasks, and outperforms baselines on multimodal math and logical reasoning benchmarks. Our results demonstrate that agentic capabilities significantly enhance both accuracy and interpretability of reward models.

05

DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

Recent unified multimodal large language models (MLLMs) have shown impressive capabilities, incorporating chain-of-thought (CoT) reasoning for enhanced text-to-image generation. However, existing approaches remain limited, either treating the model merely as a standalone generator or relying on abstract textual planning. To this end, we propose Draft-as-CoT (DraCo), a novel interleaved reasoning paradigm that fully leverages both textual and visual contents in CoT for better planning and verification. Our method first generates a low-resolution draft image as preview, providing more concrete and structural visual planning and guidance. Then, we employ the model's inherent understanding capability to verify potential semantic misalignments between the draft and input prompt, and performs refinement through selective corrections with super-resolution. In this way, our approach addresses two fundamental challenges: the coarse-grained nature of textual planning and the difficulty in generating rare attribute combinations. To support training, we curate DraCo-240K, aiming to enhance three atomic capabilities spanning general correction, instance manipulation, and layout reorganization. Supported by DraCo-CFG, a specialized classifier-free guidance (CFG) strategy for interleaved reasoning, DraCo achieves a tremendous increase on GenEval (+8%), Imagine-Bench (+0.91), and GenEval++ (+3%), significantly outperforming direct generation and other generation methods empowered by CoT.

06

SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs

Extreme low-bit quantization is critical for efficiently deploying Large Language Models (LLMs), yet it often leads to severe performance degradation at 2-bits and even 4-bits (e.g., MXFP4). We present SignRoundV2, a post-training quantization framework that is highly effective even without mixed-precision. SignRoundV2 introduces (1) a fast sensitivity metric that combines gradient information with quantization-induced deviations to guide layer-wise bit allocation, and (2) a lightweight pre-tuning search for quantization scales to improve extremely low-bit quantization. These components allow SignRoundV2 to close the gap with full-precision models. Extensive experiments indicate that our method sustains competitive accuracy for LLMs, achieving production-grade performance with about 1 percent variance at 4-5 bits and strong results even at 2 bits. The implementation is available at https://github.com/intel/auto-round.