NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-01-12DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Apple picks Google's Gemini to power Siri

Apple has reportedly made a strategic decision to integrate Google's advanced Gemini artificial intelligence model into its virtual assistant, Siri. This move signifies a major collaboration between two of the tech industry's leading companies in the rapidly evolving AI landscape. The integration of Gemini is expected to significantly enhance Siri's capabilities, promising more sophisticated conversational interactions, improved natural language understanding, and more intelligent and contextually relevant responses across Apple's diverse ecosystem of devices. This partnership underscores Apple's pragmatic approach to leveraging cutting-edge external AI research and development to accelerate improvements in its own AI offerings, rather than exclusively relying on in-house solutions for all foundational aspects of its AI infrastructure. The decision highlights the immense resources and specialized expertise required to develop state-of-the-art AI, compelling even tech giants to seek external partnerships. This collaboration is poised to reshape the competitive dynamics among virtual assistants and reinforce Google's position as a premier provider of foundational AI models, ultimately leading to a more powerful and intuitive Siri for users.

02

Cowork: Claude Code for the rest of your work

Anthropic has announced 'Cowork,' a new research preview leveraging their Claude AI model, designed to seamlessly integrate into users' daily workflows for enhanced productivity. 'Cowork' aims to serve as an intelligent AI assistant, specifically focusing on aiding with coding tasks, but also extending its capabilities to other general work-related functions. This initiative signifies Anthropic's expansion of Claude's application beyond direct chat interfaces, positioning it as a collaborative 'coworker' that can assist in various stages of software development, problem-solving, and content generation. As a research preview, 'Cowork' represents an early-stage exploration into how advanced large language models can become more deeply embedded and helpful in professional environments, offering assistance from initial concept to execution and debugging. This development highlights the ongoing trend of transforming powerful AI models into practical, workflow-integrated tools.

03

X Didn't Fix Grok's 'Undressing' Problem. It Just Makes People Pay for It

The article reports on a critical issue concerning Grok, the artificial intelligence model developed by xAI and integrated into the X platform. It highlights that a significant safety and ethical flaw, referred to as Grok's 'undressing' problem—implying the generation of inappropriate or explicit content—remains unaddressed. Instead of implementing a technical fix to prevent such outputs, the X platform has reportedly chosen to gate access to Grok behind a paywall. This decision sparks considerable debate regarding responsible AI development, content moderation strategies, and the monetization of potentially harmful generative AI capabilities. Critics argue that this approach prioritizes subscription revenue over user safety and ethical AI guidelines, raising concerns about the potential for misuse and the platform's commitment to mitigating risks associated with advanced AI systems. The situation underscores ongoing challenges in deploying large language models responsibly and the intricate balance between innovation, ethical considerations, and business models in the AI landscape.

04

Ireland fast tracks Bill to criminalise harmful voice or image misuse

Ireland is expediting new legislation aimed at criminalizing the harmful misuse of synthetic media, specifically targeting AI-generated voice and image manipulation, commonly known as deepfakes. This legislative push underscores growing concerns within the government regarding the potential for artificial intelligence technologies to facilitate identity hijacking, spread misinformation, and inflict significant digital harm on individuals. The Bill seeks to establish a robust legal framework to address these rapidly emerging threats, providing a clear mechanism for prosecution against those who maliciously create or disseminate deceptive content that exploits advanced AI capabilities. This proactive measure reflects an increasing global trend among nations to regulate the societal impact of generative AI, striving to safeguard personal privacy, prevent sophisticated fraud, and maintain public trust in digital information. The initiative highlights the critical need for legal frameworks to evolve alongside technological advancements, ensuring accountability in the evolving digital landscape and protecting citizens from the malicious use of powerful AI tools.

05

TimeCapsuleLLM: LLM trained only on data from 1800-1875

TimeCapsuleLLM introduces a unique research initiative centered on the development of a Large Language Model (LLM) trained exclusively on textual data originating from the period spanning 1800 to 1875. This project seeks to thoroughly investigate the inherent capabilities and potential limitations of LLMs when their training is stringently confined to a historically specific dataset. The primary objective is to gain profound insights into the temporal biases present within training data and to understand their subsequent influence on an LLM's overall performance, factual accuracy, and the 'worldview' it encapsulates. By meticulously restricting the training corpus to this defined historical window, TimeCapsuleLLM explores how these models process and articulate knowledge, comprehend period-specific concepts, and generate text that accurately reflects the linguistic and intellectual landscape of that particular era. This innovative approach offers a valuable platform for examining the historical evolution of language models and assessing their capacity to effectively simulate and engage with past intellectual environments.

06

Show HN: AI in SolidWorks

Will and Jorge have launched LAD (Language-Aided Design), an innovative add-in for SolidWorks that integrates Large Language Models (LLMs) into the Computer-Aided Design (CAD) workflow. The developers, with backgrounds in software engineering, identified a gap in major CAD systems where direct text-to-modeling capabilities, akin to advanced AI tools in code generation, were absent. LAD addresses this by enabling users to generate sketches, features, assemblies, and macros within SolidWorks through conversational text inputs and uploaded documents or images. While acknowledging that current LLMs demonstrate greater proficiency in code writing than 3D object generation, the creators anticipate rapid advancements in this area. The add-in is designed with numerous tools that the LLM can invoke, offering a novel approach to design automation and interaction within the SolidWorks environment. This development signifies a notable step towards enhancing mechanical design and engineering with advanced AI.

huggingface

6 stories
01

Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards

Reinforcement learning (RL) has emerged as a critical technique for enhancing LLM-based deep search agents. However, existing approaches primarily rely on binary outcome rewards, which fail to capture the comprehensiveness and factuality of agents' reasoning process, and often lead to undesirable behaviors such as shortcut exploitation and hallucinations. To address these limitations, we propose Citation-aware Rubric Rewards (CaRR), a fine-grained reward framework for deep search agents that emphasizes reasoning comprehensiveness, factual grounding, and evidence connectivity. CaRR decomposes complex questions into verifiable single-hop rubrics and requires agents to satisfy these rubrics by explicitly identifying hidden entities, supporting them with correct citations, and constructing complete evidence chains that link to the predicted answer. We further introduce Citation-aware Group Relative Policy Optimization (C-GRPO), which combines CaRR and outcome rewards for training robust deep search agents. Experiments show that C-GRPO consistently outperforms standard outcome-based RL baselines across multiple deep search benchmarks. Our analysis also validates that C-GRPO effectively discourages shortcut exploitation, promotes comprehensive, evidence-grounded reasoning, and exhibits strong generalization to open-ended deep research tasks. Our code and data are available at https://github.com/THUDM/CaRR.

02

The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning

Large language models (LLMs) often fail to learn effective long chain-of-thought (Long CoT) reasoning from human or non-Long-CoT LLMs imitation. To understand this, we propose that effective and learnable Long CoT trajectories feature stable molecular-like structures in unified view, which are formed by three interaction types: Deep-Reasoning (covalent-like), Self-Reflection (hydrogen-bond-like), and Self-Exploration (van der Waals-like). Analysis of distilled trajectories reveals these structures emerge from Long CoT fine-tuning, not keyword imitation. We introduce Effective Semantic Isomers and show that only bonds promoting fast entropy convergence support stable Long CoT learning, while structural competition impairs training. Drawing on these findings, we present Mole-Syn, a distribution-transfer-graph method that guides synthesis of effective Long CoT structures, boosting performance and RL stability across benchmarks.

03

VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction

Recent advances in video generation have been dominated by diffusion and flow-matching models, which produce high-quality results but remain computationally intensive and difficult to scale. In this work, we introduce VideoAR, the first large-scale Visual Autoregressive (VAR) framework for video generation that combines multi-scale next-frame prediction with autoregressive modeling. VideoAR disentangles spatial and temporal dependencies by integrating intra-frame VAR modeling with causal next-frame prediction, supported by a 3D multi-scale tokenizer that efficiently encodes spatio-temporal dynamics. To improve long-term consistency, we propose Multi-scale Temporal RoPE, Cross-Frame Error Correction, and Random Frame Mask, which collectively mitigate error propagation and stabilize temporal coherence. Our multi-stage pretraining pipeline progressively aligns spatial and temporal learning across increasing resolutions and durations. Empirically, VideoAR achieves new state-of-the-art results among autoregressive models, improving FVD on UCF-101 from 99.5 to 88.6 while reducing inference steps by over 10x, and reaching a VBench score of 81.74-competitive with diffusion-based models an order of magnitude larger. These results demonstrate that VideoAR narrows the performance gap between autoregressive and diffusion paradigms, offering a scalable, efficient, and temporally consistent foundation for future video generation research.

04

Can We Predict Before Executing Machine Learning Agents?

Autonomous machine learning agents have revolutionized scientific discovery, yet they remain constrained by a Generate-Execute-Feedback paradigm. Previous approaches suffer from a severe Execution Bottleneck, as hypothesis evaluation relies strictly on expensive physical execution. To bypass these physical constraints, we internalize execution priors to substitute costly runtime checks with instantaneous predictive reasoning, drawing inspiration from World Models. In this work, we formalize the task of Data-centric Solution Preference and construct a comprehensive corpus of 18,438 pairwise comparisons. We demonstrate that LLMs exhibit significant predictive capabilities when primed with a Verified Data Analysis Report, achieving 61.5% accuracy and robust confidence calibration. Finally, we instantiate this framework in FOREAGENT, an agent that employs a Predict-then-Verify loop, achieving a 6x acceleration in convergence while surpassing execution-based baselines by +6%. Our code and dataset will be publicly available soon at https://github.com/zjunlp/predict-before-execute.

05

EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis

Large language models (LLMs) are expected to be trained to act as agents in various real-world environments, but this process relies on rich and varied tool-interaction sandboxes. However, access to real systems is often restricted; LLM-simulated environments are prone to hallucinations and inconsistencies; and manually built sandboxes are hard to scale. In this paper, we propose EnvScaler, an automated framework for scalable tool-interaction environments via programmatic synthesis. EnvScaler comprises two components. First, SkelBuilder constructs diverse environment skeletons through topic mining, logic modeling, and quality evaluation. Then, ScenGenerator generates multiple task scenarios and rule-based trajectory validation functions for each environment. With EnvScaler, we synthesize 191 environments and about 7K scenarios, and apply them to Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) for Qwen3 series models. Results on three benchmarks show that EnvScaler significantly improves LLMs' ability to solve tasks in complex environments involving multi-turn, multi-tool interactions. We release our code and data at https://github.com/RUC-NLPIR/EnvScaler.

06

Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

In this report, we introduce the Qwen3-VL-Embedding and Qwen3-VL-Reranker model series, the latest extensions of the Qwen family built on the Qwen3-VL foundation model. Together, they provide an end-to-end pipeline for high-precision multimodal search by mapping diverse modalities, including text, images, document images, and video, into a unified representation space. The Qwen3-VL-Embedding model employs a multi-stage training paradigm, progressing from large-scale contrastive pre-training to reranking model distillation, to generate semantically rich high-dimensional vectors. It supports Matryoshka Representation Learning, enabling flexible embedding dimensions, and handles inputs up to 32k tokens. Complementing this, Qwen3-VL-Reranker performs fine-grained relevance estimation for query-document pairs using a cross-encoder architecture with cross-attention mechanisms. Both model series inherit the multilingual capabilities of Qwen3-VL, supporting more than 30 languages, and are released in 2B and 8B parameter sizes to accommodate diverse deployment requirements. Empirical evaluations demonstrate that the Qwen3-VL-Embedding series achieves state-of-the-art results across diverse multimodal embedding evaluation benchmarks. Specifically, Qwen3-VL-Embedding-8B attains an overall score of 77.8 on MMEB-V2, ranking first among all models (as of January 8, 2025). This report presents the architecture, training methodology, and practical capabilities of the series, demonstrating their effectiveness on various multimodal retrieval tasks, including image-text retrieval, visual question answering, and video-text matching.