NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2025-12-11中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

GPT-5.2

OpenAI has officially announced the introduction of its newest large language model, GPT-5.2, marking a significant advancement in the field of artificial intelligence. This release is accompanied by comprehensive documentation accessible on the OpenAI platform and a detailed system card, which outlines the model's technical specifications, intended capabilities, and safety implementations. While specific performance metrics and novel features are expected to be detailed further, the launch of GPT-5.2 suggests substantial improvements over previous iterations, likely focusing on enhanced reasoning, broader knowledge integration, and potentially expanded multimodal functionalities. The system card is crucial for understanding the model's design principles, its limitations, and the robust safety measures adopted to ensure responsible deployment across various applications. This development reinforces OpenAI's commitment to pushing the boundaries of AI technology and making sophisticated models available to researchers and developers globally.

02

Disney making $1B investment in OpenAI, will allow characters on Sora AI

The Walt Disney Company has announced a significant $1 billion investment in OpenAI, signaling a strategic partnership aimed at integrating advanced artificial intelligence into its expansive entertainment ecosystem. This collaboration is specifically highlighted by Disney's intent to permit the use of its iconic characters on OpenAI's Sora AI platform. Sora, known for its text-to-video generative capabilities, is poised to revolutionize content creation within the media giant. This move suggests Disney is exploring innovative methods for storytelling, animation, and digital content production, leveraging AI to bring its beloved characters to life in novel ways. The substantial investment underscores the growing intersection between traditional entertainment and cutting-edge AI technologies, positioning Disney as a leader in adopting generative AI for creative endeavors. For OpenAI, this partnership provides significant capital and a high-profile application for its advanced models, potentially accelerating the development and refinement of its video generation capabilities. This strategic alliance is expected to pave the way for new forms of interactive experiences and media content, redefining how audiences engage with Disney's intellectual property through AI-powered platforms.

03

Last quarter I rolled out Microsoft Copilot to 4k employees

A recent post highlights a significant enterprise-wide deployment of Microsoft Copilot, an advanced AI assistant, to approximately 4,000 employees within the last fiscal quarter. This initiative underscores the increasing trend of integrating sophisticated artificial intelligence tools into large corporate environments to enhance productivity and streamline operational workflows. Such a large-scale rollout demonstrates an organizational commitment to leveraging generative AI technologies for practical business applications across diverse functions, including content creation, data analysis, and communication management. The successful deployment on this scale suggests a strategic investment in digital transformation, aiming to foster improved efficiency and adaptability within the workforce. While specific metrics on the impact are not detailed, the substantial number of users indicates a robust adoption strategy and a belief in the potential of AI to revolutionize daily tasks and support decision-making processes in a corporate setting.

04

Rivian Unveils Custom Silicon, R2 Lidar Roadmap, and Universal Hands Free

Rivian, the electric vehicle manufacturer, has announced significant advancements in its autonomous driving capabilities with the unveiling of custom-designed silicon, a detailed roadmap for its R2 Lidar technology, and a new 'Universal Hands Free' system. These developments are central to Rivian's next-generation autonomy platform, aiming to enhance the safety, performance, and scalability of its advanced driver-assistance systems (ADAS). The custom silicon is expected to provide superior processing power and efficiency for real-time sensor data interpretation and decision-making, crucial for complex driving scenarios. The R2 Lidar roadmap outlines the evolution of Rivian's perception systems, emphasizing improved range, resolution, and reliability, which are vital for robust environmental sensing. Furthermore, the Universal Hands Free feature suggests a more refined and broadly applicable semi-autonomous driving experience, potentially expanding the operational design domain of its hands-free driving functions. These strategic technological investments underscore Rivian's commitment to developing sophisticated in-house solutions to accelerate its progress in the competitive autonomous vehicle landscape, promising a more integrated and advanced user experience for its future lineup.

05

A Developer Accidentally Found CSAM in AI Data. Google Banned Him for It

A significant incident has recently highlighted severe challenges in AI data integrity and content moderation, as a developer reportedly discovered Child Sexual Abuse Material (CSAM) within an AI dataset. This accidental finding underscores the immense difficulties associated with thoroughly curating and sanitizing the vast quantities of data essential for training advanced artificial intelligence models. Following this unintended discovery, Google implemented a ban on the developer from its platforms, a decision that has ignited considerable debate. The situation raises critical questions regarding the ethical responsibilities of AI developers when encountering illegal content, the efficacy of current data filtering mechanisms employed by leading technology firms, and the adequacy of platform policies concerning user accounts in such sensitive contexts. There are concerns that such punitive actions could potentially deter researchers and developers from disclosing crucial vulnerabilities or illicit content within AI ecosystems, thereby impeding collective efforts to enhance data safety and promote ethical AI development. This case accentuates the urgent necessity for robust data governance frameworks, proactive content screening technologies, and transparent, supportive protocols for developers who encounter prohibited material, ensuring accountability while fostering a secure environment for reporting.

06

Show HN: GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus

The GPULlama3.java project unveils a significant advancement in deploying Llama models, specifically enabling their compilation to PTX/OpenCL for accelerated execution on GPUs. This capability is now seamlessly integrated into Quarkus, a popular Java framework renowned for its cloud-native and serverless application development. The core objective of this project is to drastically improve the performance and operational efficiency of running large language models within Java environments by effectively utilizing GPU hardware. Detailed instructions guide users through the entire setup process, encompassing the download and configuration of TornadoVM, the establishment of project-specific environment paths, building the project with Maven, and finally, executing the Llama model. This integration marks a crucial step forward for efficient AI model deployment and execution within Java-centric ecosystems, providing developers with a streamlined methodology to leverage the robust computational power of GPUs for demanding AI tasks. It particularly emphasizes TornadoVM's role in compiling and executing Java code across diverse hardware, including GPUs.

huggingface

6 stories
01

HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models

Vision-Language-Action (VLA) models have recently enabled robotic manipulation by grounding visual and linguistic cues into actions. However, most VLAs assume the Markov property, relying only on the current observation and thus suffering from temporal myopia that degrades long-horizon coherence. In this work, we view motion as a more compact and informative representation of temporal context and world dynamics, capturing inter-state changes while filtering static pixel-level noise. Building on this idea, we propose HiF-VLA (Hindsight, Insight, and Foresight for VLAs), a unified framework that leverages motion for bidirectional temporal reasoning. HiF-VLA encodes past dynamics through hindsight priors, anticipates future motion via foresight reasoning, and integrates both through a hindsight-modulated joint expert to enable a ''think-while-acting'' paradigm for long-horizon manipulation. As a result, HiF-VLA surpasses strong baselines on LIBERO-Long and CALVIN ABC-D benchmarks, while incurring negligible additional inference latency. Furthermore, HiF-VLA achieves substantial improvements in real-world long-horizon manipulation tasks, demonstrating its broad effectiveness in practical robotic settings.

02

UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving

Autonomous driving (AD) systems struggle in long-tail scenarios due to limited world knowledge and weak visual dynamic modeling. Existing vision-language-action (VLA)-based methods cannot leverage unlabeled videos for visual causal learning, while world model-based methods lack reasoning capabilities from large language models. In this paper, we construct multiple specialized datasets providing reasoning and planning annotations for complex scenarios. Then, a unified Understanding-Generation-Planning framework, named UniUGP, is proposed to synergize scene reasoning, future video generation, and trajectory planning through a hybrid expert architecture. By integrating pre-trained VLMs and video generation models, UniUGP leverages visual dynamics and semantic reasoning to enhance planning performance. Taking multi-frame observations and language instructions as input, it produces interpretable chain-of-thought reasoning, physically consistent trajectories, and coherent future videos. We introduce a four-stage training strategy that progressively builds these capabilities across multiple existing AD datasets, along with the proposed specialized datasets. Experiments demonstrate state-of-the-art performance in perception, reasoning, and decision-making, with superior generalization to challenging long-tail situations.

03

Composing Concepts from Images and Videos via Concept-prompt Binding

Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls short in accurately extracting complex concepts from visual inputs and flexibly combining concepts from both images and videos. We introduce Bind & Compose, a one-shot method that enables flexible visual concept composition by binding visual concepts with corresponding prompt tokens and composing the target prompt with bound tokens from various sources. It adopts a hierarchical binder structure for cross-attention conditioning in Diffusion Transformers to encode visual concepts into corresponding prompt tokens for accurate decomposition of complex visual concepts. To improve concept-token binding accuracy, we design a Diversify-and-Absorb Mechanism that uses an extra absorbent token to eliminate the impact of concept-irrelevant details when training with diversified prompts. To enhance the compatibility between image and video concepts, we present a Temporal Disentanglement Strategy that decouples the training process of video concepts into two stages with a dual-branch binder structure for temporal modeling. Evaluations demonstrate that our method achieves superior concept consistency, prompt fidelity, and motion quality over existing approaches, opening up new possibilities for visual creativity.

04

IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting

Recent advances in multimodal large language models (MLLMs) have led to impressive progress across various benchmarks. However, their capability in understanding infrared images remains unexplored. To address this gap, we introduce IF-Bench, the first high-quality benchmark designed for evaluating multimodal understanding of infrared images. IF-Bench consists of 499 images sourced from 23 infrared datasets and 680 carefully curated visual question-answer pairs, covering 10 essential dimensions of image understanding. Based on this benchmark, we systematically evaluate over 40 open-source and closed-source MLLMs, employing cyclic evaluation, bilingual assessment, and hybrid judgment strategies to enhance the reliability of the results. Our analysis reveals how model scale, architecture, and inference paradigms affect infrared image comprehension, providing valuable insights for this area. Furthermore, we propose a training-free generative visual prompting (GenViP) method, which leverages advanced image editing models to translate infrared images into semantically and spatially aligned RGB counterparts, thereby mitigating domain distribution shifts. Extensive experiments demonstrate that our method consistently yields significant performance improvements across a wide range of MLLMs. The benchmark and code are available at https://github.com/casiatao/IF-Bench.

05

Learning Unmasking Policies for Diffusion Language Models

Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference. One particularly successful variant is masked discrete diffusion, in which a buffer filled with special mask tokens is progressively replaced with tokens sampled from the model's vocabulary. Efficiency can be gained by unmasking several tokens in parallel, but doing too many at once risks degrading the generation quality. Thus, one critical design aspect of dLLMs is the sampling procedure that selects, at each step of the diffusion process, which tokens to replace. Indeed, recent work has found that heuristic strategies such as confidence thresholding lead to both higher quality and token throughput compared to random unmasking. However, such heuristics have downsides: they require manual tuning, and we observe that their performance degrades with larger buffer sizes. In this work, we instead propose to train sampling procedures using reinforcement learning. Specifically, we formalize masked diffusion sampling as a Markov decision process in which the dLLM serves as the environment, and propose a lightweight policy architecture based on a single-layer transformer that maps dLLM token confidences to unmasking decisions. Our experiments show that these trained policies match the performance of state-of-the-art heuristics when combined with semi-autoregressive generation, while outperforming them in the full diffusion setting. We also examine the transferability of these policies, finding that they can generalize to new underlying dLLMs and longer sequence lengths. However, we also observe that their performance degrades when applied to out-of-domain data, and that fine-grained tuning of the accuracy-efficiency trade-off can be challenging with our approach.

06

StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation

The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone. To address this challenge, we present StereoWorld, an end-to-end framework that repurposes a pretrained video generator for high-fidelity monocular-to-stereo video generation. Our framework jointly conditions the model on the monocular video input while explicitly supervising the generation with a geometry-aware regularization to ensure 3D structural fidelity. A spatio-temporal tiling scheme is further integrated to enable efficient, high-resolution synthesis. To enable large-scale training and evaluation, we curate a high-definition stereo video dataset containing over 11M frames aligned to natural human interpupillary distance (IPD). Extensive experiments demonstrate that StereoWorld substantially outperforms prior methods, generating stereo videos with superior visual fidelity and geometric consistency. The project webpage is available at https://ke-xing.github.io/StereoWorld/.