NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-05-04中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Let's Talk about LLMs

This Hacker News entry, titled 'Let's Talk about LLMs', sets the stage for an open discussion concerning Large Language Models, a pivotal and swiftly evolving domain within artificial intelligence. A comprehensive dialogue on this subject typically explores the multifaceted aspects of LLM technology, ranging from their foundational architectural innovations and the diverse practical applications spanning various industries to the critical ethical considerations inherent in their development and widespread deployment. Essential discussion points often encompass the formidable challenges associated with model scalability, the extensive data requirements for effective training, and the continuous research efforts dedicated to identifying and mitigating inherent biases while ensuring the responsible application of AI. Moreover, such conversations frequently extend to analyzing the profound societal impact of LLMs, their potential to fundamentally transform communication paradigms and automate complex tasks, and the significant economic implications for numerous job sectors. This particular post likely serves as an invitation for the tech community to engage in a collaborative exchange of insights regarding the current advancements, prospective future trajectories, and the broader societal and technical ramifications of these increasingly powerful artificial intelligence systems, fostering a deeper understanding of their influence and future potential.

02

White House Considers Vetting A.I. Models Before They Are Released

The White House is reportedly exploring a policy initiative to implement a mandatory vetting process for artificial intelligence models before their public release. This consideration underscores an increasing governmental focus on the potential societal impacts and inherent risks associated with advanced AI technologies. The proposed framework aims to establish rigorous pre-deployment checks to address critical concerns such as AI safety, algorithmic bias, data privacy, and the potential for misuse. Such a vetting system would likely involve a multi-stakeholder approach, drawing on expertise from government agencies, industry leaders, academic researchers, and ethics organizations to develop comprehensive evaluation criteria and regulatory guidelines. This move signals a significant step towards federal oversight in the rapidly evolving AI landscape, reflecting a commitment to fostering responsible innovation while mitigating unforeseen consequences as AI systems become more powerful and pervasive.

03

How OpenAI delivers low-latency voice AI at scale

OpenAI is detailing its comprehensive strategy for deploying low-latency voice AI solutions that can operate effectively at a global scale. The technical focus addresses the critical challenge of ensuring rapid response times, essential for natural human-computer interaction, while simultaneously handling high user demand. This involves sophisticated optimizations across the entire AI pipeline, from model design to inference and deployment. Key elements likely include the development of highly efficient neural network architectures specifically tailored for real-time audio processing, coupled with advanced techniques for model quantization and compression. Furthermore, the infrastructure relies on robust, distributed computing systems, employing specialized hardware accelerators and intelligent resource allocation to manage concurrent requests without performance degradation. This initiative underscores OpenAI's dedication to making advanced voice AI practical and widely available, enabling seamless integration into applications requiring instant vocal interaction, such as intelligent assistants, dictation services, and immersive entertainment. The described methods aim to overcome significant engineering hurdles in delivering high-quality, real-time AI experiences to a broad user base.

04

Usage-based pricing killing your vibe, here's how to roll your own local AI

This article addresses the growing financial concerns associated with usage-based pricing models common in cloud-hosted artificial intelligence services, which can lead to unpredictable and escalating costs. It advocates for a strategic shift towards deploying and managing AI solutions locally, offering users enhanced control over operational expenses and data privacy. The piece likely delves into various methods and tools for establishing a self-hosted AI environment, with a particular focus on implementing local AI coding agents. This approach not only helps mitigate the financial uncertainties tied to per-query or per-token charges but also significantly improves data security by retaining sensitive information within an on-premise infrastructure. By promoting the adoption of local AI setups, the article aims to guide technical users in building and customizing their own AI capabilities, prioritizing cost-effectiveness, operational independence, and robust data governance as alternatives to external service dependencies.

05

1966 Ford Mustang Converted into a Tesla with Working 'Full Self-Driving'

An innovative automotive project has successfully transformed a classic 1966 Ford Mustang into an electric vehicle, powered by Tesla components and remarkably equipped with a fully functional 'Full Self-Driving' (FSD) system. This unique conversion marries vintage American muscle car aesthetics with cutting-edge autonomous technology. The integration of Tesla's advanced FSD capabilities into a non-Tesla platform, particularly a decades-old vehicle, represents a significant engineering achievement. This project demonstrates the potential for retrofitting sophisticated AI-driven autonomous systems into a wider range of vehicles, challenging conventional notions of automotive modernization and autonomous driving deployment. It highlights the complex technical considerations involved in adapting modern software and hardware to legacy mechanical systems, pushing the boundaries of vehicle customization and intelligent transportation.

06

Talking to strangers at the gym

The article details a personal experiment involving the initiation of conversations with 35 strangers within a gym environment. This systematic approach aimed to explore various facets of unsolicited human interaction, communication strategies, and social dynamics in a controlled public space. The qualitative observations gathered from these engagements provide anecdotal yet valuable insights into human receptiveness, conversational flow, and the subtle cues that govern social acceptance or rejection. While not quantitative, this study implicitly contributes empirical data points relevant to behavioral modeling. These findings have potential applications in advancing the design of more sophisticated human-computer interaction systems, refining social robotics for nuanced contextual understanding, and informing AI agent development with enhanced social intelligence. The experiment underscores the importance of real-world social data for machine learning models focused on understanding and replicating complex human communicative patterns, particularly in areas like natural language processing for dynamic dialogue systems and the ethical considerations in social engineering for beneficial applications.

huggingface

6 stories
01

UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors

Recent progress has shown that video diffusion models (VDMs) can be repurposed for diverse multimodal graphics tasks. However, existing methods often train separate models for each problem setting, which fixes the input-output mapping and limits the modeling of correlations across modalities. We present UniVidX, a unified multimodal framework that leverages VDM priors for versatile video generation. UniVidX formulates pixel-aligned tasks as conditional generation in a shared multimodal space, adapts to modality-specific distributions while preserving the backbone's native priors, and promotes cross-modal consistency during synthesis. It is built on three key designs. Stochastic Condition Masking (SCM) randomly partitions modalities into clean conditions and noisy targets during training, enabling omni-directional conditional generation instead of fixed mappings. Decoupled Gated LoRA (DGL) introduces per-modality LoRAs that are activated when a modality serves as the generation target, preserving the strong priors of the VDM. Cross-Modal Self-Attention (CMSA) shares keys and values across modalities while keeping modality-specific queries, facilitating information exchange and inter-modal alignment. We instantiate UniVidX in two domains: UniVid-Intrinsic, for RGB videos and intrinsic maps including albedo, irradiance, and normal; and UniVid-Alpha, for blended RGB videos and their constituent RGBA layers. Experiments show that both models achieve performance competitive with state-of-the-art methods across distinct tasks and generalize robustly to in-the-wild scenarios, even when trained on fewer than 1,000 videos. Project page: https://houyuanchen111.github.io/UniVidX.github.io/

02

Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction

Agentic web search increasingly faces two distinct demands: deep reasoning over a single target, and structured aggregation across many entities and heterogeneous sources. Current systems struggle on both fronts. Breadth-oriented tasks demand schema-aligned outputs with wide coverage and cross-entity consistency, while depth-oriented tasks require coherent reasoning over long, branching search trajectories. We introduce Web2BigTable, a multi-agent framework for web-to-table search that supports both regimes. Web2BigTable adopts a bi-level architecture in which an upper-level orchestrator decomposes the task into sub-problems and lower-level worker agents solve them in parallel. Through a closed-loop run--verify--reflect process, the framework jointly improves decomposition and execution over time via persistent, human-readable external memory, with self-evolving updates to each single-agent. During execution, workers coordinate through a shared workspace that makes partial findings visible, allowing them to reduce redundant exploration, reconcile conflicting evidence, and adapt to emerging coverage gaps. Web2BigTable sets a new state of the art on WideSearch, reaching an Avg@4 Success Rate of 38.50 (7.5times the second best at 5.10), Row F1 of 63.53 (+25.03 over the second best), and Item F1 of 80.12 (+14.42 over the second best). It also generalises to depth-oriented search on XBench-DeepSearch, achieving 73.0 accuracy. Code is available at https://github.com/web2bigtable/web2bigtable.

03

Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies

Generalist robot policies increasingly benefit from large-scale pretraining, but offline data alone is insufficient for robust real-world deployment. Deployed robots encounter distribution shifts, long-tail failures, task variations, and human correction opportunities that fixed demonstration datasets cannot fully capture. We present Learning While Deploying (LWD), a fleet-scale offline-to-online reinforcement learning framework for continual post-training of generalist Vision-Language-Action (VLA) policies. Starting from a pretrained VLA policy, LWD closes the loop between deployment, shared physical experience, policy improvement, and redeployment by using autonomous rollouts and human interventions collected across a robot fleet. To stabilize learning from heterogeneous, sparse-reward fleet data, LWD combines Distributional Implicit Value Learning (DIVL) for robust value estimation with Q-learning via Adjoint Matching (QAM) for policy extraction in flow-based VLA action generators. We validate LWD on a fleet of 16 dual-arm robots across eight real-world manipulation tasks, including semantic grocery restocking and 3--5 minute long-horizon tasks. A single generalist policy improves as fleet experience accumulates, reaching an average success rate of 95%, with the largest gains on long-horizon tasks.

04

Let ViT Speak: Generative Language-Image Pre-training

In this paper, we present Generative Language-Image Pre-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers (ViTs) designed for multimodal large language models (MLLMs). To better align vision encoders with the autoregressive nature of LLMs, GenLIP trains a ViT to predict language tokens directly from visual tokens using a standard language modeling objective, without contrastive batch construction or an additional text decoder. This design offers three key advantages: (1) Simplicity: a single transformer jointly models visual and textual tokens; (2) Scalability: it scales effectively with both data and model size; and (3) Performance: it achieves competitive or superior results across diverse multimodal benchmarks. Trained on 8B samples from Recap-DataComp-1B, GenLIP matches or surpasses strong baselines despite using substantially less pretraining data. After continued pretraining on multi-resolution images at native aspect ratios, GenLIP further improves on detail-sensitive tasks such as OCR and chart understanding, making it a strong foundation for vision encoders in MLLMs.

05

Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance

Large Language Model (LLM) Red-Teaming, which proactively identifies vulnerabilities of LLMs, is an essential process for ensuring safety. Finding effective and diverse attacks in red-teaming is important, but achieving both is challenging. Generative Flow Networks (GFNs) that perform distribution matching are a promising methods, but they are notorious for training instability and mode collapse. In particular, unstable rewards in red-teaming accelerate mode collapse. We propose Stable-GFN (S-GFN), which eliminates partition function Z estimation in GFN and reduces training instability. S-GFN avoids Z-estimation through pairwise comparisons and employs a robust masking methodology against noisy rewards. Additionally, we propose a fluency stabilizer to prevent the model from getting stuck in local optima that produce gibberish. S-GFN provides more stable training while maintaining the optimal policy of GFN. We demonstrate the overwhelming attack performance and diversity of S-GFN across various settings.

06

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating high-level semantics and low-level details in a fully entangled manner. This is suboptimal for talking head synthesis: while audio and facial motion are semantically correlated, their low-level realizations (acoustic signals and visual textures) follow distinct rendering processes. Enforcing joint modeling across all levels causes unnecessary entanglement and reduces efficiency. We propose Talker-T2AV, an autoregressive diffusion framework where high-level cross-modal modeling occurs in a shared backbone, while low-level refinement uses modality-specific decoders. A shared autoregressive language model jointly reasons over audio and video in a unified patch-level token space. Two lightweight diffusion transformer heads decode the hidden states into frame-level audio and video latents. Experiments on talking portrait benchmarks show Talker-T2AV outperforms dual-branch baselines in lip-sync accuracy, video quality, and audio quality, achieving stronger cross-modal consistency than cascaded pipelines.