NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2025-12-15中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

It seems that OpenAI is scraping [certificate transparency] logs

Recent observations indicate that OpenAI, a prominent artificial intelligence research organization, may be actively scraping Certificate Transparency (CT) logs. CT logs serve as public, auditable records of all SSL/TLS certificates issued by Certificate Authorities, primarily designed to bolster internet security and detect fraudulent certificate issuances. OpenAI's engagement in this activity prompts inquiry into their extensive data acquisition strategies. While the precise objective remains conjectural, potential motivations include accumulating exhaustive data on newly registered websites, discerning domain ownership trends, or enriching datasets essential for training sophisticated AI models. This information could significantly contribute to comprehending the evolving web environment, refining web-crawling efficiencies, or augmenting the factual accuracy and real-time knowledge of large language models. This practice underscores the diverse and often innovative methodologies AI companies utilize to secure vast quantities of varied data, crucial for their technological progress and research endeavors.

02

I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me

This article, titled 'I'm Kenyan. I don't write like ChatGPT, ChatGPT writes like me,' delves into the distinctiveness of human writing, particularly from a specific cultural perspective, in contrast to the often generalized output of large language models (LLMs). The author posits that their unique, culturally-inflected writing style is an original expression, asserting that if an AI like ChatGPT produces similar text, it is due to the AI learning from diverse human-generated content, not the other way around. The piece challenges the idea that human communication is becoming homogenized by AI, instead highlighting how LLMs are trained to mimic existing human patterns. It underscores the critical importance of a wide array of training data, including contributions from varied cultural backgrounds, for AI to accurately reflect the richness and specificity of human language. Ultimately, the essay champions the authenticity and originality of human thought and expression amidst the growing influence of artificial intelligence.

03

We Put Flock Under Surveillance: Go Make Them Behave Differently [video]

The Hacker News entry 'We Put Flock Under Surveillance: Go Make Them Behave Differently' refers to a video presentation that likely explores a system or research initiative centered on the observation and modification of group behavior. Without explicit details from the video itself, the title suggests a technological application where a 'flock' - potentially referring to autonomous agents, animal groups, or even human collectives - is subjected to continuous monitoring. The primary objective appears to be the subsequent implementation of strategies designed to induce specific behavioral changes within the observed group. This project likely integrates sophisticated surveillance techniques, potentially leveraging computer vision, sensor networks, and data analytics to gather comprehensive information on collective actions. Furthermore, it would involve the development and application of control mechanisms or AI-driven interventions to influence and direct the 'flock's' behavior towards desired outcomes. The presentation likely discusses the technical challenges, methodologies, and potential implications of such advanced behavioral control systems.

04

If AI replaces workers, should it also pay taxes?

The discussion centers on the burgeoning debate regarding the economic and societal impact of Artificial Intelligence on the global workforce. As AI technologies become more sophisticated, their capacity to automate tasks traditionally performed by humans raises significant questions about job displacement and future economic models. A key proposal emerging from this discourse is the concept of implementing a "robot tax" or "AI tax." This tax, which would be levied on AI systems or companies utilizing them to replace human labor, aims to mitigate the adverse effects of automation, such as increased unemployment and income inequality. Proponents argue that such a tax could fund social safety nets, retraining programs for displaced workers, or even universal basic income initiatives, ensuring a more equitable distribution of the wealth generated by AI. Opponents, however, caution that taxing AI could stifle innovation, slow technological progress, and potentially disadvantage economies that adopt such policies. The debate underscores the urgent need for policymakers to address the intricate challenges posed by advanced AI and its transformative potential on the future of work and fiscal policy.

05

Falcon 9 rocket launches Starlink satellites before making 550th SpaceX landing

SpaceX achieved a notable milestone with the successful launch of a Falcon 9 rocket, deploying another batch of Starlink internet satellites into low Earth orbit. Following the orbital delivery, the Falcon 9 first-stage booster executed a precision autonomous landing, marking the 550th successful recovery for SpaceX across its operational history. This accomplishment further solidifies SpaceX's leadership in reusable rocket technology, a fundamental pillar for drastically reducing launch costs and increasing the frequency of space missions. The continued expansion of the Starlink constellation is critical for providing high-speed, low-latency broadband internet access globally, especially in remote and underserved regions. Such missions exemplify the operational reliability and advanced engineering capabilities of the Falcon 9 program, crucial for both commercial satellite deployment and the long-term vision of making humanity multi-planetary. The consistent success in reusability is a testament to the robust control systems and sophisticated automated operations integral to modern spaceflight.

06

There Are No Cows in Louis Pasteur's Crypt

The article, titled 'There Are No Cows in Louis Pasteur's Crypt,' delves into a historical examination of Louis Pasteur's contributions, potentially challenging common historical narratives or popular misconceptions surrounding his work. By critically analyzing historical evidence, the piece likely highlights the importance of fact-checking and the rigorous pursuit of scientific truth, themes highly relevant to contemporary research methodologies. In the context of AI and machine learning, this historical perspective can serve as a valuable reminder of the necessity for data validation, algorithm transparency, and the continuous evaluation of models to prevent the propagation of errors or biases. It underscores the foundational principle that scientific progress, whether in historical microbiology or advanced AI development, relies heavily on objective verification and the courage to re-examine established 'truths' in light of new insights or clearer historical understanding. The narrative implicitly encourages a disciplined approach to knowledge building across all scientific disciplines.

huggingface

6 stories
01

DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry

Reliable interpretation of multimodal data in dentistry is essential for automated oral healthcare, yet current multimodal large language models (MLLMs) struggle to capture fine-grained dental visual details and lack sufficient reasoning ability for precise diagnosis. To address these limitations, we present DentalGPT, a specialized dental MLLM developed through high-quality domain knowledge injection and reinforcement learning. Specifically, the largest annotated multimodal dataset for dentistry to date was constructed by aggregating over 120k dental images paired with detailed descriptions that highlight diagnostically relevant visual features, making it the multimodal dataset with the most extensive collection of dental images to date. Training on this dataset significantly enhances the MLLM's visual understanding of dental conditions, while the subsequent reinforcement learning stage further strengthens its capability for multimodal complex reasoning. Comprehensive evaluations on intraoral and panoramic benchmarks, along with dental subsets of medical VQA benchmarks, show that DentalGPT achieves superior performance in disease classification and dental VQA tasks, outperforming many state-of-the-art MLLMs despite having only 7B parameters. These results demonstrate that high-quality dental data combined with staged adaptation provides an effective pathway for building capable and domain-specialized dental MLLMs.

02

V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties

Large-scale video generation models have shown remarkable potential in modeling photorealistic appearance and lighting interactions in real-world scenes. However, a closed-loop framework that jointly understands intrinsic scene properties (e.g., albedo, normal, material, and irradiance), leverages them for video synthesis, and supports editable intrinsic representations remains unexplored. We present V-RGBX, the first end-to-end framework for intrinsic-aware video editing. V-RGBX unifies three key capabilities: (1) video inverse rendering into intrinsic channels, (2) photorealistic video synthesis from these intrinsic representations, and (3) keyframe-based video editing conditioned on intrinsic channels. At the core of V-RGBX is an interleaved conditioning mechanism that enables intuitive, physically grounded video editing through user-selected keyframes, supporting flexible manipulation of any intrinsic modality. Extensive qualitative and quantitative results show that V-RGBX produces temporally consistent, photorealistic videos while propagating keyframe edits across sequences in a physically plausible manner. We demonstrate its effectiveness in diverse applications, including object appearance editing and scene-level relighting, surpassing the performance of prior methods.

03

Sliding Window Attention Adaptation

The self-attention mechanism in Transformer-based Large Language Models (LLMs) scales quadratically with input length, making long-context inference expensive. Sliding window attention (SWA) reduces this cost to linear complexity, but naively enabling complete SWA at inference-time for models pretrained with full attention (FA) causes severe long-context performance degradation due to training-inference mismatch. This makes us wonder: Can FA-pretrained LLMs be well adapted to SWA without pretraining? We investigate this by proposing Sliding Window Attention Adaptation (SWAA), a set of practical recipes that combine five methods for better adaptation: (1) applying SWA only during prefilling; (2) preserving "sink" tokens; (3) interleaving FA/SWA layers; (4) chain-of-thought (CoT); and (5) fine-tuning. Our experiments show that SWA adaptation is feasible while non-trivial: no single method suffices, yet specific synergistic combinations effectively recover the original long-context performance. We further analyze the performance-efficiency trade-offs of different SWAA configurations and provide recommended recipes for diverse scenarios. Our code is available at https://github.com/yuyijiong/sliding-window-attention-adaptation

04

Exploring MLLM-Diffusion Information Transfer with MetaCanvas

Multimodal learning has rapidly advanced visual understanding, largely via multimodal large language models (MLLMs) that use powerful LLMs as cognitive cores. In visual generation, however, these powerful core models are typically reduced to global text encoders for diffusion models, leaving most of their reasoning and planning ability unused. This creates a gap: current multimodal LLMs can parse complex layouts, attributes, and knowledge-intensive scenes, yet struggle to generate images or videos with equally precise and structured control. We propose MetaCanvas, a lightweight framework that lets MLLMs reason and plan directly in spatial and spatiotemporal latent spaces and interface tightly with diffusion generators. We empirically implement MetaCanvas on three different diffusion backbones and evaluate it across six tasks, including text-to-image generation, text/image-to-video generation, image/video editing, and in-context video generation, each requiring precise layouts, robust attribute binding, and reasoning-intensive control. MetaCanvas consistently outperforms global-conditioning baselines, suggesting that treating MLLMs as latent-space planners is a promising direction for narrowing the gap between multimodal understanding and generation.

05

LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator

We propose LEO-RobotAgent, a general-purpose language-driven intelligent agent framework for robots. Under this framework, LLMs can operate different types of robots to complete unpredictable complex tasks across various scenarios. This framework features strong generalization, robustness, and efficiency. The application-level system built around it can fully enhance bidirectional human-robot intent understanding and lower the threshold for human-robot interaction. Regarding robot task planning, the vast majority of existing studies focus on the application of large models in single-task scenarios and for single robot types. These algorithms often have complex structures and lack generalizability. Thus, the proposed LEO-RobotAgent framework is designed with a streamlined structure as much as possible, enabling large models to independently think, plan, and act within this clear framework. We provide a modular and easily registrable toolset, allowing large models to flexibly call various tools to meet different requirements. Meanwhile, the framework incorporates a human-robot interaction mechanism, enabling the algorithm to collaborate with humans like a partner. Experiments have verified that this framework can be easily adapted to mainstream robot platforms including unmanned aerial vehicles (UAVs), robotic arms, and wheeled robot, and efficiently execute a variety of carefully designed tasks with different complexity levels. Our code is available at https://github.com/LegendLeoChen/LEO-RobotAgent.

06

Fairy2i: Training Complex LLMs from Real LLMs with All Parameters in {pm 1, pm i}

Large language models (LLMs) have revolutionized artificial intelligence, yet their massive memory and computational demands necessitate aggressive quantization, increasingly pushing representations toward the theoretical limit of a single bit. While complex-valued LLMs, such as iFairy, offer a superior chance for low-bit representation compared to real-valued counterparts, they require training from scratch, preventing the utilization of the vast ecosystem of pre-trained real-valued foundation models. Here we present Fairy2i, a universal framework that transforms pre-trained real-valued layers into an equivalent widely-linear complex form, enabling extremely low-bit quantization while reusing existing checkpoints. By proving a lossless mathematical equivalence between real and widely-linear maps, we convert standard Transformers into the complex domain and employ a phase-aware quantization scheme with a highly efficient codebook of fourth roots of unity. Furthermore, we introduce a recursive residual quantization mechanism that iteratively minimizes quantization error, allowing inference to proceed via efficient multiplication-free accumulation. We demonstrate that Fairy2i restores the performance of LLaMA-2 7B at an effective 2-bit precision to levels nearly comparable with full-precision baselines, significantly outperforming state-of-the-art real-valued binary and ternary quantization methods. This work bridges the gap between the representational efficiency of complex-valued arithmetic and the practical utility of pre-trained models, paving a new way for efficient inference on commodity hardware.