NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-09-30DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Sora 2

OpenAI has introduced Sora 2, the latest iteration of its text-to-video generation model, building upon the capabilities of its predecessor. Sora 2 is designed to create highly realistic and imaginative video scenes purely from text descriptions, demonstrating significant advancements in understanding complex prompts, character consistency, and dynamic scene generation. The release is accompanied by a 'system card,' outlining the model's technical specifications, potential applications, and, crucially, safety considerations and ethical implications associated with advanced generative AI. This development underscores OpenAI's ongoing commitment to pushing the boundaries of multimodal AI, enabling users to translate abstract textual ideas into visually compelling moving images with unprecedented fidelity and control, while also addressing responsible deployment.

02

Claude Agent SDK for Python

Anthropic has officially launched the Claude Agent SDK for Python, a new software development kit aimed at empowering developers to build and deploy sophisticated AI agents utilizing Anthropic's Claude large language models. This SDK provides a comprehensive suite of tools and functionalities designed to simplify the complex process of integrating advanced conversational and reasoning capabilities into agent-based applications. By offering a Pythonic interface, the SDK enables developers to programmatically interact with Claude, facilitating the creation of intelligent agents that can understand natural language, engage in complex reasoning, make informed decisions, and execute actions within various environments. This release is expected to accelerate the development of practical AI agent applications across diverse sectors, allowing for more efficient automation, enhanced problem-solving, and novel interactive experiences. It represents a significant step towards making powerful AI agent technology more accessible and easier to implement for a broad range of development needs, fostering innovation in the AI ecosystem.

03

Comprehension debt: A ticking time bomb of LLM-generated code

The growing dependency on Large Language Models (LLMs) for generating code introduces a critical, yet often overlooked, challenge termed "comprehension debt." This concept, akin to technical debt, refers to the accumulated difficulty and time investment required for human developers to fully understand, maintain, and debug software components predominantly authored by AI. While LLMs offer significant advantages in accelerating initial development phases and boosting productivity, the article warns that a reduced cognitive load during generation can lead to a substantial increase in cognitive burden during later stages of the software lifecycle. This "ticking time bomb" threatens future software projects with potential slowdowns in maintenance, higher risks of introducing new bugs due to incomplete understanding, and a general decline in the long-term sustainability and quality of codebases. The implication is that without strategic interventions, organizations risk accumulating an unsustainable level of comprehension debt, which could severely impede innovation and operational efficiency as these AI-generated systems mature and require complex modifications or extensive troubleshooting. Addressing this requires a re-evaluation of developer workflows, potentially integrating new tools or methodologies to enhance human understanding and oversight of LLM-generated code from its inception.

04

Extract-0: A specialized language model for document information extraction

Extract-0 introduces a novel specialized language model designed specifically for document information extraction tasks. This model addresses the common challenges associated with accurately and efficiently extracting structured data from unstructured or semi-structured documents. By leveraging advanced natural language processing and deep learning techniques, Extract-0 aims to significantly improve the precision and recall of information retrieval, making it particularly useful for applications requiring automated data capture from various document types, such as invoices, contracts, and research papers. Its specialized architecture allows for fine-tuned performance on complex extraction scenarios, moving beyond general-purpose models to offer enhanced capabilities in targeted information extraction. The development signifies a crucial step forward in creating more adaptable and precise AI tools for business and research workflows that heavily rely on advanced document understanding and data digitization.

05

What is 'compute-in-memory' and why is it important for AI?

Compute-in-memory (CIM) represents a transformative architectural paradigm designed to address the inherent inefficiencies of traditional von Neumann architectures, where data must constantly shuttle between a central processing unit and discrete memory modules. This incessant data movement, known as the 'von Neumann bottleneck,' consumes significant energy and introduces latency, severely limiting the performance and power efficiency of modern computing systems. CIM tackles this challenge by integrating computational capabilities directly into or very close to memory units, allowing data processing to occur where the data resides. This novel approach is particularly vital for Artificial Intelligence applications, which are inherently data-intensive and computationally demanding. By minimizing the energy-intensive and time-consuming transfer of large datasets, CIM can dramatically accelerate neural network inference and training, enhancing overall AI system performance. It promises substantial improvements in energy efficiency and computational speed, making it an attractive solution for deploying complex AI models on resource-constrained edge devices, mobile platforms, and for enabling more powerful high-performance AI accelerators. As AI workloads continue to grow in complexity and scale, CIM offers a critical pathway toward achieving the necessary computational throughput and power efficiency for the next generation of intelligent systems, facilitating broader and more efficient AI adoption across various domains.

06

AI will happily design the wrong thing for you

The article, "AI will happily design the wrong thing for you," critically examines the inherent limitations of artificial intelligence when applied to design processes. It posits that while AI tools can efficiently generate numerous design solutions based on given parameters, they frequently lack the deeper contextual understanding, nuanced human empathy, and comprehensive grasp of underlying user intent required for truly effective and appropriate outcomes. The core argument is that AI, without significant human oversight and ethical considerations, may optimize for technical correctness or superficial metrics rather than actual suitability, potentially leading to designs that are functionally flawed, ethically problematic, or simply irrelevant to human needs. The author stresses the importance of human designers maintaining a pivotal role in defining objectives, evaluating AI outputs for relevance and impact, and integrating the critical ethical and social dimensions that AI currently struggles to comprehend. This perspective underscores the necessity of a collaborative human-AI approach, advocating for AI to serve as a powerful assistant rather than an autonomous decision-maker in complex creative and problem-solving tasks, particularly in fields like user experience and product design. This ensures that innovation remains aligned with human values and practical efficacy.

GitHub

2 stories
01

MoneyPrinterTurbo

MoneyPrinterTurbo is an open-source project for fully automated short video generation. Users provide a video theme or keywords, and the system automatically generates video scripts, selects visual materials, creates subtitles, and adds background music, culminating in a high-definition short video. Key features include a robust MVC architecture with both API and web interfaces, supporting AI-driven or custom script creation, and multiple HD video formats (9:16 and 16:9). It offers batch video generation, adjustable segment durations, and multi-language support (Chinese and English). The platform integrates various large language models such as OpenAI, Moonshot, Azure, DeepSeek, and more for advanced content generation, along with customizable voice synthesis and subtitle options. It utilizes high-quality, copyright-free video materials and allows for local asset integration, providing a comprehensive tool for streamlined video content creation.

02

openpilot

openpilot is an open-source operating system designed for robotics, specifically enhancing driver assistance systems in over 300 supported vehicles. Developed by comma.ai, it transforms compatible cars into semi-autonomous vehicles by providing advanced driving features. The system operates on dedicated hardware like the comma 3X device, requiring specific software installation and a car harness for integration. Beyond its core functionality, openpilot emphasizes robust safety and testing protocols, adhering to ISO26262 guidelines and incorporating extensive software-in-the-loop and hardware-in-the-loop testing, including a custom C-based safety model within its panda hardware. The project fosters an active development community, welcoming contributions via GitHub, and provides detailed documentation and development tools. It also offers various branches, from stable releases to bleeding-edge nightly builds, catering to different user needs. openpilot leverages user-consented driving data uploads to continuously train and improve its models, enhancing its performance across the fleet. While offering advanced capabilities, it is explicitly stated as alpha-quality software for research purposes, stressing user responsibility for compliance with local regulations and the absence of warranty.

huggingface

6 stories
01

From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones

We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned using pre-annotated, expert-generated human feedback. In this work, we introduce Reinforcement Learning from Human Interaction (RLHI), a paradigm that learns directly from in-the-wild user conversations. We develop two complementary methods: (1) RLHI with User-Guided Rewrites, which revises unsatisfactory model outputs based on users' natural-language follow-up responses, (2) RLHI with User-Based Rewards, which learns via a reward model conditioned on knowledge of the user's long-term interaction history (termed persona). Together, these methods link long-term user personas to turn-level preferences via persona-conditioned preference optimization. Trained on conversations derived from WildChat, both RLHI variants outperform strong baselines in personalization and instruction-following, and similar feedback enhances performance on reasoning benchmarks. These results suggest organic human interaction offers scalable, effective supervision for personalized alignment.

02

SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video alignment at a remarkably fast speed, deployable on RTX 5090 GPU. Two core designs ensure our efficient, effective and long video generation: (1) Linear DiT: We leverage linear attention as the core operation, which is more efficient than vanilla attention given the large number of tokens processed in video generation. (2) Constant-Memory KV cache for Block Linear Attention: we design block-wise autoregressive approach for long video generation by employing a constant-memory state, derived from the cumulative properties of linear attention. This KV cache provides the Linear DiT with global context at a fixed memory cost, eliminating the need for a traditional KV cache and enabling efficient, minute-long video generation. In addition, we explore effective data filters and model training strategies, narrowing the training cost to 12 days on 64 H100 GPUs, which is only 1% of the cost of MovieGen. Given its low cost, SANA-Video achieves competitive performance compared to modern state-of-the-art small diffusion models (e.g., Wan 2.1-1.3B and SkyReel-V2-1.3B) while being 16x faster in measured latency. Moreover, SANA-Video can be deployed on RTX 5090 GPUs with NVFP4 precision, accelerating the inference speed of generating a 5-second 720p video from 71s to 29s (2.4x speedup). In summary, SANA-Video enables low-cost, high-quality video generation.

03

Scaling Generalist Data-Analytic Agents

Data-analytic agents are emerging as a key catalyst for automated scientific discovery and for the vision of Innovating AI. Current approaches, however, rely heavily on prompt engineering over proprietary models, while open-source models struggle to face diverse-format, large-scale data files and long-horizon, multi-step reasoning that real-world analytics demands. This paper introduces DataMind, a scalable data synthesis and agent training recipe designed to build generalist data-analytic agents. DataMind tackles three key challenges in building open-source data-analytic agents, including insufficient data resources, improper training strategy, and unstable code-based multi-turn rollout. Concretely, DataMind applies 1) a fine-grained task taxonomy and a recursive easy-to-hard task composition mechanism to increase the diversity and difficulty of synthesized queries; 2) a knowledge-augmented trajectory sampling strategy followed by model-based and rule-based filtering; 3) a dynamically adjustable training objective combining both SFT and RL losses; 4) a memory-frugal and stable code-based multi-turn rollout framework. Built on DataMind, we curate DataMind-12K, a high-quality trajectory set spanning diverse domains, task categories, and data file formats for data-analytic tasks. Trained on DataMind-12K, our DataMind-14B achieves state-of-the-art with an average score of 71.16% on multiple data analysis benchmarks, outperforming the strongest proprietary baselines DeepSeek-V3.1 and GPT-5. Our DataMind-7B also performs best among all open-source models with a score of 68.10%. We also incorporate some empirical insights gained from our exploratory trials into the analysis experiments, aiming to provide actionable insights about agentic training for the community. We will release DataMind-12K and DataMind-7B,14B for the community's future research.

04

InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation

Long-sequence processing is a critical capability for modern large language models. However, the self-attention mechanism in the standard Transformer architecture faces severe computational and memory bottlenecks when processing long sequences. While trainable sparse attention methods offer a promising solution, existing approaches such as NSA introduce excessive extra parameters and disrupt the conventional pretrain-on-short, finetune-on-long workflow, resulting in slow convergence and difficulty in acceleration. To overcome these limitations, we introduce dense-sparse switchable attention framework, termed as InfLLM-V2. InfLLM-V2 is a trainable sparse attention that seamlessly adapts models from short to long sequences. Specifically, InfLLM-V2 reuses dense attention parameters through parameter-free architecture modification, maintaining consistency between short and long sequence processing. Additionally, InfLLM-V2 ensures computational efficiency across all sequence lengths, by using dense attention for short inputs and smoothly transitioning to sparse attention for long sequences. To achieve practical acceleration, we further introduce an efficient implementation of InfLLM-V2 that significantly reduces the computational overhead. Our experiments on long-context understanding and chain-of-thought reasoning demonstrate that InfLLM-V2 is 4times faster than dense attention while retaining 98.1% and 99.7% of the performance, respectively. Based on the InfLLM-V2 framework, we have trained and open-sourced MiniCPM4.1 (https://huggingface.co/openbmb/MiniCPM4.1-8B), a hybrid reasoning model, providing a reproducible implementation for the research community.

05

HunyuanImage 3.0 Technical Report

We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with its image generation module publicly available. The achievement of HunyuanImage 3.0 relies on several key components, including meticulous data curation, advanced architecture design, a native Chain-of-Thoughts schema, progressive model pre-training, aggressive model post-training, and an efficient infrastructure that enables large-scale training and inference. With these advancements, we successfully trained a Mixture-of-Experts (MoE) model comprising over 80 billion parameters in total, with 13 billion parameters activated per token during inference, making it the largest and most powerful open-source image generative model to date. We conducted extensive experiments and the results of automatic and human evaluation of text-image alignment and visual quality demonstrate that HunyuanImage 3.0 rivals previous state-of-the-art models. By releasing the code and weights of HunyuanImage 3.0, we aim to enable the community to explore new ideas with a state-of-the-art foundation model, fostering a dynamic and vibrant multimodal ecosystem. All open source assets are publicly available at https://github.com/Tencent-Hunyuan/HunyuanImage-3.0

06

The Era of Real-World Human Interaction: RL from User Conversations

We posit that to achieve continual model improvement and multifaceted alignment, future models must learn from natural human interaction. Current conversational models are aligned using pre-annotated, expert-generated human feedback. In this work, we introduce Reinforcement Learning from Human Interaction (RLHI), a paradigm that learns directly from in-the-wild user conversations. We develop two complementary methods: (1) RLHI with User-Guided Rewrites, which revises unsatisfactory model outputs based on users' natural-language follow-up responses, (2) RLHI with User-Based Rewards, which learns via a reward model conditioned on knowledge of the user's long-term interaction history (termed persona). Together, these methods link long-term user personas to turn-level preferences via persona-conditioned preference optimization. Trained on conversations derived from WildChat, both RLHI variants outperform strong baselines in personalization and instruction-following, and similar feedback enhances performance on reasoning benchmarks. These results suggest organic human interaction offers scalable, effective supervision for personalized alignment.