NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-03-02DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Language Model Contains Personality Subnetworks

Recent research indicates that large language models (LLMs) may contain distinct 'personality subnetworks' within their complex architectures. This novel finding suggests that specific components of an LLM are primarily responsible for generating text that aligns with particular personality traits, moving beyond the traditional view of LLMs as monolithic systems for text generation. The study likely identifies these subnetworks through advanced interpretability techniques, examining how different layers or sets of neurons contribute to the model's expressive range in terms of personality. This discovery holds significant implications for the field of AI, particularly in understanding and controlling emergent behaviors in LLMs, mitigating biases related to persona generation, and potentially enabling more granular control over an AI's conversational style and output. The existence of such subnetworks could pave the way for developing more nuanced and ethically aligned AI assistants that can adapt their communication based on defined personality profiles, offering new avenues for personalized user interactions and advanced AI development.

02

New iPad Air, powered by M4

Apple has unveiled the latest generation of its iPad Air, now powered by the advanced M4 chip, marking a significant upgrade in performance and capabilities for the popular tablet line. The integration of the M4 chip brings substantial enhancements in CPU and GPU performance, critical for demanding applications, professional workflows, and high-fidelity gaming. Crucially, the M4 chip features an improved Neural Engine, designed to accelerate on-device artificial intelligence and machine learning tasks, making the new iPad Air a more powerful platform for generative AI applications and advanced computational photography. This move positions the iPad Air as a more formidable contender in the portable computing market, offering users a blend of power efficiency and robust performance.

03

If AI writes code, should the session be part of the commit?

The central question explored is the integration of AI-assisted code generation into standard software development practices, specifically whether the interactive sessions with AI tools should be recorded and included within version control commits. This inquiry underscores a critical discussion around the transparency and provenance of code in an increasingly AI-driven development environment. Incorporating AI interaction logs into commit histories could provide invaluable context for debugging, understanding design decisions, and tracking the complete lifecycle of code, thereby enhancing maintainability and accountability. Such a paradigm shift in workflow could redefine best practices for modern software engineering, addressing concerns related to code ownership, future debugging efforts, and the auditing of AI-generated components. The topic further prompts consideration of the necessary tooling and standards to effectively capture and manage these AI-assisted development artifacts, influencing developer workflows and the overall governance of projects leveraging artificial intelligence for code production.

04

Parallel coding agents with tmux and Markdown specs

A recent article introduces an innovative methodology for orchestrating artificial intelligence-powered coding agents, emphasizing parallel execution and structured task definition. This system effectively leverages tmux, a versatile terminal multiplexer, to manage multiple concurrent agent environments, thereby enabling agents to operate simultaneously on diverse programming tasks. Critical task specifications and detailed project requirements are meticulously outlined using Markdown, providing a clear, human-readable, and easily parseable format for agent instruction. This approach is designed to significantly enhance efficiency in software development by facilitating the parallel handling of various coding processes, including code generation, refactoring, debugging, and testing. By integrating widely adopted and robust tools such as tmux and Markdown, the framework presents a practical and accessible solution for developers aiming to integrate AI agents into complex software projects. The focus on parallel processing and standardized communication via Markdown not only fosters a more organized and scalable workflow but also promises to accelerate development cycles and potentially elevate code quality through distributed intelligent automation.

05

Civis Romanus Sum

The article, bearing the evocative Latin title "Civis Romanus Sum" (I am a Roman citizen), appears to delve into a contemporary philosophical and ethical discussion concerning the evolving role and potential status of artificial intelligence within human society. While the literal content provided is minimal, the title, when contextualized in a modern technological discourse, suggests an exploration of concepts such as digital citizenship, the rights and responsibilities of advanced AI entities, and their integration into established societal frameworks. It likely draws metaphorical parallels between the historical declaration of Roman citizenship, which conferred specific rights and protections, and the ongoing debate surrounding AI personhood, autonomy, and governance. The piece may critically analyze existing ethical guidelines for AI development, propose future legal or social structures for human-AI coexistence, and stimulate broader reflections on what it means to be a 'citizen' in an increasingly AI-driven world. This perspective positions the discussion as a crucial contribution to the field of AI ethics and societal impact.

06

Show HN: Omni – Open-source workplace search and chat, built on Postgres

Omni is a newly unveiled open-source workplace search and chat platform, presented as a self-hosted, more extensible, and cost-effective alternative to existing commercial solutions like Glean. It is engineered for small to mid-size teams and boasts seamless integration with popular enterprise applications such as Google Drive, Gmail, Slack, and Confluence. A core architectural innovation of Omni is its exclusive reliance on Postgres, specifically utilizing ParadeDB for advanced BM25 indexing and pgvector for HNSW vector indexing, thereby obviating the need for separate Elasticsearch instances or dedicated vector databases. This PostgreSQL-centric approach simplifies deployment and maintenance, requiring only a single `docker compose up` command for setup. The platform's robust search capabilities are powered by a hybrid search mechanism, intelligently combining results from both its BM25 and HNSW vector indices to deliver highly relevant and comprehensive information retrieval across all integrated data sources. Additionally, Omni is designed to support integration with Large Language Models (LLMs) for enhanced chat and search functionalities.

huggingface

6 stories
01

CUDA Agent: Large-Scale Agentic RL for High-Performance CUDA Kernel Generation

GPU kernel optimization is fundamental to modern deep learning but remains a highly specialized task requiring deep hardware expertise. Despite strong performance in general programming, large language models (LLMs) remain uncompetitive with compiler-based systems such as torch.compile for CUDA kernel generation. Existing CUDA code generation approaches either rely on training-free refinement or fine-tune models within fixed multi-turn execution-feedback loops, but both paradigms fail to fundamentally improve the model's intrinsic CUDA optimization ability, resulting in limited performance gains. We present CUDA Agent, a large-scale agentic reinforcement learning system that develops CUDA kernel expertise through three components: a scalable data synthesis pipeline, a skill-augmented CUDA development environment with automated verification and profiling to provide reliable reward signals, and reinforcement learning algorithmic techniques enabling stable training. CUDA Agent achieves state-of-the-art results on KernelBench, delivering 100%, 100%, and 92% faster rate over torch.compile on KernelBench Level-1, Level-2, and Level-3 splits, outperforming the strongest proprietary models such as Claude Opus 4.5 and Gemini 3 Pro by about 40% on the hardest Level-3 setting.

02

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding

This paper addresses the critical and underexplored challenge of long video understanding with low computational budgets. We propose LongVideo-R1, an active, reasoning-equipped multimodal large language model (MLLM) agent designed for efficient video context navigation, avoiding the redundancy of exhaustive search. At the core of LongVideo-R1 lies a reasoning module that leverages high-level visual cues to infer the most informative video clip for subsequent processing. During inference, the agent initiates traversal from top-level visual summaries and iteratively refines its focus, immediately halting the exploration process upon acquiring sufficient knowledge to answer the query. To facilitate training, we first extract hierarchical video captions from CGBench, a video corpus with grounding annotations, and guide GPT-5 to generate 33K high-quality chain-of-thought-with-tool trajectories. The LongVideo-R1 agent is fine-tuned upon the Qwen-3-8B model through a two-stage paradigm: supervised fine-tuning (SFT) followed by reinforcement learning (RL), where RL employs a specifically designed reward function to maximize selective and efficient clip navigation. Experiments on multiple long video benchmarks validate the effectiveness of name, which enjoys superior tradeoff between QA accuracy and efficiency. All curated data and source code are provided in the supplementary material and will be made publicly available. Code and data are available at: https://github.com/qiujihao19/LongVideo-R1

03

CL4SE: A Context Learning Benchmark For Software Engineering Tasks

Context engineering has emerged as a pivotal paradigm for unlocking the potential of Large Language Models (LLMs) in Software Engineering (SE) tasks, enabling performance gains at test time without model fine-tuning. Despite its success, existing research lacks a systematic taxonomy of SE-specific context types and a dedicated benchmark to quantify the heterogeneous effects of different contexts across core SE workflows. To address this gap, we propose CL4SE (Context Learning for Software Engineering), a comprehensive benchmark featuring a fine-grained taxonomy of four SE-oriented context types (interpretable examples, project-specific context, procedural decision-making context, and positive & negative context), each mapped to a representative task (code generation, code summarization, code review, and patch correctness assessment). We construct high-quality datasets comprising over 13,000 samples from more than 30 open-source projects and evaluate five mainstream LLMs across nine metrics. Extensive experiments demonstrate that context learning yields an average performance improvement of 24.7% across all tasks. Specifically, procedural context boosts code review performance by up to 33% (Qwen3-Max), mixed positive-negative context improves patch assessment by 30% (DeepSeek-V3), project-specific context increases code summarization BLEU by 14.78% (GPT-Oss-120B), and interpretable examples enhance code generation PASS@1 by 5.72% (DeepSeek-V3). CL4SE establishes the first standardized evaluation framework for SE context learning, provides actionable empirical insights into task-specific context design, and releases a large-scale dataset to facilitate reproducible research in this domain.

04

Mode Seeking meets Mean Seeking for Fast Long Video Generation

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. To address this, we propose a training paradigm where Mode Seeking meets Mean Seeking, decoupling local fidelity from long-term coherence based on a unified representation via a Decoupled Diffusion Transformer. Our approach utilizes a global Flow Matching head trained via supervised learning on long videos to capture narrative structure, while simultaneously employing a local Distribution Matching head that aligns sliding windows to a frozen short-video teacher via a mode-seeking reverse-KL divergence. This strategy enables the synthesis of minute-scale videos that learns long-range coherence and motions from limited long videos via supervised flow matching, while inheriting local realism by aligning every sliding-window segment of the student to a frozen short-video teacher, resulting in a few-step fast long video generator. Evaluations show that our method effectively closes the fidelity-horizon gap by jointly improving local sharpness, motion and long-range consistency. Project website: https://primecai.github.io/mmm/.

05

DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference

Vision-language models (VLMs) have achieved remarkable multimodal understanding and reasoning capabilities, yet remain computationally expensive due to dense visual tokenization. Existing efficiency approaches either merge redundant visual tokens or drop them progressively in language backbone, often trading accuracy for speed. In this work, we propose DUET-VLM, a versatile plug-and-play dual compression framework that consists of (a) vision-only redundancy aware compression of vision encoder's output into information-preserving tokens, followed by (b) layer-wise, salient text-guided dropping of visual tokens within the language backbone to progressively prune less informative tokens. This coordinated token management enables aggressive compression while retaining critical semantics. On LLaVA-1.5-7B, our approach maintains over 99% of baseline accuracy with 67% fewer tokens, and still retains >97% even at 89% reduction. With this dual-stage compression during training, it achieves 99.7% accuracy at 67% and 97.6% at 89%, surpassing prior SoTA visual token reduction methods across multiple benchmarks. When integrated into Video-LLaVA-7B, it even surpasses the baseline -- achieving >100% accuracy with a substantial 53.1% token reduction and retaining 97.6% accuracy under an extreme 93.4% setting. These results highlight end-to-end training with DUET-VLM, enabling robust adaptation to reduced visual (image/video) input without sacrificing accuracy, producing compact yet semantically rich representations within the same computational budget. Our code is available at https://github.com/AMD-AGI/DUET-VLM.

06

Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators

Generative retrieval has emerged as a powerful paradigm for LLM-based recommendation. However, industrial recommender systems often benefit from restricting the output space to a constrained subset of items based on business logic (e.g. enforcing content freshness or product category), which standard autoregressive decoding cannot natively support. Moreover, existing constrained decoding methods that make use of prefix trees (Tries) incur severe latency penalties on hardware accelerators (TPUs/GPUs). In this work, we introduce STATIC (Sparse Transition Matrix-Accelerated Trie Index for Constrained Decoding), an efficient and scalable constrained decoding technique designed specifically for high-throughput LLM-based generative retrieval on TPUs/GPUs. By flattening the prefix tree into a static Compressed Sparse Row (CSR) matrix, we transform irregular tree traversals into fully vectorized sparse matrix operations, unlocking massive efficiency gains on hardware accelerators. We deploy STATIC on a large-scale industrial video recommendation platform serving billions of users. STATIC produces significant product metric impact with minimal latency overhead (0.033 ms per step and 0.25% of inference time), achieving a 948x speedup over a CPU trie implementation and a 47-1033x speedup over a hardware-accelerated binary-search baseline. Furthermore, the runtime overhead of STATIC remains extremely low across a wide range of practical configurations. To the best of our knowledge, STATIC enables the first production-scale deployment of strictly constrained generative retrieval. In addition, evaluation on academic benchmarks demonstrates that STATIC can considerably improve cold-start performance for generative retrieval. Our code is available at https://github.com/youtube/static-constraint-decoding.