NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-03-20ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Attention Residuals

The 'Attention Residuals' project, developed by MoonshotAI and hosted on GitHub, presents a novel architectural paradigm aimed at enhancing the robustness and computational efficiency of deep learning models, especially those built upon the pervasive transformer architecture. This innovative approach integrates residual connections directly into or in close conjunction with attention mechanisms, addressing critical challenges such as training instability and ensuring more effective gradient propagation through deeper networks. The core idea is to leverage the benefits of residual learning to stabilize and accelerate the training of models that heavily rely on complex attention patterns, thereby potentially unlocking superior performance in tasks like natural language understanding, generation, and other sequence-to-sequence problems. This work signifies a contribution towards optimizing the foundational components of modern AI, offering a blueprint for future scalable and high-performing model designs in areas such as large language models and multimodal AI systems.

02

Launch HN: Sitefire (YC W26) – Automating actions to improve AI visibility

Sitefire, founded by Vincent and Jochen (YC W26) with backgrounds in reinforcement learning and optimization from Stanford, has launched a platform designed to enhance brands' visibility in AI search. The initiative stems from observing a decline in traffic for marketing teams due to Google's AI Overviews. Sitefire aims to move beyond anecdotal evidence and address the challenges of AI search optimization through a data-driven approach. Their system not only monitors but also actively improves traffic derived from AI search engines. A key insight driving their platform is the understanding that unlike traditional search, AI search engines often expand a single user prompt into multiple fan-out queries. Sitefire's core value proposition lies in automating actions to adapt to these evolving AI search paradigms, offering a concrete solution for brands struggling to maintain or improve their online presence in the era of AI-powered information retrieval.

03

MacBook M5 Pro and Qwen3.5 = Local AI Security System

The development of a "Local AI Security System" integrating a MacBook M5 Pro with the Qwen3.5 large language model marks a notable evolution in edge computing and privacy-preserving AI. This innovative approach focuses on deploying sophisticated artificial intelligence capabilities directly on-device, thereby minimizing reliance on cloud infrastructure. By processing sensitive data locally, the system inherently enhances user privacy and security, reducing exposure to potential data breaches or unauthorized surveillance inherent in cloud-centric models. The high-performance M5 Pro chip is instrumental in enabling efficient, real-time inference of advanced models like Qwen3.5, which is critical for demanding security applications requiring immediate analysis and response. Potential applications range from intelligent local surveillance and anomaly detection to secure personal AI assistants, all operating with improved latency and robust offline functionality. This initiative underscores a growing trend toward empowering local devices with advanced AI, fostering greater autonomy, trustworthiness, and resilience in security-critical environments.

04

Wikipedia RFC on banning LLM contributions

Wikipedia has initiated a crucial Request for Comment (RFC) to address the contentious issue of integrating or prohibiting contributions generated by Large Language Models (LLMs) within its articles. This community-driven discussion aims to formulate a definitive policy regarding the acceptable and unacceptable uses of AI tools for content creation on the platform. The ongoing debate centers on several critical concerns, including maintaining editorial accuracy, ensuring rigorous verifiability, and preventing the potential spread of AI-hallucinated or misleading content. Participants are also discussing how to preserve the human-centric ethos of collaborative knowledge building that defines Wikipedia. Stakeholders are weighing the perceived benefits of LLMs in streamlining article creation against the inherent risks associated with automated content, such as challenges in proper attribution, originality, and the dilution of essential editorial oversight. The outcome of this RFC is anticipated to establish crucial guidelines for editors globally, setting a significant precedent for how major online platforms manage the increasing influx of AI-generated text, impacting Wikipedia's long-term integrity and quality.

05

Thousands have swooned over this MAGA dream girl. She's made with AI

A recent phenomenon highlights the growing impact of artificial intelligence in creating convincing digital personas, as a figure dubbed a 'MAGA dream girl,' entirely generated by AI, has captivated thousands online. This development underscores the sophisticated capabilities of generative AI technologies to produce photorealistic images that can garner significant public attention and engagement. The creation and widespread reception of such an AI-fabricated entity raise crucial questions regarding authenticity, the future of online identity, and the potential for AI-generated content to influence social and political landscapes. It also demonstrates the increasing blurring lines between human and artificial creations in the digital sphere, prompting discussions on media literacy and the ethical implications of deploying advanced AI for social interaction and influence. The incident serves as a prominent example of how readily AI-driven synthetic media can be integrated into online culture, challenging traditional notions of personhood and online interaction.

06

Google Search is now using AI to replace headlines

Google Search has begun integrating artificial intelligence to modify or replace news headlines displayed in its search results, signaling a significant shift in how information is presented to users. This development, hinted at as an experimental phase, suggests Google's exploration into using AI for potentially enhancing the relevance or engagement of news content. While the exact motivations are still emerging, this move could aim to optimize click-through rates or personalize headline presentation. However, it also raises important questions regarding journalistic integrity, the potential for altered context, and the transparency of AI-generated content. The implementation marks a notable instance of AI directly influencing how users perceive and interact with news articles at the crucial discovery phase within one of the world's largest information platforms, potentially setting a precedent for future content curation practices.

huggingface

6 stories
01

Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation

We introduce Nemotron-Cascade 2, an open 30B MoE model with 3B activated parameters that delivers best-in-class reasoning and strong agentic capabilities. Despite its compact size, its mathematical and coding reasoning performance approaches that of frontier open models. It is the second open-weight LLM, after DeepSeekV3.2-Speciale-671B-A37B, to achieve Gold Medal-level performance in the 2025 International Mathematical Olympiad (IMO), the International Olympiad in Informatics (IOI), and the ICPC World Finals, demonstrating remarkably high intelligence density with 20x fewer parameters. In contrast to Nemotron-Cascade 1, the key technical advancements are as follows. After SFT on a meticulously curated dataset, we substantially expand Cascade RL to cover a much broader spectrum of reasoning and agentic domains. Furthermore, we introduce multi-domain on-policy distillation from the strongest intermediate teacher models for each domain throughout the Cascade RL process, allowing us to efficiently recover benchmark regressions and sustain strong performance gains along the way. We release the collection of model checkpoint and training data.

02

Memento-Skills: Let Agents Design Agents

We introduce Memento-Skills, a generalist, continually-learnable LLM agent system that functions as an agent-designing agent: it autonomously constructs, adapts, and improves task-specific agents through experience. The system is built on a memory-based reinforcement learning framework with stateful prompts, where reusable skills (stored as structured markdown files) serve as persistent, evolving memory. These skills encode both behaviour and context, enabling the agent to carry forward knowledge across interactions. Starting from simple elementary skills (like Web search and terminal operations), the agent continually improves via the Read--Write Reflective Learning mechanism introduced in Memento~2~wang2025memento2. In the read phase, a behaviour-trainable skill router selects the most relevant skill conditioned on the current stateful prompt; in the write phase, the agent updates and expands its skill library based on new experience. This closed-loop design enables continual learning without updating LLM parameters, as all adaptation is realised through the evolution of externalised skills and prompts. Unlike prior approaches that rely on human-designed agents, Memento-Skills enables a generalist agent to design agents end-to-end for new tasks. Through iterative skill generation and refinement, the system progressively improves its own capabilities. Experiments on the General AI Assistants benchmark and Humanity's Last Exam demonstrate sustained gains, achieving 26.2% and 116.2% relative improvements in overall accuracy, respectively. Code is available at https://github.com/Memento-Teams/Memento-Skills.

03

3DreamBooth: High-Fidelity 3D Subject-Driven Video Generation Model

Creating dynamic, view-consistent videos of customized subjects is highly sought after for a wide range of emerging applications, including immersive VR/AR, virtual production, and next-generation e-commerce. However, despite rapid progress in subject-driven video generation, existing methods predominantly treat subjects as 2D entities, focusing on transferring identity through single-view visual features or textual prompts. Because real-world subjects are inherently 3D, applying these 2D-centric approaches to 3D object customization reveals a fundamental limitation: they lack the comprehensive spatial priors necessary to reconstruct the 3D geometry. Consequently, when synthesizing novel views, they must rely on generating plausible but arbitrary details for unseen regions, rather than preserving the true 3D identity. Achieving genuine 3D-aware customization remains challenging due to the scarcity of multi-view video datasets. While one might attempt to fine-tune models on limited video sequences, this often leads to temporal overfitting. To resolve these issues, we introduce a novel framework for 3D-aware video customization, comprising 3DreamBooth and 3Dapter. 3DreamBooth decouples spatial geometry from temporal motion through a 1-frame optimization paradigm. By restricting updates to spatial representations, it effectively bakes a robust 3D prior into the model without the need for exhaustive video-based training. To enhance fine-grained textures and accelerate convergence, we incorporate 3Dapter, a visual conditioning module. Following single-view pre-training, 3Dapter undergoes multi-view joint optimization with the main generation branch via an asymmetrical conditioning strategy. This design allows the module to act as a dynamic selective router, querying view-specific geometric hints from a minimal reference set. Project page: https://ko-lani.github.io/3DreamBooth.

04

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e.g., VLM features or structural conditions) to mitigate these issues, this reliance severely bottlenecks model robustness and generalization. To overcome this limitation, we present SAMA (factorized Semantic Anchoring and Motion Alignment), a framework that factorizes video editing into semantic anchoring and motion modeling. First, we introduce Semantic Anchoring, which establishes a reliable visual anchor by jointly predicting semantic tokens and video latents at sparse anchor frames, enabling purely instruction-aware structural planning. Second, Motion Alignment pre-trains the same backbone on motion-centric video restoration pretext tasks (cube inpainting, speed perturbation, and tube shuffle), enabling the model to internalize temporal dynamics directly from raw videos. SAMA is optimized with a two-stage pipeline: a factorized pre-training stage that learns inherent semantic-motion representations without paired video-instruction editing data, followed by supervised fine-tuning on paired editing data. Remarkably, the factorized pre-training alone already yields strong zero-shot video editing ability, validating the proposed factorization. SAMA achieves state-of-the-art performance among open-source models and is competitive with leading commercial systems (e.g., Kling-Omni). Code, models, and datasets will be released.

05

LVOmniBench: Pioneering Long Audio-Video Understanding Evaluation for Omnimodal LLMs

Recent advancements in omnimodal large language models (OmniLLMs) have significantly improved the comprehension of audio and video inputs. However, current evaluations primarily focus on short audio and video clips ranging from 10 seconds to 5 minutes, failing to reflect the demands of real-world applications, where videos typically run for tens of minutes. To address this critical gap, we introduce LVOmniBench, a new benchmark designed specifically for the cross-modal comprehension of long-form audio and video. This dataset comprises high-quality videos sourced from open platforms that feature rich audio-visual dynamics. Through rigorous manual selection and annotation, LVOmniBench comprises 275 videos, ranging in duration from 10 to 90 minutes, and 1,014 question-answer (QA) pairs. LVOmniBench aims to rigorously evaluate the capabilities of OmniLLMs across domains, including long-term memory, temporal localization, fine-grained understanding, and multimodal perception. Our extensive evaluation reveals that current OmniLLMs encounter significant challenges when processing extended audio-visual inputs. Open-source models generally achieve accuracies below 35%, whereas the Gemini 3 Pro reaches a peak accuracy of approximately 65%. We anticipate that this dataset, along with our empirical findings, will stimulate further research and the development of advanced models capable of resolving complex cross-modal understanding problems within long-form audio-visual contexts.

06

Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens

Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal architectures. However, current discrete generation methods remain limited to low-dimensional latent tokens (typically 8-32 dims), sacrificing the semantic richness essential for understanding. While high-dimensional pretrained representations (768-1024 dims) could bridge this gap, their discrete generation poses fundamental challenges. In this paper, we present Cubic Discrete Diffusion (CubiD), the first discrete generation model for high-dimensional representations. CubiD performs fine-grained masking throughout the high-dimensional discrete representation -- any dimension at any position can be masked and predicted from partial observations. This enables the model to learn rich correlations both within and across spatial positions, with the number of generation steps fixed at T regardless of feature dimensionality, where T ll hwd. On ImageNet-256, CubiD achieves state-of-the-art discrete generation with strong scaling behavior from 900M to 3.7B parameters. Crucially, we validate that these discretized tokens preserve original representation capabilities, demonstrating that the same discrete tokens can effectively serve both understanding and generation tasks. We hope this work will inspire future research toward unified multimodal architectures. Code is available at: https://github.com/YuqingWang1029/CubiD.