NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-20ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Qwen3.6-Max-Preview: Smarter, Sharper, Still Evolving

Qwen has unveiled a preview of its latest large language model, Qwen3.6-Max-Preview, signaling a continued commitment to advancing AI capabilities with promises of enhanced intelligence and sharper performance. This iteration is positioned as a significant step in the model's evolution, suggesting improvements in areas such as reasoning, comprehension, and generation, crucial for complex natural language processing applications. While specific technical details regarding its architecture, training methodologies, or comprehensive benchmark results are eagerly awaited, the "Max-Preview" designation implies a more robust and capable version than its predecessors. The "Still Evolving" tagline highlights the dynamic nature of AI research, indicating that development is ongoing and further optimizations or features may be integrated into subsequent releases. This strategic preview underscores the competitive innovation within the large language model space, aiming to empower developers and researchers with more sophisticated tools for building advanced AI applications and contributing to the broader landscape of generative AI.

02

Kimi K2.6: Advancing open-source coding

Kimi K2.6 has been announced as a significant update aimed at advancing the landscape of open-source coding. This new version is positioned to introduce enhanced capabilities and tools specifically designed to empower developers within the open-source community. While detailed technical specifications were not extensively provided in the initial brief, the release underscores a strategic focus on improving efficiency and innovation in collaborative coding environments. It suggests potential improvements in areas such as intelligent code assistance, automated code generation, or streamlined integration with existing open-source workflows. The 'advancing open-source coding' theme implies that Kimi K2.6 is tailored to address common challenges faced by open-source contributors, potentially offering more intuitive development experiences, improved code quality, and faster project iteration cycles. This update is anticipated to solidify Kimi's role as a vital resource for developers seeking robust, accessible, and community-driven solutions for their coding endeavors, thereby fostering greater participation and productivity in global open-source initiatives.

03

Deezer says 44% of songs uploaded to its platform daily are AI-generated

Music streaming service Deezer has revealed a striking trend in content creation, reporting that approximately 44% of all songs uploaded to its platform on a daily basis are generated by artificial intelligence. This significant figure underscores the accelerating influence of generative AI technologies within the music industry, presenting both opportunities and considerable challenges for streaming platforms and artists alike. The surge in AI-created music raises critical questions regarding content authenticity, intellectual property rights, and fair compensation mechanisms for human artists. For platforms like Deezer, it necessitates the development of advanced detection and classification systems to manage the influx of AI-generated content, potentially leading to new strategies for content curation and monetization. This development signals a pivotal shift in how music is produced, distributed, and consumed, highlighting the need for industry-wide discussions on ethical guidelines and regulatory frameworks to navigate the evolving landscape of AI-powered creative works. The statistic from Deezer serves as a potent indicator of the profound impact AI is having on creative industries, pushing stakeholders to adapt to a future where machines play a substantial role in artistic output.

04

OpenClaw isn't fooling me. I remember MS-DOS

The blog post, titled "OpenClaw isn't fooling me. I remember MS-DOS," reflects the author's skeptical view on an emergent AI agent solution, likely referred to as "OpenClaw." By drawing a historical parallel to MS-DOS, the author suggests potential concerns regarding its proprietary nature, complexity, or lack of transparency, which might limit user control in contemporary AI implementations. The accompanying article on flyingpenguin.com aims to guide users in constructing an "OpenClaw-free" alternative, emphasizing the development of secure, always-on local AI agents. This initiative aligns with a broader industry interest in decentralized AI processing and the pursuit of greater autonomy and data privacy within AI applications. The central argument advocates for building AI agents that operate independently and securely within a user's local environment, contrasting with perceived drawbacks of existing or developing centralized AI frameworks. The discussion seeks to empower users to create custom AI solutions that prioritize openness and local operational integrity over potentially restrictive external platforms.

05

NSA is using Anthropic's Mythos despite blacklist

The National Security Agency (NSA) is reportedly utilizing Anthropic's advanced artificial intelligence platform, Mythos, in its operations, a significant development that comes despite an apparent internal or governmental blacklist. This situation underscores the accelerating integration of sophisticated AI models within critical intelligence infrastructure, signaling a strategic imperative for the NSA to leverage cutting-edge technology even when it conflicts with existing policy or procurement restrictions. The decision to employ Mythos suggests the NSA perceives substantial operational advantages or unique capabilities offered by Anthropic's AI, which are deemed essential for its national security and intelligence gathering missions. This deployment raises important questions regarding the effectiveness of such blacklists, the procedural frameworks for technology adoption within sensitive government agencies, and the broader implications for AI governance and oversight across the defense and intelligence sectors. The continued reliance on commercial AI tools by a top intelligence agency highlights the dual-use nature of advanced AI technologies and the complex balance required between innovation, national security imperatives, and regulatory compliance.

06

AI chatbots could be making you stupider

The rapid proliferation of AI chatbots is raising concerns about their potential negative impact on human cognitive abilities, suggesting that increased reliance could diminish critical thinking and problem-solving skills. The central argument posits that as individuals offload cognitive tasks and information synthesis to AI systems, their own capacities for independent thought, analytical reasoning, and memory retention may degrade. This phenomenon, often described as 'cognitive outsourcing' or 'digital amnesia,' necessitates a careful re-evaluation of human-AI interaction paradigms. Experts are exploring the long-term implications for education, professional development, and societal intelligence, advocating for mindful AI integration strategies. There is a growing call for research into the neurocognitive effects of sustained AI use, aiming to guide the development of AI tools that foster intellectual growth and augment human intelligence, rather than inadvertently leading to cognitive decline.

huggingface

6 stories
01

Qwen3.5-Omni Technical Report

In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-plus achieves SOTA results across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding. Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (MoE) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis, often caused by encoding efficiency discrepancies between text and speech tokenizers, we introduce ARIA. ARIA dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding and speech generation across 10 languages with human-like emotional nuance. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding.

02

PersonaVLM: Long-Term Personalized Multimodal LLMs

Multimodal Large Language Models (MLLMs) serve as daily assistants for millions. However, their ability to generate responses aligned with individual preferences remains limited. Prior approaches enable only static, single-turn personalization through input augmentation or output alignment, and thus fail to capture users' evolving preferences and personality over time (see Fig.1). In this paper, we introduce PersonaVLM, an innovative personalized multimodal agent framework designed for long-term personalization. It transforms a general-purpose MLLM into a personalized assistant by integrating three key capabilities: (a) Remembering: It proactively extracts and summarizes chronological multimodal memories from interactions, consolidating them into a personalized database. (b) Reasoning: It conducts multi-turn reasoning by retrieving and integrating relevant memories from the database. (c) Response Alignment: It infers the user's evolving personality throughout long-term interactions to ensure outputs remain aligned with their unique characteristics. For evaluation, we establish Persona-MME, a comprehensive benchmark comprising over 2,000 curated interaction cases, designed to assess long-term MLLM personalization across seven key aspects and 14 fine-grained tasks. Extensive experiments validate our method's effectiveness, improving the baseline by 22.4% (Persona-MME) and 9.8% (PERSONAMEM) under a 128k context, while outperforming GPT-4o by 5.2% and 2.0%, respectively. Project page: https://PersonaVLM.github.io.

03

Hierarchical Codec Diffusion for Video-to-Speech Generation

Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the hierarchical nature of speech, which spans coarse speaker-aware semantics to fine-grained prosodic details. This oversight hinders direct alignment between visual and speech features at specific hierarchical levels during property matching. In this paper, leveraging the hierarchical structure of Residual Vector Quantization (RVQ)-based codec, we propose HiCoDiT, a novel Hierarchical Codec Diffusion Transformer that exploits the inherent hierarchy of discrete speech tokens to achieve strong audio-visual alignment. Specifically, since lower-level tokens encode coarse speaker-aware semantics and higher-level tokens capture fine-grained prosody, HiCoDiT employs low-level and high-level blocks to generate tokens at different levels. The low-level blocks condition on lip-synchronized motion and facial identity to capture speaker-aware content, while the high-level blocks use facial expression to modulate prosodic dynamics. Finally, to enable more effective coarse-to-fine conditioning, we propose a dual-scale adaptive instance layer normalization that jointly captures global vocal style through channel-wise normalization and local prosody dynamics through temporal-wise normalization. Extensive experiments demonstrate that HiCoDiT outperforms baselines in fidelity and expressiveness, highlighting the potential of discrete modelling for VTS. The code and speech demo are both available at https://github.com/Jiaxin-Ye/HiCoDiT.

04

Elucidating the SNR-t Bias of Diffusion Probabilistic Models

Diffusion Probabilistic Models have demonstrated remarkable performance across a wide range of generative tasks. However, we have observed that these models often suffer from a Signal-to-Noise Ratio-timestep (SNR-t) bias. This bias refers to the misalignment between the SNR of the denoising sample and its corresponding timestep during the inference phase. Specifically, during training, the SNR of a sample is strictly coupled with its timestep. However, this correspondence is disrupted during inference, leading to error accumulation and impairing the generation quality. We provide comprehensive empirical evidence and theoretical analysis to substantiate this phenomenon and propose a simple yet effective differential correction method to mitigate the SNR-t bias. Recognizing that diffusion models typically reconstruct low-frequency components before focusing on high-frequency details during the reverse denoising process, we decompose samples into various frequency components and apply differential correction to each component individually. Extensive experiments show that our approach significantly improves the generation quality of various diffusion models (IDDPM, ADM, DDIM, A-DPM, EA-DPM, EDM, PFGM++, and FLUX) on datasets of various resolutions with negligible computational overhead. The code is at https://github.com/AMAP-ML/DCW.

05

DiPO: Disentangled Perplexity Policy Optimization for Fine-grained Exploration-Exploitation Trade-Off

Reinforcement Learning with Verifiable Rewards (RLVR) has catalyzed significant advances in the reasoning capabilities of Large Language Models (LLMs). However, effectively managing the exploration and exploitation trade-off remains a critical challenge. In this paper, we fully analyze the exploration and exploitation dilemma of extremely hard and easy samples during the training and propose a new fine-grained trade-off mechanism. Concretely, we introduce a perplexity space disentangling strategy that divides the sample space into distinct exploration (high perplexity) and exploitation (low perplexity) subspaces, thereby mining fine-grained samples requiring exploration-exploitation trade-off. Subsequently, we propose a bidirectional reward allocation mechanism with a minimum impact on verification rewards to implement perplexity-guided exploration and exploitation, enabling more stable policy optimization. Finally, we have evaluated our method on two mainstream tasks: mathematical reasoning and function calling, and experimental results demonstrate the superiority of the proposed method, confirming its effectiveness in enhancing LLM performance by fine-grained exploration-exploitation trade-off.

06

Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems

Retrieval-Augmented Generation (RAG) systems critically depend on effective document chunking strategies to balance retrieval quality, latency, and operational cost. Traditional chunking approaches, such as fixed-size, rule-based, or fully agentic chunking, often suffer from high token consumption, redundant text generation, limited scalability, and poor debuggability, especially for large-scale web content ingestion. In this paper, we propose Web Retrieval-Aware Chunking (W-RAC), a novel, cost-efficient chunking framework designed specifically for web-based documents. W-RAC decouples text extraction from semantic chunk planning by representing parsed web content as structured, ID-addressable units and leveraging large language models (LLMs) only for retrieval-aware grouping decisions rather than text generation. This significantly reduces token usage, eliminates hallucination risks, and improves system observability.Experimental analysis and architectural comparison demonstrate that W-RAC achieves comparable or better retrieval performance than traditional chunking approaches while reducing chunking-related LLM costs by an order of magnitude.