NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-11-27ENGLISH EDITION
This issue
—
All time
—

Hacker News

5 stories
01

Replace your boss before they replace you

The concept 'Replace your boss before they replace you,' presented by replaceyourboss.ai, introduces a provocative perspective on the increasing integration of artificial intelligence into professional environments. This initiative appears to advocate for individuals to proactively adopt and leverage AI technologies to automate tasks, enhance productivity, and potentially streamline operational workflows that are traditionally within the purview of managerial roles. The underlying premise suggests that advanced AI tools are becoming capable of handling complex decision-making, resource allocation, and project management, thereby empowering employees to operate with greater autonomy and efficiency. By embracing these AI-driven solutions, professionals are encouraged to future-proof their careers and adapt to an evolving workforce landscape where AI plays a pivotal role in optimizing performance and redefining traditional hierarchical structures. This platform likely provides tools or insights into how AI agents can perform tasks that mimic or augment managerial functions, urging a strategic shift in how human capital interacts with intelligent automation.

02

We're losing our voice to LLMs

The article "We're losing our voice to LLMs" addresses a growing concern regarding the pervasive influence of Large Language Models (LLMs) on content creation and communication. It argues that increasing reliance on AI for generating text risks diluting the unique human voice, leading to a homogenization of expression across various platforms. This trend may diminish individual writing styles, authentic perspectives, and the nuanced emotional depth traditionally found in human-authored content. The piece likely delves into the broader implications for creativity, personal identity, and the future landscape of digital interaction, suggesting that while LLMs offer efficiency, they inadvertently standardize communication. Concerns are raised about the potential erosion of critical thinking, the loss of distinctive narratives, and the challenge of discerning original human thought amidst a surge of AI-generated content. The author prompts reflection on balancing AI's benefits with the imperative to preserve genuine human authorship and unique individual contributions.

03

Show HN: Runprompt – run .prompt files from the command line

Runprompt is an innovative single-file Python script introduced to streamline the execution of Large Language Model (LLM) prompts directly from the command line. Developed with inspiration from Google's Dotprompt format, which integrates frontmatter with Handlebars templates, Runprompt simplifies complex LLM interactions into a more manageable, programmatic workflow. The tool allows users to treat prompts as 'first-class programs,' enabling powerful Unix-style piping and sequential chaining of multiple prompts for sophisticated operations. Its core capabilities include robust templating, the generation of structured outputs defined by clear schemas, and the ability to integrate prompts seamlessly. This approach facilitates various applications, such as performing sentiment analysis by piping raw text into a pre-configured prompt that outputs a JSON-formatted result. Runprompt offers a developer-centric solution for enhanced LLM prompt engineering, aiming to provide a simpler, more efficient method for running and managing LLM interactions in a command-line environment.

04

The current state of the theory that GPL propagates to AI models

The article delves into the current understanding and ongoing theoretical discussions concerning the propagation of the GNU General Public License (GPL) to artificial intelligence models. A central point of contention is whether an AI model, specifically one trained using code or datasets licensed under the GPL, becomes subject to the GPL's copyleft stipulations. This raises critical legal and intellectual property questions regarding whether a trained model constitutes a "derivative work" of its training data in a way that triggers the GPL's viral clause, potentially requiring the model itself to be released under an open-source license. The implications are substantial for companies and developers aiming to commercialize AI technologies built upon components that might fall under GPL. Navigating this complex intersection of software licensing, copyright law, and the unique operational aspects of AI models is essential for ensuring legal compliance, mitigating risks, and promoting sustainable AI development practices. The debate highlights the need for clarity as AI continues to integrate deeply with open-source ecosystems.

05

Show HN: Era – Open-source local sandbox for AI agents

ERA is an open-source, local sandbox solution designed to address the critical security challenge of isolating AI agents, particularly in light of vulnerabilities like jailbreaking that could enable cyber attacks. Developed as a response to incidents where AI models, such as Claude, have been exploited to execute malicious code, ERA provides microVM-based sandboxing with hardware-level security. This approach offers enhanced isolation for AI-generated code, positioning it as a significantly safer alternative to traditional containerization methods. By leveraging hardware-backed security, ERA ensures that potential threats originating from AI agent activities remain contained, preventing them from compromising the host system. The project, available on GitHub, aims to establish a robust and secure environment for developing and deploying AI agents, inviting community feedback and contributions to further strengthen its capabilities in safeguarding AI-driven operations.

GitHub

2 stories
01

TrendRadar

TrendRadar is a lightweight, easily deployable hotspot assistant that aggregates news from over 11 mainstream platforms, including Zhihu, Douyin, and Weibo. It offers intelligent push strategies (daily, current, incremental) and precise content filtering based on user-defined keywords, eliminating information overload. The platform features real-time trend analysis, a personalized hotspot algorithm, and supports multi-channel notifications via WeChat Work, Feishu, DingTalk, Telegram, Email, ntfy, Bark, and Slack. Unique to v3.0.0 is an AI intelligent analysis system, powered by the Model Context Protocol (MCP), enabling natural language queries and 13 analytical tools for deep news data insights. Deployment is streamlined via GitHub Actions and Docker, with data persistence for historical records. TrendRadar is designed for investors, content creators, and general users seeking to monitor market trends, track public sentiment, or acquire tailored information efficiently, reducing reliance on algorithmic recommendations.

02

Agent Development Kit (ADK) for Go

The Agent Development Kit (ADK) for Go is an open-source, code-first toolkit designed to streamline the building, evaluation, and deployment of sophisticated AI agents. It applies robust software development principles to AI agent creation, offering a flexible and modular framework for orchestrating workflows from simple tasks to complex multi-agent systems. While optimized for Google's Gemini, ADK is model-agnostic and deployment-agnostic, ensuring compatibility across various AI models and deployment environments, including integration with other frameworks. This Go version specifically leverages Go's strengths in concurrency and performance, making it ideal for developers creating cloud-native agent applications. Key features include an idiomatic Go design, a rich tool ecosystem for diverse agent capabilities, code-first development for ultimate flexibility and testability, and strong support for containerization and cloud-native environments like Google Cloud Run. This enables developers to build scalable and robust AI solutions with granular control.

huggingface

6 stories
01

Latent Collaboration in Multi-Agent Systems

Multi-agent systems (MAS) extend large language models (LLMs) from independent single-model reasoning to coordinative system-level intelligence. While existing LLM agents depend on text-based mediation for reasoning and communication, we take a step forward by enabling models to collaborate directly within the continuous latent space. We introduce LatentMAS, an end-to-end training-free framework that enables pure latent collaboration among LLM agents. In LatentMAS, each agent first performs auto-regressive latent thoughts generation through last-layer hidden embeddings. A shared latent working memory then preserves and transfers each agent's internal representations, ensuring lossless information exchange. We provide theoretical analyses establishing that LatentMAS attains higher expressiveness and lossless information preservation with substantially lower complexity than vanilla text-based MAS. In addition, empirical evaluations across 9 comprehensive benchmarks spanning math and science reasoning, commonsense understanding, and code generation show that LatentMAS consistently outperforms strong single-model and text-based MAS baselines, achieving up to 14.6% higher accuracy, reducing output token usage by 70.8%-83.7%, and providing 4x-4.3x faster end-to-end inference. These results demonstrate that our new latent collaboration framework enhances system-level reasoning quality while offering substantial efficiency gains without any additional training. Code and data are fully open-sourced at https://github.com/Gen-Verse/LatentMAS.

02

Inferix: A Block-Diffusion based Next-Generation Inference Engine for World Simulation

World models serve as core simulators for fields such as agentic AI, embodied AI, and gaming, capable of generating long, physically realistic, and interactive high-quality videos. Moreover, scaling these models could unlock emergent capabilities in visual perception, understanding, and reasoning, paving the way for a new paradigm that moves beyond current LLM-centric vision foundation models. A key breakthrough empowering them is the semi-autoregressive (block-diffusion) decoding paradigm, which merges the strengths of diffusion and autoregressive methods by generating video tokens in block-applying diffusion within each block while conditioning on previous ones, resulting in more coherent and stable video sequences. Crucially, it overcomes limitations of standard video diffusion by reintroducing LLM-style KV Cache management, enabling efficient, variable-length, and high-quality generation. Therefore, Inferix is specifically designed as a next-generation inference engine to enable immersive world synthesis through optimized semi-autoregressive decoding processes. This dedicated focus on world simulation distinctly sets it apart from systems engineered for high-concurrency scenarios (like vLLM or SGLang) and from classic video diffusion models (such as xDiTs). Inferix further enhances its offering with interactive video streaming and profiling, enabling real-time interaction and realistic simulation to accurately model world dynamics. Additionally, it supports efficient benchmarking through seamless integration of LV-Bench, a new fine-grained evaluation benchmark tailored for minute-long video generation scenarios. We hope the community will work together to advance Inferix and foster world model exploration.

03

Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy

The synthesis of synchronized audio-visual content is a key challenge in generative AI, with open-source models facing challenges in robust audio-video alignment. Our analysis reveals that this issue is rooted in three fundamental challenges of the joint diffusion process: (1) Correspondence Drift, where concurrently evolving noisy latents impede stable learning of alignment; (2) inefficient global attention mechanisms that fail to capture fine-grained temporal cues; and (3) the intra-modal bias of conventional Classifier-Free Guidance (CFG), which enhances conditionality but not cross-modal synchronization. To overcome these challenges, we introduce Harmony, a novel framework that mechanistically enforces audio-visual synchronization. We first propose a Cross-Task Synergy training paradigm to mitigate drift by leveraging strong supervisory signals from audio-driven video and video-driven audio generation tasks. Then, we design a Global-Local Decoupled Interaction Module for efficient and precise temporal-style alignment. Finally, we present a novel Synchronization-Enhanced CFG (SyncCFG) that explicitly isolates and amplifies the alignment signal during inference. Extensive experiments demonstrate that Harmony establishes a new state-of-the-art, significantly outperforming existing methods in both generation fidelity and, critically, in achieving fine-grained audio-visual synchronization.

04

Monet: Reasoning in Latent Visual Space Beyond Images and Language

"Thinking with images" has emerged as an effective paradigm for advancing visual reasoning, extending beyond text-only chains of thought by injecting visual evidence into intermediate reasoning steps. However, existing methods fall short of human-like abstract visual thinking, as their flexibility is fundamentally limited by external tools. In this work, we introduce Monet, a training framework that enables multimodal large language models (MLLMs) to reason directly within the latent visual space by generating continuous embeddings that function as intermediate visual thoughts. We identify two core challenges in training MLLMs for latent visual reasoning: high computational cost in latent-vision alignment and insufficient supervision over latent embeddings, and address them with a three-stage distillation-based supervised fine-tuning (SFT) pipeline. We further reveal a limitation of applying GRPO to latent reasoning: it primarily enhances text-based reasoning rather than latent reasoning. To overcome this, we propose VLPO (Visual-latent Policy Optimization), a reinforcement learning method that explicitly incorporates latent embeddings into policy gradient updates. To support SFT, we construct Monet-SFT-125K, a high-quality text-image interleaved CoT dataset containing 125K real-world, chart, OCR, and geometry CoTs. Our model, Monet-7B, shows consistent gains across real-world perception and reasoning benchmarks and exhibits strong out-of-distribution generalization on challenging abstract visual reasoning tasks. We also empirically analyze the role of each training component and discuss our early unsuccessful attempts, providing insights for future developments in visual latent reasoning. Our model, data, and code are available at https://github.com/NOVAglow646/Monet.

05

MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots

Grounding natural-language instructions into continuous control for quadruped robots remains a fundamental challenge in vision language action. Existing methods struggle to bridge high-level semantic reasoning and low-level actuation, leading to unstable grounding and weak generalization in the real world. To address these issues, we present MobileVLA-R1, a unified vision-language-action framework that enables explicit reasoning and continuous control for quadruped robots. We construct MobileVLA-CoT, a large-scale dataset of multi-granularity chain-of-thought (CoT) for embodied trajectories, providing structured reasoning supervision for alignment. Built upon this foundation, we introduce a two-stage training paradigm that combines supervised CoT alignment with GRPO reinforcement learning to enhance reasoning consistency, control stability, and long-horizon execution. Extensive evaluations on VLN and VLA tasks demonstrate superior performance over strong baselines, with approximately a 5% improvement. Real-world deployment on a quadruped robot validates robust performance in complex environments. Code: https://github.com/AIGeeksGroup/MobileVLA-R1. Website: https://aigeeksgroup.github.io/MobileVLA-R1.

06

Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma

Reinforcement Learning from Human Feedback (RLHF) is widely used for aligning large language models, yet practitioners face a persistent puzzle: improving safety often reduces fairness, scaling to diverse populations becomes computationally intractable, and making systems robust often amplifies majority biases. We formalize this tension as the Alignment Trilemma: no RLHF system can simultaneously achieve (i) epsilon-representativeness across diverse human values, (ii) polynomial tractability in sample and compute complexity, and (iii) delta-robustness against adversarial perturbations and distribution shift. Through a complexity-theoretic analysis integrating statistical learning theory and robust optimization, we prove that achieving both representativeness (epsilon <= 0.01) and robustness (delta <= 0.001) for global-scale populations requires Omega(2^{d_context}) operations, which is super-polynomial in the context dimensionality. We show that current RLHF implementations resolve this trilemma by sacrificing representativeness: they collect only 10^3--10^4 samples from homogeneous annotator pools while 10^7--10^8 samples are needed for true global representation. Our framework provides a unified explanation for documented RLHF pathologies including preference collapse, sycophancy, and systematic bias amplification. We conclude with concrete directions for navigating these fundamental trade-offs through strategic relaxations of alignment requirements.