NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-11-14ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

AI World Clocks

The "AI World Clocks" project, available at clocks.brianmoore.com, showcases an innovative and continuous digital art installation where a completely new clock face is rendered every minute. This real-time generative process is powered by an intricate system involving nine distinct artificial intelligence models working in concert. Each AI model contributes uniquely to the visual output, resulting in a constantly evolving series of timepieces that blend various aesthetic styles and interpretations. The initiative serves as a compelling demonstration of advanced generative AI capabilities in producing a continuous stream of novel visual content. By continuously creating and displaying unique clock designs, the project effectively explores the dynamic interplay between art, technology, and AI-driven creativity. It provides valuable insights into how multiple AI systems can be effectively integrated to deliver a perpetually changing visual experience, pushing the boundaries of automated design, real-time artistic generation, and the practical application of diverse AI models in a creative context. This ongoing experiment highlights the potential of AI for dynamic content creation and visual experimentation beyond static outputs.

02

Structured Outputs on the Claude Developer Platform (API)

Anthropic has officially rolled out support for structured outputs on its Claude Developer Platform API, a pivotal enhancement designed to provide developers with predictable and machine-readable responses from its powerful large language models. This capability allows AI applications to reliably generate outputs in specified formats such as JSON, XML, or YAML, moving beyond the inherent variability of free-form text generation. This development is crucial for integrating LLMs into complex programmatic workflows, facilitating tasks like automated data extraction, content generation adhering to predefined schemas, and the development of more robust and reliable AI agents. By enabling developers to precisely define and enforce output formats, the Claude platform significantly minimizes parsing errors, streamlines subsequent data processing, and substantially improves the overall consistency and efficiency of AI-powered applications. This strategic move is poised to accelerate the creation of more sophisticated, dependable, and deeply integrated AI solutions across various industries.

03

AGI fantasy is a blocker to actual engineering

This article contends that the prevailing

04

Show HN: Chirp – Local Windows dictation with ParakeetV3 no executable required

Chirp is a new local dictation application designed for Windows users operating in restricted environments where `.exe` installations are prohibited and cloud-based speech services are blocked. This tool enables accurate and fast dictation without relying on a GPU or transmitting audio data to external cloud servers. Chirp leverages NVIDIA’s ParakeetV3 model, specifically the Parakeet TDT 0.6B v3 ONNX bundle, and is engineered to run entirely locally using Python, with `uv` managing its processes. This innovative solution offers a viable and accessible alternative to conventional Windows dictation options or GPU-intensive setups, particularly in locked-down systems. The project highlights its performance, which is comparable to advanced models like Whisper-large-v3, demonstrating similar word error rates, and emphasizes its ease of deployment for anyone capable of executing Python scripts.

05

I think nobody wants AI in Firefox, Mozilla

Recent discussions and expressed user sentiment indicate a notable resistance or lack of enthusiasm towards the integration of artificial intelligence functionalities within the Mozilla Firefox web browser. The core sentiment, encapsulated by the statement, "I think nobody wants AI in Firefox, Mozilla," suggests that users may prioritize aspects such as browser performance, privacy, and simplicity over the perceived benefits of AI-driven features. This trend highlights a significant challenge for browser developers like Mozilla, who must carefully balance innovation with user expectations and potential concerns regarding data handling, system resource consumption, and the overall user experience. The prevailing sentiment underscores a segment of the user base that values a minimalist and privacy-respecting browsing environment, prompting a critical evaluation of how AI enhancements align with these core user values and the browser's established identity.

06

Nvidia is gearing up to sell servers instead of just GPUs and components

Nvidia is reportedly initiating a significant strategic shift, transitioning from primarily selling Graphics Processing Units (GPUs) and related components to offering complete Artificial Intelligence (AI) servers. This move, highlighted by J.P. Morgan, signals a strong push towards vertical integration within the AI hardware ecosystem. This "master plan" by CEO Jensen Huang is expected to substantially boost Nvidia's profit margins by capturing a larger share of the value chain. By delivering fully integrated AI server solutions, potentially starting with platforms like "Vera Rubin," Nvidia aims to provide comprehensive, optimized systems directly to customers, rather than just supplying the underlying hardware. This strategic evolution positions Nvidia as a more holistic solution provider in the burgeoning AI infrastructure market, competing more directly with server manufacturers while leveraging its dominant position in AI accelerators. This could streamline deployment for AI developers and enterprises, offering a more tightly coupled hardware-software stack.

GitHub

2 stories
01

TrendRadar

TrendRadar is an open-source, lightweight hotspot assistant designed to provide users with relevant news and information quickly, with deployment taking as little as 30 seconds. It aggregates trending topics from over 11 major platforms, including Zhihu, Douyin, Weibo, and Baidu. The system features intelligent push strategies (daily, current, incremental), precise content filtering using custom keywords, and advanced hotspot trend analysis to track news evolution. A personalized algorithm sorts aggregated content based on rank, frequency, and hotness. It supports real-time notifications across multiple channels like WeChat Work, Feishu, DingTalk, Telegram, Email, and ntfy, alongside multi-device web reports via GitHub Pages. A significant V3.0.0 update introduced AI intelligent analysis powered by the Model Context Protocol (MCP), enabling natural language queries and deep data insights with 13 analytical tools. TrendRadar empowers users to proactively obtain desired information, reducing reliance on platform-specific algorithms, making it ideal for investors, self-media creators, and PR professionals.

02

Agent Development Kit (ADK) for Go

The Agent Development Kit (ADK) for Go is an open-source, code-first toolkit designed to streamline the building, evaluation, and deployment of sophisticated AI agents. It applies robust software development principles to agent creation, offering a flexible and modular framework for orchestrating workflows from simple tasks to complex multi-agent systems. While optimized for Google's Gemini, ADK is model-agnostic and deployment-agnostic, ensuring broad compatibility. The Go version specifically leverages Go's strengths in concurrency and performance, making it ideal for cloud-native agent applications. Key features include idiomatic Go design, a rich tool ecosystem for diverse agent capabilities, code-first development for ultimate flexibility and testability, and robust support for modular multi-agent systems and cloud-native deployment environments like Google Cloud Run.

huggingface

6 stories
01

UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist

While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce UniVA, an open-source, omni-capable multi-agent framework for next-generation video generalists that unifies video understanding, segmentation, editing, and generation into cohesive workflows. UniVA employs a Plan-and-Act dual-agent architecture that drives a highly automated and proactive workflow: a planner agent interprets user intentions and decomposes them into structured video-processing steps, while executor agents execute these through modular, MCP-based tool servers (for analysis, generation, editing, tracking, etc.). Through a hierarchical multi-level memory (global knowledge, task context, and user-specific preferences), UniVA sustains long-horizon reasoning, contextual continuity, and inter-agent communication, enabling interactive and self-reflective video creation with full traceability. This design enables iterative and any-conditioned video workflows (e.g., text/image/video-conditioned generation rightarrow multi-round editing rightarrow object segmentation rightarrow compositional synthesis) that were previously cumbersome to achieve with single-purpose models or monolithic video-language models. We also introduce UniVA-Bench, a benchmark suite of multi-step video tasks spanning understanding, editing, segmentation, and generation, to rigorously evaluate such agentic video systems. Both UniVA and UniVA-Bench are fully open-sourced, aiming to catalyze research on interactive, agentic, and general-purpose video intelligence for the next generation of multimodal AI systems. (https://univa.online/)

02

PAN: A World Model for General, Interactable, and Long-Horizon World Simulation

A world model enables an intelligent agent to imagine, predict, and reason about how the world evolves in response to its actions, and accordingly to plan and strategize. While recent video generation models produce realistic visual sequences, they typically operate in the prompt-to-full-video manner without causal control, interactivity, or long-horizon consistency required for purposeful reasoning. Existing world modeling efforts, on the other hand, often focus on restricted domains (e.g., physical, game, or 3D-scene dynamics) with limited depth and controllability, and struggle to generalize across diverse environments and interaction formats. In this work, we introduce PAN, a general, interactable, and long-horizon world model that predicts future world states through high-quality video simulation conditioned on history and natural language actions. PAN employs the Generative Latent Prediction (GLP) architecture that combines an autoregressive latent dynamics backbone based on a large language model (LLM), which grounds simulation in extensive text-based knowledge and enables conditioning on language-specified actions, with a video diffusion decoder that reconstructs perceptually detailed and temporally coherent visual observations, to achieve a unification between latent space reasoning (imagination) and realizable world dynamics (reality). Trained on large-scale video-action pairs spanning diverse domains, PAN supports open-domain, action-conditioned simulation with coherent, long-term dynamics. Extensive experiments show that PAN achieves strong performance in action-conditioned world simulation, long-horizon forecasting, and simulative reasoning compared to other video generators and world models, taking a step towards general world models that enable predictive simulation of future world states for reasoning and acting.

03

Solving a Million-Step LLM Task with Zero Errors

LLMs have achieved remarkable breakthroughs in reasoning, insights, and tool use, but chaining these abilities into extended processes at the scale of those routinely executed by humans, organizations, and societies has remained out of reach. The models have a persistent error rate that prevents scale-up: for instance, recent experiments in the Towers of Hanoi benchmark domain showed that the process inevitably becomes derailed after at most a few hundred steps. Thus, although LLM research is often still benchmarked on tasks with relatively few dependent logical steps, there is increasing attention on the ability (or inability) of LLMs to perform long range tasks. This paper describes MAKER, the first system that successfully solves a task with over one million LLM steps with zero errors, and, in principle, scales far beyond this level. The approach relies on an extreme decomposition of a task into subtasks, each of which can be tackled by focused microagents. The high level of modularity resulting from the decomposition allows error correction to be applied at each step through an efficient multi-agent voting scheme. This combination of extreme decomposition and error correction makes scaling possible. Thus, the results suggest that instead of relying on continual improvement of current LLMs, massively decomposed agentic processes (MDAPs) may provide a way to efficiently solve problems at the level of organizations and societies.

04

Black-Box On-Policy Distillation of Large Language Models

Black-box distillation creates student large language models (LLMs) by learning from a proprietary teacher model's text outputs alone, without access to its internal logits or parameters. In this work, we introduce Generative Adversarial Distillation (GAD), which enables on-policy and black-box distillation. GAD frames the student LLM as a generator and trains a discriminator to distinguish its responses from the teacher LLM's, creating a minimax game. The discriminator acts as an on-policy reward model that co-evolves with the student, providing stable, adaptive feedback. Experimental results show that GAD consistently surpasses the commonly used sequence-level knowledge distillation. In particular, Qwen2.5-14B-Instruct (student) trained with GAD becomes comparable to its teacher, GPT-5-Chat, on the LMSYS-Chat automatic evaluation. The results establish GAD as a promising and effective paradigm for black-box LLM distillation.

05

One Small Step in Latent, One Giant Leap for Pixels: Fast Latent Upscale Adapter for Your Diffusion Models

Diffusion models struggle to scale beyond their training resolutions, as direct high-resolution sampling is slow and costly, while post-hoc image super-resolution (ISR) introduces artifacts and additional latency by operating after decoding. We present the Latent Upscaler Adapter (LUA), a lightweight module that performs super-resolution directly on the generator's latent code before the final VAE decoding step. LUA integrates as a drop-in component, requiring no modifications to the base model or additional diffusion stages, and enables high-resolution synthesis through a single feed-forward pass in latent space. A shared Swin-style backbone with scale-specific pixel-shuffle heads supports 2x and 4x factors and remains compatible with image-space SR baselines, achieving comparable perceptual quality with nearly 3x lower decoding and upscaling time (adding only +0.42 s for 1024 px generation from 512 px, compared to 1.87 s for pixel-space SR using the same SwinIR architecture). Furthermore, LUA shows strong generalization across the latent spaces of different VAEs, making it easy to deploy without retraining from scratch for each new decoder. Extensive experiments demonstrate that LUA closely matches the fidelity of native high-resolution generation while offering a practical and efficient path to scalable, high-fidelity image synthesis in modern diffusion pipelines.

06

Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following

Recent progress in large language models (LLMs) has led to impressive performance on a range of tasks, yet advanced instruction following (IF)-especially for complex, multi-turn, and system-prompted instructions-remains a significant challenge. Rigorous evaluation and effective training for such capabilities are hindered by the lack of high-quality, human-annotated benchmarks and reliable, interpretable reward signals. In this work, we introduce AdvancedIF (we will release this benchmark soon), a comprehensive benchmark featuring over 1,600 prompts and expert-curated rubrics that assess LLMs ability to follow complex, multi-turn, and system-level instructions. We further propose RIFL (Rubric-based Instruction-Following Learning), a novel post-training pipeline that leverages rubric generation, a finetuned rubric verifier, and reward shaping to enable effective reinforcement learning for instruction following. Extensive experiments demonstrate that RIFL substantially improves the instruction-following abilities of LLMs, achieving a 6.7% absolute gain on AdvancedIF and strong results on public benchmarks. Our ablation studies confirm the effectiveness of each component in RIFL. This work establishes rubrics as a powerful tool for both training and evaluating advanced IF in LLMs, paving the way for more capable and reliable AI systems.