NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-15ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference

Google's Gemma 4 model has achieved a significant milestone, demonstrating native execution capabilities directly on Apple's iPhone devices. This groundbreaking development enables full offline AI inference, allowing the sophisticated large language model to perform complex artificial intelligence tasks without requiring an internet connection or reliance on cloud-based servers. This advancement is crucial for enhancing privacy, reducing latency, and improving the accessibility of powerful AI applications for mobile users, bypassing the need for continuous data transfer to remote servers. The ability to run Gemma 4 natively on a smartphone highlights ongoing progress in optimizing large language models for resource-constrained edge devices, potentially opening new avenues for personalized AI applications, enhanced data privacy, and improved performance by eliminating network latency. This breakthrough underscores the increasing viability of bringing advanced AI capabilities directly to consumers' hands, independent of external infrastructure, marking a key evolution in mobile AI and edge computing, and setting a new precedent for on-device machine learning performance.

02

CPUs Aren't Dead. Gemma2B Out Scored GPT-3.5 Turbo on Test That Made It Famous

A significant report highlights that Gemma2B, an open-source language model, has demonstrably outperformed OpenAI's GPT-3.5 Turbo on a widely recognized benchmark test, thereby reigniting discussions around the necessity of specialized hardware for advanced AI workloads. This achievement underscores the continued viability and surprising capabilities of Central Processing Units (CPUs) in executing sophisticated AI inference tasks, pushing back against the narrative that only high-end Graphical Processing Units (GPUs) are suitable. The successful benchmark performance by Gemma2B suggests substantial advancements in model architecture and optimization techniques that enable powerful large language models to run efficiently on more conventional computing infrastructure. This development is crucial for democratizing access to powerful AI, reducing the barriers of entry associated with expensive specialized hardware, and fostering broader innovation across various sectors. The results reinforce the dynamic evolution of AI model development, emphasizing the potential for highly optimized models to deliver top-tier performance on diverse hardware platforms, making AI more accessible and resource-efficient.

03

Gemini Robotics-ER 1.6

Google DeepMind has announced Gemini Robotics-ER 1.6, signaling a significant update in its ongoing initiative to integrate advanced artificial intelligence capabilities, particularly from the Gemini model, into robotic systems. While the brief announcement does not delve into specific technical details or the full scope of enhancements, this version release indicates continuous progress in the field of embodied AI. The update likely introduces improvements in areas such as robot learning, task execution, enhanced perception, and autonomous decision-making, aiming to create more adaptable and capable robotic agents. This development reinforces DeepMind's commitment to advancing AI-driven robotics, pushing the boundaries of how intelligent systems interact with and learn from the physical world. Such advancements are critical for developing robots that can operate effectively in complex, unstructured environments, fostering the development of more general-purpose AI within physical forms and potentially impacting various sectors requiring sophisticated automation.

04

Show HN: Tier – Adaptive tool routing that makes small LLMs 10pt more accurate

Tier is an innovative project introduced on Hacker News that focuses on enhancing the accuracy of smaller Large Language Models (LLMs) by up to 10 percentage points. This improvement is achieved through an adaptive tool routing mechanism. The core idea behind Tier is to intelligently select and utilize external tools or functions based on the context of a given task, effectively augmenting the capabilities of less powerful LLMs. By dynamically routing queries to appropriate tools, Tier allows these models to overcome their inherent limitations, leading to more precise and reliable outputs. This approach represents a significant step towards making smaller, more efficient LLMs competitive with their larger counterparts, opening avenues for broader deployment in resource-constrained environments while maintaining high performance.

05

AI-Assisted Cognition Endangers Human Development

The article "AI-Assisted Cognition Endangers Human Development" presents a critical viewpoint on the increasing reliance on artificial intelligence for cognitive tasks. It argues that while AI offers significant advantages in efficiency and data processing, an overdependence on these tools could paradoxically hinder the natural evolution of human intellectual capabilities. The piece likely delves into how constantly available AI assistance for activities such as critical thinking, problem-solving, and memory recall might lead to an atrophy of these fundamental cognitive skills over an extended period. This potential dependency could impact various aspects of life, including educational progress, professional efficacy, and personal growth, thereby diminishing the inherent human capacity for deep learning, creative innovation, and independent thought. The author presumably raises concerns about the long-term societal ramifications, suggesting that future generations, accustomed to outsourced cognition, may struggle with complex, novel challenges without technological aid. The article advocates for a balanced approach, emphasizing the necessity of nurturing human cognitive resilience alongside advancements in AI to ensure technology remains an augmenting tool rather than a substitute for essential human intellectual development.

06

AI ruling prompts warnings from US lawyers: Your chats could be used against you

A recent legal ruling pertaining to artificial intelligence has triggered significant warnings from US legal professionals, advising individuals that their digital communications, particularly those involving interactions with AI systems, could potentially be admissible as evidence in court. This development signals an evolving legal landscape where the privacy of content generated by or exchanged through AI platforms may no longer be assumed. Lawyers are urging the public to exercise heightened caution regarding the nature of their online chats and AI engagements, emphasizing that these interactions could carry unforeseen legal consequences. The ruling highlights the critical intersection of advanced technology, data privacy, and judicial interpretation, suggesting a profound shift in how digital information is perceived and leveraged within the legal system. This situation underscores an increasing imperative for clearer regulatory frameworks and greater user awareness concerning data sovereignty and the potential evidentiary weight of AI-processed communications in future litigation.

huggingface

6 stories
01

KnowRL: Boosting LLM Reasoning via Reinforcement Learning with Minimal-Sufficient Knowledge Guidance

RLVR improves reasoning in large language models, but its effectiveness is often limited by severe reward sparsity on hard problems. Recent hint-based RL methods mitigate sparsity by injecting partial solutions or abstract templates, yet they typically scale guidance by adding more tokens, which introduce redundancy, inconsistency, and extra training overhead. We propose KnowRL (Knowledge-Guided Reinforcement Learning), an RL training framework that treats hint design as a minimal-sufficient guidance problem. During RL training, KnowRL decomposes guidance into atomic knowledge points (KPs) and uses Constrained Subset Search (CSS) to construct compact, interaction-aware subsets for training. We further identify a pruning interaction paradox -- removing one KP may help while removing multiple such KPs can hurt -- and explicitly optimize for robust subset curation under this dependency structure. We train KnowRL-Nemotron-1.5B from OpenMath-Nemotron-1.5B. Across eight reasoning benchmarks at the 1.5B scale, KnowRL-Nemotron-1.5B consistently outperforms strong RL and hinting baselines. Without KP hints at inference, KnowRL-Nemotron-1.5B reaches 70.08 average accuracy, already surpassing Nemotron-1.5B by +9.63 points; with selected KPs, performance improves to 74.16, establishing a new state of the art at this scale. The model, curated training data, and code are publicly available at https://github.com/Hasuer/KnowRL.

02

ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents

GUI agents drive applications through their visual interfaces instead of programmatic APIs, interacting with arbitrary software via taps, swipes, and keystrokes, reaching a long tail of applications that CLI-based agents cannot. Yet progress in this area is bottlenecked less by modeling capacity than by the absence of a coherent full-stack infrastructure: online RL training suffers from environment instability and closed pipelines, evaluation protocols drift silently across works, and trained agents rarely reach real users on real devices. We present ClawGUI, an open-source framework addressing these three gaps within a single harness. ClawGUI-RL provides the first open-source GUI agent RL infrastructure with validated support for both parallel virtual environments and real physical devices, integrating GiGPO with a Process Reward Model for dense step-level supervision. ClawGUI-Eval enforces a fully standardized evaluation pipeline across 6 benchmarks and 11+ models, achieving 95.8% reproduction against official baselines. ClawGUI-Agent brings trained agents to Android, HarmonyOS, and iOS through 12+ chat platforms with hybrid CLI-GUI control and persistent personalized memory. Trained end to end within this pipeline, ClawGUI-2B achieves 17.1% Success Rate on MobileWorld GUI-Only, outperforming the same-scale MAI-UI-2B baseline by 6.0%.

03

Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning

We describe the pre-training, post-training, and quantization of Nemotron 3 Super, a 120 billion (active 12 billion) parameter hybrid Mamba-Attention Mixture-of-Experts model. Nemotron 3 Super is the first model in the Nemotron 3 family to 1) be pre-trained in NVFP4, 2) leverage LatentMoE, a new Mixture-of-Experts architecture that optimizes for both accuracy per FLOP and accuracy per parameter, and 3) include MTP layers for inference acceleration through native speculative decoding. We pre-trained Nemotron 3 Super on 25 trillion tokens followed by post-training using supervised fine tuning (SFT) and reinforcement learning (RL). The final model supports up to 1M context length and achieves comparable accuracy on common benchmarks, while also achieving up to 2.2x and 7.5x higher inference throughput compared to GPT-OSS-120B and Qwen3.5-122B, respectively. Nemotron 3 Super datasets, along with the base, post-trained, and quantized checkpoints, are open-sourced on HuggingFace.

04

Lyra 2.0: Explorable Generative 3D Worlds

Recent advances in video generation enable a new paradigm for 3D scene creation: generating camera-controlled videos that simulate scene walkthroughs, then lifting them to 3D via feed-forward reconstruction techniques. This generative reconstruction approach combines the visual fidelity and creative capacity of video models with 3D outputs ready for real-time rendering and simulation. Scaling to large, complex environments requires 3D-consistent video generation over long camera trajectories with large viewpoint changes and location revisits, a setting where current video models degrade quickly. Existing methods for long-horizon generation are fundamentally limited by two forms of degradation: spatial forgetting and temporal drifting. As exploration proceeds, previously observed regions fall outside the model's temporal context, forcing the model to hallucinate structures when revisited. Meanwhile, autoregressive generation accumulates small synthesis errors over time, gradually distorting scene appearance and geometry. We present Lyra 2.0, a framework for generating persistent, explorable 3D worlds at scale. To address spatial forgetting, we maintain per-frame 3D geometry and use it solely for information routing -- retrieving relevant past frames and establishing dense correspondences with the target viewpoints -- while relying on the generative prior for appearance synthesis. To address temporal drifting, we train with self-augmented histories that expose the model to its own degraded outputs, teaching it to correct drift rather than propagate it. Together, these enable substantially longer and 3D-consistent video trajectories, which we leverage to fine-tune feed-forward reconstruction models that reliably recover high-quality 3D scenes.

05

Toward Autonomous Long-Horizon Engineering for ML Research

Autonomous AI research has advanced rapidly, but long-horizon ML research engineering remains difficult: agents must sustain coherent progress across task comprehension, environment setup, implementation, experimentation, and debugging over hours or days. We introduce AiScientist, a system for autonomous long-horizon engineering for ML research built on a simple principle: strong long-horizon performance requires both structured orchestration and durable state continuity. To this end, AiScientist combines hierarchical orchestration with a permission-scoped File-as-Bus workspace: a top-level Orchestrator maintains stage-level control through concise summaries and a workspace map, while specialized agents repeatedly re-ground on durable artifacts such as analyses, plans, code, and experimental evidence rather than relying primarily on conversational handoffs, yielding thin control over thick state. Across two complementary benchmarks, AiScientist improves PaperBench score by 10.54 points on average over the best matched baseline and achieves 81.82 Any Medal% on MLE-Bench Lite. Ablation studies further show that File-as-Bus protocol is a key driver of performance, reducing PaperBench by 6.41 points and MLE-Bench Lite by 31.82 points when removed. These results suggest that long-horizon ML research engineering is a systems problem of coordinating specialized work over durable project state, rather than a purely local reasoning problem.

06

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and how it is organized. A de facto standard approach to video tokenization is to represent a video as a spatiotemporal 3D grid of tokens, each capturing the corresponding local information in the original signal. This requires the downstream model that consumes the tokens, e.e., a text-to-video model, to learn to predict all low-level details "pixel-by-pixel" irrespective of the video's inherent complexity, leading to high learning complexity. We present VideoFlexTok, which represents videos with a variable-length sequence of tokens structured in a coarse-to-fine manner -- where the first tokens (emergently) capture abstract information, such as semantics and motion, and later tokens add fine-grained details. The generative flow decoder enables realistic video reconstructions from any token count. This representation structure allows adapting the token count according to downstream needs and encoding videos longer than the baselines with the same budget. We evaluate VideoFlexTok on class- and text-to-video generative tasks and show that it leads to more efficient training compared to 3D grid tokens, e.g., achieving comparable generation quality (gFVD and ViCLIP Score) with a 5x smaller model (1.1B vs 5.2B). Finally, we demonstrate how VideoFlexTok can enable long video generation without prohibitive computational cost by training a text-to-video model on 10-second 81-frame videos with only 672 tokens, 8x fewer than a comparable 3D grid tokenizer.