NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-01ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Grok 4.3

x.ai has officially announced Grok 4.3, representing the latest iteration of its flagship large language model, now accessible via its developer documentation portal. This release signifies a continued commitment to advancing conversational AI and generative capabilities, offering an updated model for a diverse range of applications. Grok 4.3 is anticipated to deliver notable improvements in key areas such as reasoning, factual accuracy, and complex problem-solving, building upon the foundation of its predecessors. While specific enhancements are detailed in the accompanying documentation, the update is expected to provide developers with more powerful and efficient tools for building intelligent systems. The primary focus is on empowering the developer community, facilitating seamless integration of advanced AI functionalities into their platforms. Through the x.ai API, developers can leverage Grok 4.3's expanded capabilities for use cases spanning intelligent chatbots, sophisticated content creation, and data analysis. This release aims to further democratize access to cutting-edge large language model technology, fostering innovation across the AI landscape.

02

Uber torches 2026 AI budget on Claude Code in four months

Uber has reportedly expended its entire Artificial Intelligence budget allocated for 2026 on 'Claude Code' within an exceptionally short timeframe of just four months. This rapid and substantial investment highlights Uber's aggressive strategy in integrating advanced AI capabilities, potentially indicating a significant pivot towards leveraging large language models or sophisticated code generation tools for its operational needs and product development. The swift depletion of a future year's budget on a single AI solution, specifically mentioning 'Claude Code' (likely referring to Anthropic's Claude AI for coding tasks), underscores the company's commitment to cutting-edge AI adoption. Industry analysts suggest this move could accelerate Uber's internal development cycles, enhance its platform's intelligence, or optimize various aspects of its business through automated code generation and advanced analytical processing. However, such a concentrated expenditure also raises questions about future AI investments and the long-term strategic implications of committing such a significant portion of its resources to a singular technology provider or solution so far in advance.

03

Advanced Quantization Algorithm for LLMs

Intel has unveiled `auto-round`, an advanced quantization algorithm specifically engineered to optimize Large Language Models (LLMs) for efficient deployment. This innovative method aims to significantly enhance the performance and reduce the memory footprint of LLM inference, particularly on Intel CPUs and GPUs, by converting high-precision model weights to lower-precision formats, such as 4-bit. A core objective of `auto-round` is to meticulously preserve model accuracy during this compression process, addressing a common challenge in quantization techniques. By facilitating the deployment of large-scale AI models in environments with limited resources and accelerating inference speeds, `auto-round` is poised to make powerful LLMs more accessible and practical for a wider range of applications. This initiative highlights Intel's ongoing efforts in AI hardware-software co-optimization, fostering the development and widespread adoption of cutting-edge AI technologies by lowering operational costs and energy consumption.

04

Spotify adds 'Verified' badges to distinguish human artists from AI

Spotify has announced the implementation of 'Verified' badges to its platform, a new feature designed to clearly distinguish musical artists who are human from those who utilize artificial intelligence in their creative process. This strategic move by the streaming giant directly addresses growing concerns within the music industry regarding the proliferation of AI-generated content and the potential for confusion among listeners. The introduction of these badges is a response to the increasing challenge of identifying original human artistry amidst a rapidly expanding volume of synthetic music. By providing a distinct indicator, Spotify aims to enhance transparency for its user base, allowing fans to easily identify and support human artists while navigating the evolving digital soundscape. This initiative underscores the platform's commitment to preserving the authenticity of artistic expression and protecting intellectual property. It sets a precedent for how digital music services can adapt to and manage the integration of AI, ensuring trust and clarity in an era where AI's capabilities in music creation are rapidly advancing and influencing the industry's future.

05

Show HN: AI CAD Harness

Zach, a co-founder of Adam, has unveiled their new AI CAD Harness, marking a significant advancement from their earlier text-to-CAD/3D experiments. Responding to feedback from mechanical engineers who expressed a need for integrated assistance rather than opaque 3D model generation, Adam developed a solution that directly integrates with popular CAD platforms such as Onshape and Autodesk Fusion. This "harness" operates as an intelligent AI agent, designed to analyze and comprehend the existing feature tree of a part directly within the CAD software. Critically, it then agentically modifies the design, granting users complete transparency and granular control over the modifications. This approach represents a strategic pivot towards embedding AI as an assistive tool within established professional design workflows, moving beyond mere generative applications to a more interactive and controlled AI-powered design augmentation. The beta version of this innovative tool is currently available, poised to streamline and enhance mechanical design processes.

06

The Gay Jailbreak Technique

This article introduces 'The Gay Jailbreak Technique,' a method ostensibly designed to bypass or circumvent restrictions typically implemented in AI systems, particularly large language models. While specific details of the technique are not provided in the content, the term 'jailbreak' in AI contexts generally refers to prompt engineering strategies or vulnerabilities exploited to elicit responses that models are otherwise programmed to refuse, often due to safety guidelines or content policies. Such techniques are explored by researchers and enthusiasts to test the boundaries of AI capabilities and limitations, evaluate robustness, and identify potential ethical or security vulnerabilities in advanced AI systems. The mention of 'Gay' in the title might suggest a particular methodology, community origin, or a unique approach to this challenging area of AI interaction and control. This technique contributes to the ongoing discourse on AI safety, alignment, and the challenges of deploying highly capable, yet controlled, generative AI.

huggingface

6 stories
01

Heterogeneous Scientific Foundation Model Collaboration

Agentic large language model systems have demonstrated strong capabilities. However, their reliance on language as the universal interface fundamentally limits their applicability to many real-world problems, especially in scientific domains where domain-specific foundation models have been developed to address specialized tasks beyond natural language. In this work, we introduce Eywa, a heterogeneous agentic framework designed to extend language-centric systems to a broader class of scientific foundation models. The key idea of Eywa is to augment domain-specific foundation models with a language-model-based reasoning interface, enabling language models to guide inference over non-linguistic data modalities. This design allows predictive foundation models, which are typically optimized for specialized data and tasks, to participate in higher-level reasoning and decision-making processes within agentic systems. Eywa can serve as a drop-in replacement for a single-agent pipeline (EywaAgent) or be integrated into existing multi-agent systems by replacing traditional agents with specialized agents (EywaMAS). We further investigate a planning-based orchestration framework in which a planner dynamically coordinates traditional agents and Eywa agents to solve complex tasks across heterogeneous data modalities (EywaOrchestra). We evaluate Eywa across a diverse set of scientific domains spanning physical, life, and social sciences. Experimental results demonstrate that Eywa improves performance on tasks involving structured and domain-specific data, while reducing reliance on language-based reasoning through effective collaboration with specialized foundation models.

02

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still struggle with spatial reasoning, persistent state, long-horizon consistency, and causal understanding. We argue that the field should move beyond appearance synthesis toward intelligent visual generation: plausible visuals grounded in structure, dynamics, domain knowledge, and causal relations. To frame this shift, we introduce a five-level taxonomy: Atomic Generation, Conditional Generation, In-Context Generation, Agentic Generation, and World-Modeling Generation, progressing from passive renderers to interactive, agentic, world-aware generators. We analyze key technical drivers, including flow matching, unified understanding-and-generation models, improved visual representations, post-training, reward modeling, data curation, synthetic data distillation, and sampling acceleration. We further show that current evaluations often overestimate progress by emphasizing perceptual quality while missing structural, temporal, and causal failures. By combining benchmark review, in-the-wild stress tests, and expert-constrained case studies, this roadmap offers a capability-centered lens for understanding, evaluating, and advancing the next generation of intelligent visual generation systems.

03

Co-Evolving Policy Distillation

RLVR and OPD have become standard paradigms for post-training. We provide a unified analysis of these two paradigms in consolidating multiple expert capabilities into a single model, identifying capability loss in different ways: mixed RLVR suffers from inter-capability divergence cost, while the pipeline of first training experts and then performing OPD, though avoiding divergence, fails to fully absorb teacher capabilities due to large behavioral pattern gaps between teacher and student. We propose Co-Evolving Policy Distillation (CoPD), which encourages parallel training of experts and introduces OPD during each expert's ongoing RLVR training rather than after complete expert training, with experts serving as mutual teachers (making OPD bidirectional) to co-evolve. This enables more consistent behavioral patterns among experts while maintaining sufficient complementary knowledge throughout. Experiments validate that CoPD achieves all-in-one integration of text, image, and video reasoning capabilities, significantly outperforming strong baselines such as mixed RLVR and MOPD, and even surpassing domain-specific experts. The model parallel training pattern offered by CoPD may inspire a novel training scaling paradigm.

04

ExoActor: Exocentric Video Generation as Generalizable Interactive Humanoid Control

Humanoid control systems have made significant progress in recent years, yet modeling fluent interaction-rich behavior between a robot, its surrounding environment, and task-relevant objects remains a fundamental challenge. This difficulty arises from the need to jointly capture spatial context, temporal dynamics, robot actions, and task intent at scale, which is a poor match to conventional supervision. We propose ExoActor, a novel framework that leverages the generalization capabilities of large-scale video generation models to address this problem. The key insight in ExoActor is to use third-person video generation as a unified interface for modeling interaction dynamics. Given a task instruction and scene context, ExoActor synthesizes plausible execution processes that implicitly encode coordinated interactions between robot, environment, and objects. Such video output is then transformed into executable humanoid behaviors through a pipeline that estimates human motion and executes it via a general motion controller, yielding a task-conditioned behavior sequence. To validate the proposed framework, we implement it as an end-to-end system and demonstrate its generalization to new scenarios without additional real-world data collection. Furthermore, we conclude by discussing limitations of the current implementation and outlining promising directions for future research, illustrating how ExoActor provides a scalable approach to modeling interaction-rich humanoid behaviors, potentially opening a new avenue for generative models to advance general-purpose humanoid intelligence.

05

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at release time and grade mainly the final response, making it difficult to evaluate agents against evolving workflow demand or verify whether a task was executed. We introduce Claw-Eval-Live, a live benchmark for workflow agents that separates a refreshable signal layer, updated across releases from public workflow-demand signals, from a reproducible, time-stamped release snapshot. Each release is constructed from public workflow-demand signals, with ClawHub Top-500 skills used in the current release, and materialized as controlled tasks with fixed fixtures, services, workspaces, and graders. For grading, Claw-Eval-Live records execution traces, audit logs, service state, and post-run workspace artifacts, using deterministic checks when evidence is sufficient and structured LLM judging only for semantic dimensions. The release contains 105 tasks spanning controlled business services and local workspace repair, and evaluates 13 frontier models under a shared public pass rule. Experiments reveal that reliable workflow automation remains far from solved: the leading model passes only 66.7% of tasks and no model reaches 70%. Failures are structured by task family and execution surface, with HR, management, and multi-system business workflows as persistent bottlenecks and local workspace repair comparatively easier but unsaturated. Leaderboard rank alone is insufficient because models with similar pass rates can diverge in overall completion, and task-level discrimination concentrates in a middle band of tasks. Claw-Eval-Live suggests that workflow-agent evaluation should be grounded twice, in fresh external demand and in verifiable agent action.

06

Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence

We introduce Nemotron 3 Nano Omni, the latest model in the Nemotron multimodal series and the first to natively support audio inputs alongside text, images, and video. Nemotron 3 Nano Omni delivers consistent accuracy improvements over its predecessor, Nemotron Nano V2 VL, across all modalities, enabled by advances in architecture, training data and recipes. In particular, Nemotron 3 delivers leading results in real-world document understanding, long audio-video comprehension, and agentic computer use. Built on the highly efficient Nemotron 3 Nano 30B-A3B backbone, Nemotron 3 Nano Omni further incorporates innovative multimodal token-reduction techniques to deliver substantially lower inference latency and higher throughput than other models of similar size. We are releasing model checkpoints in BF16, FP8, and FP4 formats, along with portions of the training data and codebase to facilitate further research and development.