NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-02ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Google releases Gemma 4 open models

Google has announced the release of Gemma 4, a significant update in its family of open models. This new iteration aims to further democratize access to advanced generative AI capabilities for a broad community of developers, researchers, and organizations. The Gemma 4 models are expected to offer enhanced performance, improved efficiency, and greater flexibility for various applications, ranging from sophisticated natural language understanding to content generation tasks. By making these models openly available, Google reinforces its commitment to fostering innovation within the AI ecosystem and accelerating the development of new AI-powered solutions. This release typically includes a spectrum of model sizes, allowing for optimized deployment across different computational environments, from resource-constrained devices to high-performance cloud infrastructure, thereby empowering a wider range of AI experimentation and practical implementation.

02

Qwen3.6-Plus: Towards real world agents

The release of Qwen3.6-Plus marks a significant step in the development of AI agents capable of operating effectively in real-world scenarios. This iteration of the Qwen model series focuses on enhancing the foundational capabilities required for autonomous operation, including improved reasoning, decision-making, and interaction with complex environments. The 'Plus' designation suggests substantial advancements over previous versions, specifically tailored to bridge the gap between theoretical large language models and practical, deployable AI agents. This advancement is crucial for unlocking new possibilities across various industries, enabling more sophisticated automation and intelligent assistance. The emphasis on "real world agents" highlights the commitment to creating robust, reliable, and adaptable AI systems that can handle the unpredictability and dynamism inherent in real-world applications, pushing the boundaries of what large language models can achieve in agentic roles.

03

Lemonade by AMD: a fast and open source local LLM server using GPU and NPU

AMD has introduced 'Lemonade,' an innovative open-source server solution designed for deploying Large Language Models (LLMs) locally. This project emphasizes high performance and efficiency by leveraging both Graphics Processing Units (GPUs) and Neural Processing Units (NPUs) for accelerated AI inference. Lemonade aims to provide developers and users with a fast and accessible platform to run LLMs on their own hardware, addressing growing demands for privacy-preserving AI applications and reducing reliance on cloud-based services. The open-source nature of Lemonade fosters community collaboration, allowing for greater transparency, customization, and continuous improvement in local LLM deployments. This initiative positions AMD as a key player in enabling advanced AI capabilities on client devices and edge computing environments, potentially democratizing access to powerful generative AI technologies through optimized hardware utilization. The server's architecture is engineered to maximize the computational advantages of AMD's hardware, offering a compelling alternative for those seeking robust, performant, and secure local AI solutions outside of conventional data center infrastructures. This development underscores the increasing trend towards decentralized AI processing and the critical role of hardware-agnostic, yet optimized, software platforms in realizing its full potential.

04

IBM Announces Strategic Collaboration with Arm

IBM and Arm have announced a significant strategic collaboration aimed at revolutionizing the landscape of enterprise computing. This partnership brings together IBM's extensive expertise in enterprise solutions, cloud infrastructure, and AI with Arm's leadership in semiconductor IP design, which underpins a vast array of modern computing devices. The alliance is poised to drive innovation in data center technologies, edge computing, and specialized hardware architectures designed for high-performance workloads, including artificial intelligence and machine learning applications. By combining their strengths, the companies intend to develop next-generation platforms that offer enhanced performance, energy efficiency, and security for critical business operations. This initiative signals a concerted effort to foster an open and robust ecosystem for enterprise-grade solutions, potentially influencing future standards and accelerating the adoption of new computing paradigms across various industries. The collaboration is expected to address evolving demands for scalable and adaptable computing resources, ultimately shaping how businesses leverage technology for growth and innovation in the coming years.

05

Mistral secures $830M in debt financing to fund AI data center

Mistral, a prominent AI company, has successfully secured $830 million in debt financing. This substantial capital infusion is earmarked specifically for the development and establishment of a new AI data center. The funding highlights the significant investment flowing into the artificial intelligence sector, particularly towards infrastructure critical for advancing large-scale AI models and services. This move is indicative of a strategic effort by Mistral to bolster its computational capabilities, essential for training and deploying sophisticated AI systems, and to keep pace with the rapidly evolving demands of the AI industry. The investment underscores the increasing need for dedicated high-performance computing resources to support the growth and innovation within the artificial intelligence landscape. This financial backing will enable Mistral to scale its operations and further its position as a key player in the competitive AI market, focusing on enhancing its core technological infrastructure.

06

Cursor 3

Cursor 3 represents the latest iteration of the AI-first code editor, building upon its foundational promise of enhancing developer productivity through intelligent assistance. This new version is expected to introduce significant advancements in its core AI capabilities, offering more sophisticated code generation, intelligent debugging suggestions, and improved code refactoring tools. The update aims to further integrate AI directly into the development workflow, enabling developers to write, understand, and debug code more efficiently. Key improvements are likely to include enhanced natural language understanding for prompts, faster AI response times, and a more intuitive user interface designed to seamlessly incorporate AI-driven features. Cursor 3 positions itself as a critical tool for modern software development, striving to minimize boilerplate code and accelerate the creative process for engineers across various programming languages and projects. The focus remains on making AI an indispensable partner in coding, streamlining complex tasks and fostering a more productive coding environment.

huggingface

6 stories
01

ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers

OpenClaw has rapidly established itself as a leading open-source autonomous agent runtime, offering powerful capabilities including tool integration, local file access, and shell command execution. However, these broad operational privileges introduce critical security vulnerabilities, transforming model errors into tangible system-level threats such as sensitive data leakage, privilege escalation, and malicious third-party skill execution. Existing security measures for the OpenClaw ecosystem remain highly fragmented, addressing only isolated stages of the agent lifecycle rather than providing holistic protection. To bridge this gap, we present ClawKeeper, a real-time security framework that integrates multi-dimensional protection mechanisms across three complementary architectural layers. (1) Skill-based protection operates at the instruction level, injecting structured security policies directly into the agent context to enforce environment-specific constraints and cross-platform boundaries. (2) Plugin-based protection serves as an internal runtime enforcer, providing configuration hardening, proactive threat detection, and continuous behavioral monitoring throughout the execution pipeline. (3) Watcher-based protection introduces a novel, decoupled system-level security middleware that continuously verifies agent state evolution. It enables real-time execution intervention without coupling to the agent's internal logic, supporting operations such as halting high-risk actions or enforcing human confirmation. We argue that this Watcher paradigm holds strong potential to serve as a foundational building block for securing next-generation autonomous agent systems. Extensive qualitative and quantitative evaluations demonstrate the effectiveness and robustness of ClawKeeper across diverse threat scenarios. We release our code.

02

Brevity Constraints Reverse Performance Hierarchies in Language Models

Standard evaluation protocols reveal a counterintuitive phenomenon: on 7.7% of benchmark problems spanning five datasets, larger language models underperform smaller ones by 28.4 percentage points despite 10-100x more parameters. Through systematic evaluation of 31 models (0.5B-405B parameters) across 1,485 problems, we identify the mechanism as spontaneous scale-dependent verbosity that introduces errors through overelaboration. Causal intervention experiments demonstrate this reflects correctable prompt design rather than fundamental capability limitations. Constraining large models to produce brief responses improves accuracy by 26 percentage points and reduces performance gaps by up to two-thirds. Most critically, brevity constraints completely reverse performance hierarchies on mathematical reasoning and scientific knowledge benchmarks, with large models achieving 7.7-15.9 percentage point advantages over small models -- direct inversions of the original gaps. These reversals prove large models possess superior latent capabilities that universal prompting masks. We validate findings through three independent contamination tests and demonstrate inverse scaling operates continuously across the full parameter spectrum, with dataset-specific optimal scales ranging from 0.5B to 3.0B parameters. Our results establish that maximizing large model performance requires scale-aware prompt engineering rather than universal evaluation protocols, with immediate implications for deployment: prompt adaptation simultaneously improves accuracy and reduces computational costs.

03

MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome

Recent progress in deep research systems has been impressive, but evaluation still lags behind real user needs. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process. Most also offer limited multimodal coverage, rely on synthetic tasks that do not reflect real-world query complexity, and cannot be refreshed as knowledge evolves. To address these gaps, we introduce MiroEval, a benchmark and evaluation framework for deep research systems. The benchmark comprises 100 tasks (70 text-only, 30 multimodal), all grounded in real user needs and constructed via a dual-path pipeline that supports periodic updates, enabling a live and evolving setting. The proposed evaluation suite assesses deep research systems along three complementary dimensions: adaptive synthesis quality evaluation with task-specific rubrics, agentic factuality verification via active retrieval and reasoning over both web sources and multimodal attachments, and process-centric evaluation audits how the system searches, reasons, and refines throughout its investigation. Evaluation across 13 systems yields three principal findings: the three evaluation dimensions capture complementary aspects of system capability, with each revealing distinct strengths and weaknesses across systems; process quality serves as a reliable predictor of overall outcome while revealing weaknesses invisible to output-level metrics; and multimodal tasks pose substantially greater challenges, with most systems declining by 3 to 10 points. The MiroThinker series achieves the most balanced performance, with MiroThinker-H1 ranking the highest overall in both settings. Human verification and robustness results confirm the reliability of the benchmark and evaluation framework. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents.

04

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation

Most recent advances in 3D generative modeling rely on diffusion or flow-matching formulations. We instead explore a fully autoregressive alternative and introduce GaussianGPT, a transformer-based model that directly generates 3D Gaussians via next-token prediction, thus facilitating full 3D scene generation. We first compress Gaussian primitives into a discrete latent grid using a sparse 3D convolutional autoencoder with vector quantization. The resulting tokens are serialized and modeled using a causal transformer with 3D rotary positional embedding, enabling sequential generation of spatial structure and appearance. Unlike diffusion-based methods that refine scenes holistically, our formulation constructs scenes step-by-step, naturally supporting completion, outpainting, controllable sampling via temperature, and flexible generation horizons. This formulation leverages the compositional inductive biases and scalability of autoregressive modeling while operating on explicit representations compatible with modern neural rendering pipelines, positioning autoregressive transformers as a complementary paradigm for controllable and context-aware 3D generation.

05

Universal YOCO for Efficient Depth Scaling

The rise of test-time scaling has remarkably boosted the reasoning and agentic proficiency of Large Language Models (LLMs). Yet, standard Transformers struggle to scale inference-time compute efficiently, as conventional looping strategies suffer from high computational overhead and a KV cache that inflates alongside model depth. We present Universal YOCO (YOCO-U), which combines the YOCO decoder-decoder architecture with recursive computation to achieve a synergistic effect greater than either alone. Built on the YOCO framework, YOCO-U implements a Universal Self-Decoder that performs multiple iterations via parameter sharing, while confining the iterative process to shallow, efficient-attention layers. This combination yields a favorable capability-efficiency tradeoff that neither YOCO nor recursion achieves independently. The YOCO architecture provides a constant global KV cache and linear pre-filling, while partial recursion enhances representational depth with limited overhead. Together, YOCO-U improves token utility and scaling behavior while maintaining efficient inference. Empirical results confirm that YOCO-U remains highly competitive in general and long-context benchmarks, demonstrating that the integration of efficient-attention architectures and recursive computation is a promising direction for scalable LLMs.

06

Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding

3D Visual Grounding (3D-VG) aims to localize objects in 3D scenes via natural language descriptions. While recent advancements leveraging Vision-Language Models (VLMs) have explored zero-shot possibilities, they typically suffer from a static workflow relying on preprocessed 3D point clouds, essentially degrading grounding into proposal matching. To bypass this reliance, our core motivation is to decouple the task: leveraging 2D VLMs to resolve complex spatial semantics, while relying on deterministic multi-view geometry to instantiate the 3D structure. Driven by this insight, we propose "Think, Act, Build (TAB)", a dynamic agentic framework that reformulates 3D-VG tasks as a generative 2D-to-3D reconstruction paradigm operating directly on raw RGB-D streams. Specifically, guided by a specialized 3D-VG skill, our VLM agent dynamically invokes visual tools to track and reconstruct the target across 2D frames. Crucially, to overcome the multi-view coverage deficit caused by strict VLM semantic tracking, we introduce the Semantic-Anchored Geometric Expansion, a mechanism that first anchors the target in a reference video clip and then leverages multi-view geometry to propagate its spatial location across unobserved frames. This enables the agent to "Build" the target's 3D representation by aggregating these multi-view features via camera parameters, directly mapping 2D visual cues to 3D coordinates. Furthermore, to ensure rigorous assessment, we identify flaws such as reference ambiguity and category errors in existing benchmarks and manually refine the incorrect queries. Extensive experiments on ScanRefer and Nr3D demonstrate that our framework, relying entirely on open-source models, significantly outperforms previous zero-shot methods and even surpasses fully supervised baselines.