NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-03-17ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

GPT 5.4 Mini and Nano

OpenAI has reportedly announced the introduction of new, more compact iterations of its foundational language models, designated as GPT-5.4 Mini and Nano. These new models are expected to extend the accessibility and deployment flexibility of OpenAI's advanced artificial intelligence capabilities. While specific technical details regarding their architecture, training data, and performance benchmarks are yet to be fully disclosed, the 'Mini' and 'Nano' nomenclature typically indicates a focus on optimized efficiency, reduced computational overhead, and potentially lower operational costs. This strategic expansion suggests OpenAI's ongoing commitment to developing a diversified portfolio of models tailored for various application scales, from resource-constrained environments to specialized tasks requiring lightweight, high-performance AI. The release aims to address the growing demand for smaller, more efficient large language models capable of running on edge devices or within applications where the full power of larger models like GPT-5.4 might be excessive or economically unfeasible. This move aligns with broader industry trends towards making powerful AI more ubiquitous and adaptable across a wider array of platforms and use cases.

02

Nvidia's DLSS 5 uses generative AI to boost photorealism in video games

Nvidia is advancing its Deep Learning Super Sampling (DLSS) technology with the introduction of DLSS 5, which incorporates generative AI to significantly enhance photorealism in video games. This next iteration aims to push the boundaries of visual fidelity, offering gamers more lifelike graphics and immersive experiences. By leveraging generative AI models, DLSS 5 will intelligently reconstruct and upscale game frames, filling in missing details and creating highly realistic textures and lighting effects that were previously challenging to achieve through traditional rendering techniques. This innovative approach moves beyond simple upscaling and frame generation, utilizing AI to "imagine" and create visual information, thereby improving both performance and visual quality simultaneously. While the immediate application is focused on boosting realism in gaming, Nvidia's strategic move suggests broader ambitions for generative AI's role in real-time graphics and simulations, potentially extending into professional visualization, virtual reality, and other interactive media beyond the gaming industry. This development marks a significant step in integrating advanced AI directly into the graphics pipeline, promising a new era of visual quality for digital environments.

03

Show HN: Antfly: Distributed, Multimodal Search and Memory and Graphs in Go

Antfly is a newly introduced distributed document database and search engine built in Go, designed to integrate full-text, vector, and graph search capabilities into a unified system. It offers a single-binary deployment, making it suitable for both local development and small-scale distributed multimodal search and memory applications. A significant differentiator is its native ML inference service, Termite, which facilitates vector search directly within the platform, thereby reducing the dependency on external API calls. The system supports comprehensive multimodal indexing for images, audio, and video data, alongside advanced features like MongoDB-style in-place updates and streaming Retrieval Augmented Generation (RAG). From a distributed systems perspective, Antfly employs a multi-Raft setup, leveraging etcd's library and backed by CockroachDB's Pebble storage engine. It further enhances its distributed architecture by assigning independent Raft groups to metadata and data shards, ensuring high availability and consistency in a developer-friendly package.

04

Show HN: March Madness Bracket Challenge for AI Agents Only

A novel March Madness bracket challenge has been launched, designed exclusively for AI agents rather than human participants. The system allows human users to prompt their AI agents with a URL, enabling the agents to autonomously read API documentation, register, select winners for all 63 games, and submit a bracket. A live leaderboard then tracks the performance of participating AI agents throughout the tournament. The creator encountered unique design challenges in developing an agent-first user experience. A key solution involved differentiating content delivery: agents accessing the homepage receive plain-text API instructions, while human users view the standard visual website. Early observations revealed agents attempting to browse the site using tools like Playwright instead of utilizing the API. This led to further refinements, including detecting HeadlessChrome to serve agent-specific HTML, emphasizing the importance of dedicated AI agent user experience considerations in the development process.

05

Node.js blocks PR from dev because he used Claude Code to create it

The Node.js project recently took the decision to block a pull request (PR) submitted by a developer, citing the use of 'Claude Code' for its generation. This incident brings to the forefront a burgeoning debate within the open-source community concerning the acceptance and integration of AI-generated code. The decision by Node.js maintainers signals a potential precedent or at least a significant point of discussion regarding the policies and guidelines for contributions that leverage large language models or other generative AI tools. Key considerations include code ownership, intellectual property rights, potential licensing issues, the maintainability and quality of AI-generated code, and the broader implications for project governance. This event underscores the evolving landscape of software development, where AI tools are becoming increasingly prevalent, prompting project leads to formulate clear stances on their usage and ethical considerations within collaborative, community-driven environments. It highlights the challenges in balancing innovation with established best practices and ensuring the integrity of foundational projects.

06

Unsloth Studio

Unsloth Studio represents a new development environment or platform designed to streamline and accelerate the process of fine-tuning Large Language Models (LLMs). Building upon the core Unsloth library, which is celebrated for its significant speed and memory efficiency improvements in LLM training, the Studio aims to provide a more integrated and user-friendly experience. It likely offers tools and features that abstract away complexities, allowing developers and researchers to more easily deploy, customize, and optimize LLMs for specific tasks and datasets. This initiative seeks to democratize advanced LLM fine-tuning, making it accessible to a broader audience without requiring deep expertise in low-level optimization techniques. By leveraging Unsloth's proven capabilities, the Studio is poised to dramatically reduce the computational resources and time traditionally associated with high-performance LLM development, thereby fostering innovation and faster iteration cycles in the AI community. Its introduction marks a step towards more efficient and scalable LLM application development.

huggingface

6 stories
01

AI Can Learn Scientific Taste

Great scientists have strong judgement and foresight, closely tied to what we call scientific taste. Here, we use the term to refer to the capacity to judge and propose research ideas with high potential impact. However, most relative research focuses on improving an AI scientist's executive capability, while enhancing an AI's scientific taste remains underexplored. In this work, we propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference modeling and alignment problem. For preference modeling, we train Scientific Judge on 700K field- and time-matched pairs of high- vs. low-citation papers to judge ideas. For preference alignment, using Scientific Judge as a reward model, we train a policy model, Scientific Thinker, to propose research ideas with high potential impact. Experiments show Scientific Judge outperforms SOTA LLMs (e.g., GPT-5.2, Gemini 3 Pro) and generalizes to future-year test, unseen fields, and peer-review preference. Furthermore, Scientific Thinker proposes research ideas with higher potential impact than baselines. Our findings show that AI can learn scientific taste, marking a key step toward reaching human-level AI scientists.

02

OpenSeeker: Democratizing Frontier Search Agents by Fully Open-Sourcing Training Data

Deep search capabilities have become an indispensable competency for frontier Large Language Model (LLM) agents, yet the development of high-performance search agents remains dominated by industrial giants due to a lack of transparent, high-quality training data. This persistent data scarcity has fundamentally hindered the progress of the broader research community in developing and innovating within this domain. To bridge this gap, we introduce OpenSeeker, the first fully open-source search agent (i.e., model and data) that achieves frontier-level performance through two core technical innovations: (1) Fact-grounded scalable controllable QA synthesis, which reverse-engineers the web graph via topological expansion and entity obfuscation to generate complex, multi-hop reasoning tasks with controllable coverage and complexity. (2) Denoised trajectory synthesis, which employs a retrospective summarization mechanism to denoise the trajectory, therefore promoting the teacher LLMs to generate high-quality actions. Experimental results demonstrate that OpenSeeker, trained (a single training run) on only 11.7k synthesized samples, achieves state-of-the-art performance across multiple benchmarks including BrowseComp, BrowseComp-ZH, xbench-DeepSearch, and WideSearch. Notably, trained with simple SFT, OpenSeeker significantly outperforms the second-best fully open-source agent DeepDive (e.g., 29.5% v.s. 15.3% on BrowseComp), and even surpasses industrial competitors such as Tongyi DeepResearch (trained via extensive continual pre-training, SFT, and RL) on BrowseComp-ZH (48.4% v.s. 46.7%). We fully open-source the complete training dataset and the model weights to democratize frontier search agent research and foster a more transparent, collaborative ecosystem.

03

Tri-Prompting: Video Diffusion with Unified Control over Scene, Subject, and Motion

Recent video diffusion models have made remarkable strides in visual quality, yet precise, fine-grained control remains a key bottleneck that limits practical customizability for content creation. For AI video creators, three forms of control are crucial: (i) scene composition, (ii) multi-view consistent subject customization, and (iii) camera-pose or object-motion adjustment. Existing methods typically handle these dimensions in isolation, with limited support for multi-view subject synthesis and identity preservation under arbitrary pose changes. This lack of a unified architecture makes it difficult to support versatile, jointly controllable video. We introduce Tri-Prompting, a unified framework and two-stage training paradigm that integrates scene composition, multi-view subject consistency, and motion control. Our approach leverages a dual-condition motion module driven by 3D tracking points for background scenes and downsampled RGB cues for foreground subjects. To ensure a balance between controllability and visual realism, we further propose an inference ControlNet scale schedule. Tri-Prompting supports novel workflows, including 3D-aware subject insertion into any scenes and manipulation of existing subjects in an image. Experimental results demonstrate that Tri-Prompting significantly outperforms specialized baselines such as Phantom and DaS in multi-view subject identity, 3D consistency, and motion accuracy.

04

HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification

Can AI make progress on important, unsolved mathematical problems? Large language models are now capable of sophisticated mathematical and scientific reasoning, but whether they can perform novel research is still widely debated and underexplored. We introduce HorizonMath, a benchmark of over 100 predominantly unsolved problems spanning 8 domains in computational and applied mathematics, paired with an open-source evaluation framework for automated verification. Our benchmark targets a class of problems where discovery is hard, requiring meaningful mathematical insight, but verification is computationally efficient and simple. Because these solutions are unknown, HorizonMath is immune to data contamination, and most state-of-the-art models score near 0%. Existing research-level benchmarks instead rely on formal proof verification or manual review, both of which are expensive to scale. Using this platform, we find two problems for which GPT 5.4 Pro proposes solutions that improve on the best-known published results, representing potential novel contributions (pending expert review). We release HorizonMath as an open challenge and a growing community resource, where correct solutions to problems in the unsolved problem classes could constitute novel results in the mathematical literature.

05

OxyGen: Unified KV Cache Management for Vision-Language-Action Models under Multi-Task Parallelism

Embodied AI agents increasingly require parallel execution of multiple tasks, such as manipulation, conversation, and memory construction, from shared observations under distinct time constraints. Recent Mixture-of-Transformers (MoT) Vision-Language-Action Models (VLAs) architecturally support such heterogeneous outputs, yet existing inference systems fail to achieve efficient multi-task parallelism for on-device deployment due to redundant computation and resource contention. We identify isolated KV cache management as the root cause. To address this, we propose unified KV cache management, an inference paradigm that treats KV cache as a first-class shared resource across tasks and over time. This abstraction enables two key optimizations: cross-task KV sharing eliminates redundant prefill of shared observations, while cross-frame continuous batching decouples variable-length language decoding from fixed-rate action generation across control cycles. We implement this paradigm for "pi"_{0.5}, the most popular MoT VLA, and evaluate under representative robotic configurations. OxyGen achieves up to 3.7times speedup over isolated execution, delivering over 200 tokens/s language throughput and 70 Hz action frequency simultaneously without action quality degradation.

06

Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models

Pre-trained Large Language Model (LLM) exhibits broad capabilities, yet, for specific tasks or domains their attainment of higher accuracy and more reliable reasoning generally depends on post-training through Supervised Fine-Tuning (SFT) or Reinforcement Learning (RL). Although often treated as distinct methodologies, recent theoretical and empirical developments demonstrate that SFT and RL are closely connected. This study presents a comprehensive and unified perspective on LLM post-training with SFT and RL. We first provide an in-depth overview of both techniques, examining their objectives, algorithmic structures, and data requirements. We then systematically analyze their interplay, highlighting frameworks that integrate SFT and RL, hybrid training pipelines, and methods that leverage their complementary strengths. Drawing on a representative set of recent application studies from 2023 to 2025, we identify emerging trends, characterize the rapid shift toward hybrid post-training paradigms, and distill key takeaways that clarify when and why each method is most effective. By synthesizing theoretical insights, practical methodologies, and empirical evidence, this study establishes a coherent understanding of SFT and RL within a unified framework and outlines promising directions for future research in scalable, efficient, and generalizable LLM post-training.