NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-10-14DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

GPT-5o-mini hallucinates medical residency applicant grades

A recent analysis, drawing insights from a Thalamus GME blog post, has brought to light a significant issue with OpenAI's GPT-5o-mini: its tendency to 'hallucinate' medical residency applicant grades. This flaw manifests as the generation of inaccurate or entirely fabricated academic records, posing substantial integrity risks for medical residency applications and the broader medical education ecosystem. The findings underscore persistent challenges in ensuring data fidelity and factual accuracy within Large Language Models, particularly when processing sensitive and high-stakes personal information. This incident emphasizes the imperative for robust validation mechanisms and meticulous implementation of AI tools in critical domains where precision is paramount. While AI offers immense potential for efficiency gains, its current limitations in reliable factual recall and its capacity to produce synthetic yet plausible data necessitate stringent human oversight and exhaustive verification processes, especially concerning sensitive applicant data. This serves as a crucial case study on the ethical and practical considerations for deploying advanced AI systems in professional and regulatory contexts.

02

Why the push for Agentic when models can barely follow a simple instruction?

The provided title, "Why the push for Agentic when models can barely follow a simple instruction?", encapsulates a significant debate within the artificial intelligence community concerning the strategic direction of AI development. It raises a pertinent question about the industry's fervent pursuit of complex, autonomous "agentic" AI systems, particularly when current large language models (LLMs) frequently exhibit inconsistencies and difficulties in adhering to seemingly straightforward instructions. This query underscores a perceived gap between the aspirational goals for advanced AI agents, capable of independent decision-making and task execution, and the practical, often frustrating, challenges encountered with the reliability and predictability of foundational models. The discussion implies a critical need to bolster the robustness and instruction-following capabilities of core AI models. Experts suggest that a solid and consistent foundation in basic task execution is indispensable for agentic AI to achieve effective, reliable, and trustworthy operation, questioning the scalability and utility of sophisticated architectures if fundamental issues of model comprehension and adherence remain unresolved. This ongoing dialogue emphasizes a potential misalignment in development priorities.

03

Show HN: Metorial (YC F25) – Vercel for MCP

Metorial, a YC F25 startup founded by Wen and Tobias, has introduced an integration platform designed to streamline the server-side deployment and management of AI agents. Positioned as the "Vercel for MCP," Metorial addresses critical challenges encountered when running MCP servers, such as complex Docker configurations, per-user OAuth flows, scaling concurrent sessions, and establishing observability. The platform automates these infrastructure-heavy tasks, significantly reducing the setup time required to connect AI agents with external tools and data. Metorial offers an open catalog of approximately 600 pre-configured MCP servers, including integrations with services like GitHub, Slack, Google Drive, and Salesforce, enabling users to deploy them with minimal effort. Users can also integrate custom MCP servers or modify existing ones. The platform fully manages the OAuth process, handling client ID/secret, token refresh, and ensuring isolated environments for each user, thus simplifying AI agent integration from weeks to just a few clicks.

04

Intel Announces Inference-Optimized Xe3P Graphics Card with 160GB VRAM

Intel has officially unveiled its new Xe3P Graphics Card, specifically engineered and optimized for inference workloads. This latest addition to Intel's hardware lineup features an impressive 160GB of VRAM, positioning it as a high-capacity solution for demanding AI and machine learning applications in enterprise and data center environments. The Xe3P aims to significantly accelerate data center AI deployments, offering robust performance for critical tasks such as real-time analytics, large-scale model serving, and various advanced computer vision and natural language processing inference operations. The substantial 160GB VRAM capacity is particularly noteworthy, indicating capabilities for efficiently handling larger AI models and extensive datasets during the inference phase, which is crucial for modern AI infrastructure. This strategic release underscores Intel's ongoing commitment to expanding its presence and competitiveness within the rapidly evolving AI acceleration hardware ecosystem, providing a powerful option for businesses seeking specialized hardware for their AI deployment needs.

05

NVIDIA DGX Spark In-Depth Review: A New Standard for Local AI Inference

A comprehensive in-depth review of the NVIDIA DGX Spark has been published, highlighting its potential to set a new benchmark for local AI inference solutions. This analysis delves into the architectural innovations and performance metrics of the DGX Spark system, which is designed to accelerate AI workloads directly on premises, offering significant advantages in data privacy, low latency, and operational efficiency compared to cloud-based alternatives. The review likely examines its computational power, memory configurations, software stack integrations, and overall suitability for demanding AI model deployment scenarios, such as real-time analytics, secure enterprise AI, and advanced research. Initial findings suggest that the DGX Spark delivers exceptional throughput and energy efficiency, positioning it as a critical infrastructure component for organizations seeking robust and scalable local AI capabilities. The evaluation underscores its role in democratizing high-performance AI inference, making advanced AI applications more accessible and manageable outside of large data centers.

06

Why your boss isn't worried about AI – "can't you just turn it off?"

The article, titled 'Why your boss isn't worried about AI – "can't you just turn it off?"', explores the significant disconnect between executive-level perceptions and the complex realities of integrating artificial intelligence within organizations. It addresses the common, yet often misguided, assumption among some leaders that AI systems are easily controllable and reversible, akin to conventional software. This overlooks the intricate, emergent behaviors and deeply embedded nature of advanced AI and machine learning models, which can have unforeseen impacts on operations, ethics, and strategic planning. The piece highlights that such a simplified view can impede effective AI governance, robust risk management, and necessary organizational adjustments, underscoring the critical need for enhanced AI literacy among leadership to navigate the transformative challenges and opportunities presented by rapidly evolving AI technologies.

GitHub

3 stories
01

Spring AI Alibaba

Spring AI Alibaba is an advanced agentic AI framework designed for building sophisticated ChatBot, Workflow, and Multi-agent applications. It features a graph-based multi-agent framework, allowing developers to construct complex workflows and agents, with visual debugging and Dify DSL generation capabilities. The framework is engineered for enterprise environments, offering deep integration with the Alibaba Cloud AI ecosystem, including the Aliyun Bailian platform for LLM services and RAG solutions, as well as AI observation tools like ARMS and Langfuse. It also supports enterprise-level MCP integration through Nacos MCP Registry for service discovery and routing, and leverages Higress for LLM model proxying. Spring AI Alibaba introduces specialized Plan-Act agent products like JManus for delicate plan adjustment and reuse, and DeepResearch, an agent for comprehensive research and report generation utilizing web search, crawling, and Python scripting. The platform aims to facilitate the transition of AI agents from experimental demos to production-ready solutions, emphasizing deterministic and domain-specific agent development.

02

Happy-LLM

The "Happy-LLM" project by Datawhale offers a comprehensive, free, and open-source tutorial designed to guide learners through the principles and practical implementation of Large Language Models (LLMs). This systematic curriculum delves into fundamental Natural Language Processing (NLP) methods, the architectural foundations of LLMs like the Transformer, and the intricate details of the training process, from pre-training to fine-tuning. Participants will gain a deep understanding of core concepts such as attention mechanisms and existing LLM structures, and acquire hands-on experience by implementing a complete LLaMA2 model. The tutorial also covers advanced application techniques like Retrieval-Augmented Generation (RAG) and AI Agents, equipping learners with the skills to navigate the rapidly evolving LLM landscape. Aimed at students, researchers, and enthusiasts with some programming and deep learning background, "Happy-LLM" fosters practical engagement and encourages contributions to the broader open-source AI community.

03

Welcome to Anthropic's Prompt Engineering Interactive Tutorial

This interactive tutorial by Anthropic provides a comprehensive, step-by-step guide to engineering optimal prompts for Claude AI models, including Haiku, Sonnet, and Opus. It aims to equip users with the ability to master prompt structures, identify common failure modes, apply '80/20' techniques for improvement, understand Claude's capabilities, and build effective prompts for various applications. Structured into nine chapters with practical exercises and an advanced appendix, the course progresses from fundamental concepts like clear instructions and role assignment to intermediate topics such as data separation and output formatting. Advanced sections cover hallucination avoidance and constructing complex prompts for industry-specific use cases in areas like chatbots, legal, financial, and coding services. The tutorial emphasizes hands-on experimentation through an 'Example Playground' and is also available as a user-friendly Google Sheets version.

huggingface

7 stories
01

QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs

We propose QeRL, a Quantization-enhanced Reinforcement Learning framework for large language models (LLMs). While RL is essential for LLMs' reasoning capabilities, it is resource-intensive, requiring substantial GPU memory and long rollout durations. QeRL addresses these issues by combining NVFP4 quantization with Low-Rank Adaptation (LoRA), accelerating rollout phase of RL while reducing memory overhead. Beyond efficiency, our findings show that quantization noise increases policy entropy, enhancing exploration, and enabling the discovery of better strategies during RL. To further optimize exploration, QeRL introduces an Adaptive Quantization Noise (AQN) mechanism, which dynamically adjusts noise during training. Experiments demonstrate that QeRL delivers over 1.5 times speedup in the rollout phase. Moreover, this is the first framework to enable RL training of a 32B LLM on a single H100 80GB GPU, while delivering overall speedups for RL training. It also achieves faster reward growth and higher final accuracy than 16-bit LoRA and QLoRA, while matching the performance of full-parameter fine-tuning on mathematical benchmarks such as GSM8K (90.8%) and MATH 500 (77.4%) in the 7B model. These results establish QeRL as an efficient and effective framework for RL training in LLMs.

02

Stable Video Infinity: Infinite-Length Video Generation with Error Recycling

We propose Stable Video Infinity (SVI) that is able to generate infinite-length videos with high temporal consistency, plausible scene transitions, and controllable streaming storylines. While existing long-video methods attempt to mitigate accumulated errors via handcrafted anti-drifting (e.g., modified noise scheduler, frame anchoring), they remain limited to single-prompt extrapolation, producing homogeneous scenes with repetitive motions. We identify that the fundamental challenge extends beyond error accumulation to a critical discrepancy between the training assumption (seeing clean data) and the test-time autoregressive reality (conditioning on self-generated, error-prone outputs). To bridge this hypothesis gap, SVI incorporates Error-Recycling Fine-Tuning, a new type of efficient training that recycles the Diffusion Transformer (DiT)'s self-generated errors into supervisory prompts, thereby encouraging DiT to actively identify and correct its own errors. This is achieved by injecting, collecting, and banking errors through closed-loop recycling, autoregressively learning from error-injected feedback. Specifically, we (i) inject historical errors made by DiT to intervene on clean inputs, simulating error-accumulated trajectories in flow matching; (ii) efficiently approximate predictions with one-step bidirectional integration and calculate errors with residuals; (iii) dynamically bank errors into replay memory across discretized timesteps, which are resampled for new input. SVI is able to scale videos from seconds to infinite durations with no additional inference cost, while remaining compatible with diverse conditions (e.g., audio, skeleton, and text streams). We evaluate SVI on three benchmarks, including consistent, creative, and conditional settings, thoroughly verifying its versatility and state-of-the-art role.

03

Demystifying Reinforcement Learning in Agentic Reasoning

Recently, the emergence of agentic RL has showcased that RL could also effectively improve the agentic reasoning ability of LLMs, yet the key design principles and optimal practices remain unclear. In this work, we conduct a comprehensive and systematic investigation to demystify reinforcement learning in agentic reasoning from three key perspectives: data, algorithm, and reasoning mode. We highlight our key insights: (i) Replacing stitched synthetic trajectories with real end-to-end tool-use trajectories yields a far stronger SFT initialization; high-diversity, model-aware datasets sustain exploration and markedly improve RL performance. (ii) Exploration-friendly techniques are crucial for agentic RL, such as clip higher, overlong reward shaping, and maintaining adequate policy entropy could improve the training efficiency. (iii) A deliberative strategy with fewer tool calls outperforms frequent tool calls or verbose self-reasoning, improving tool efficiency and final accuracy. Together, these simple practices consistently enhance agentic reasoning and training efficiency, achieving strong results on challenging benchmarks with smaller models, and establishing a practical baseline for future agentic RL research. Beyond these empirical insights, we further contribute a high-quality, real end-to-end agentic SFT dataset along with a high-quality RL dataset, and demonstrate the effectiveness of our insights in boosting the agentic reasoning ability of LLMs across four challenging benchmarks, including AIME2024/AIME2025, GPQA-Diamond, and LiveCodeBench-v6. With our recipes, 4B-sized models could also achieve superior agentic reasoning performance compared to 32B-sized models. Code and models: https://github.com/Gen-Verse/Open-AgentRL

04

Diffusion Transformers with Representation Autoencoders

Latent generative modeling, where a pretrained autoencoder maps pixels into a latent space for the diffusion process, has become the standard strategy for Diffusion Transformers (DiT); however, the autoencoder component has barely evolved. Most DiTs continue to rely on the original VAE encoder, which introduces several limitations: outdated backbones that compromise architectural simplicity, low-dimensional latent spaces that restrict information capacity, and weak representations that result from purely reconstruction-based training and ultimately limit generative quality. In this work, we explore replacing the VAE with pretrained representation encoders (e.g., DINO, SigLIP, MAE) paired with trained decoders, forming what we term Representation Autoencoders (RAEs). These models provide both high-quality reconstructions and semantically rich latent spaces, while allowing for a scalable transformer-based architecture. Since these latent spaces are typically high-dimensional, a key challenge is enabling diffusion transformers to operate effectively within them. We analyze the sources of this difficulty, propose theoretically motivated solutions, and validate them empirically. Our approach achieves faster convergence without auxiliary representation alignment losses. Using a DiT variant equipped with a lightweight, wide DDT head, we achieve strong image generation results on ImageNet: 1.51 FID at 256x256 (no guidance) and 1.13 at both 256x256 and 512x512 (with guidance). RAE offers clear advantages and should be the new default for diffusion transformer training.

05

Don't Just Fine-tune the Agent, Tune the Environment

Large Language Model (LLM) agents show great promise for complex, multi-turn tool-use tasks, but their development is often hampered by the extreme scarcity of high-quality training data. Supervised fine-tuning (SFT) on synthetic data leads to overfitting, whereas standard reinforcement learning (RL) struggles with a critical cold-start problem and training instability. To address these challenges, we introduce Environment Tuning, a novel training paradigm that enables agents to learn complex behaviors directly from problem instances without relying on pre-collected expert trajectories. Environment Tuning orchestrates this learning process through a structured curriculum, actionable environment augmentation that provides corrective feedback, and fine-grained progress rewards to ensure stable and efficient exploration. Using only 400 problem instances from Berkeley Function-Calling Leaderboard (BFCL) benchmark, our method not only achieves competitive in-distribution performance against strong baselines but also demonstrates superior out-of-distribution generalization, overcoming the performance collapse common to SFT-based approaches. Our work presents a paradigm shift from supervised fine-tuning on static trajectories to dynamic, environment-based exploration, paving the way for training more robust and data-efficient agents.

06

InfiniHuman: Infinite 3D Human Creation with Precise Control

Generating realistic and controllable 3D human avatars is a long-standing challenge, particularly when covering broad attribute ranges such as ethnicity, age, clothing styles, and detailed body shapes. Capturing and annotating large-scale human datasets for training generative models is prohibitively expensive and limited in scale and diversity. The central question we address in this paper is: Can existing foundation models be distilled to generate theoretically unbounded, richly annotated 3D human data? We introduce InfiniHuman, a framework that synergistically distills these models to produce richly annotated human data at minimal cost and with theoretically unlimited scalability. We propose InfiniHumanData, a fully automatic pipeline that leverages vision-language and image generation models to create a large-scale multi-modal dataset. User study shows our automatically generated identities are undistinguishable from scan renderings. InfiniHumanData contains 111K identities spanning unprecedented diversity. Each identity is annotated with multi-granularity text descriptions, multi-view RGB images, detailed clothing images, and SMPL body-shape parameters. Building on this dataset, we propose InfiniHumanGen, a diffusion-based generative pipeline conditioned on text, body shape, and clothing assets. InfiniHumanGen enables fast, realistic, and precisely controllable avatar generation. Extensive experiments demonstrate significant improvements over state-of-the-art methods in visual quality, generation speed, and controllability. Our approach enables high-quality avatar generation with fine-grained control at effectively unbounded scale through a practical and affordable solution. We will publicly release the automatic data generation pipeline, the comprehensive InfiniHumanData dataset, and the InfiniHumanGen models at https://yuxuan-xue.com/infini-human.

07

ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding

While Large Language Models (LLMs) excel at algorithmic code generation, they struggle with front-end development, where correctness is judged on rendered pixels and interaction. We present ReLook, an agentic, vision-grounded reinforcement learning framework that empowers an agent to close a robust generate--diagnose--refine loop by invoking a multimodal LLM (MLLM) as a tool. During training, the agent uses the MLLM-in-the-loop both as a visual critic--scoring code with screenshots--and as a source of actionable, vision-grounded feedback; a strict zero-reward rule for invalid renders anchors renderability and prevents reward hacking. To prevent behavioral collapse, we introduce Forced Optimization, a strict acceptance rule that admits only improving revisions, yielding monotonically better trajectories. At inference, we decouple the critic and run a lightweight, critic-free self-edit cycle, keeping latency comparable to base decoding while retaining most of the gains. Across three widely used benchmarks, ReLook consistently outperforms strong baselines in vision-grounded front-end code generation, highlighting the benefits of agentic perception, visual rewards, and training-inference decoupling.