NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-02-25ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Hacker used Anthropic's Claude chatbot to attack government agencies in Mexico

A recent report has brought to light an incident where a malicious actor exploited Anthropic's Claude chatbot to execute attacks targeting several government agencies within Mexico. This development underscores a critical and escalating concern regarding the potential misuse of advanced artificial intelligence tools for nefarious purposes. While the precise methodology of the chatbot's involvement in these attacks, such as its application in social engineering, data collection, or the generation of deceptive content, remains to be fully elucidated, the deployment of a sophisticated Large Language Model (LLM) like Claude by an attacker points to evolving cybersecurity threats. This event necessitates a rigorous review of the security protocols and ethical frameworks governing the development and deployment of generative AI platforms. It emphasizes the urgent need for AI developers to strengthen preventative measures against platform abuse, particularly when such technologies can be weaponized against critical governmental infrastructure. The incident serves as a potent reminder for the AI community and cybersecurity experts to forge closer collaborations in fortifying digital defenses against the complex and rapidly emerging landscape of AI-powered cyber threats.

02

The Pentagon Threatens Anthropic

Recent reports indicate a developing tension between the U.S. Department of Defense (The Pentagon) and leading artificial intelligence firm Anthropic, signaling potential governmental concerns over the burgeoning field of advanced AI. While specific details of the 'threat' remain undisclosed, the situation likely pertains to national security implications, the responsible development and deployment of dual-use AI technologies, or regulatory compliance. Anthropic, known for its focus on AI safety and 'Constitutional AI' principles, operates at the forefront of large language model research, making its work of significant interest to defense sectors. This development underscores the increasing scrutiny from governmental bodies on AI companies, highlighting the critical juncture where technological innovation intersects with national security imperatives and ethical considerations. The interaction suggests an evolving landscape where AI developers may face heightened pressure to align their research and product roadmaps with strategic national interests and emerging regulatory frameworks.

03

The Appeal and Reality of Recycling LoRAs with Adaptive Merging

The research paper, "The Appeal and Reality of Recycling LoRAs with Adaptive Merging," delves into the practical aspects and underlying mechanisms of reusing Low-Rank Adaptation (LoRA) modules through adaptive merging techniques. LoRAs have emerged as a cornerstone in parameter-efficient fine-tuning, significantly reducing the computational cost and storage requirements associated with adapting large pre-trained models. This study explores the potential benefits of "recycling" these fine-tuned LoRA weights, specifically focusing on methods that adaptively combine multiple LoRA modules to synthesize new models or improve existing ones. The work likely investigates scenarios where knowledge transfer or model compression can be achieved by intelligently merging LoRAs, contrasting the theoretical appeal of such an approach with the practical challenges and performance trade-offs encountered in real-world applications. It aims to provide insights into the effectiveness, limitations, and optimal strategies for integrating diverse LoRA adaptations, offering a clearer understanding of when and how adaptive merging can truly enhance the development and deployment of customized large language models or other deep learning architectures.

04

Claude Code Remote Control

The introduction of "Claude Code Remote Control" signifies a notable advancement in the capabilities of the Claude Code development environment, hinting at the availability of a programmatic interface for its AI-powered coding functionalities. While comprehensive details are hosted at code.claude.com/docs/en/remote-control, this feature is expected to empower developers with an API or SDK, facilitating the remote control, automation, and seamless integration of Claude's sophisticated code generation and assistance tools into diverse software development workflows. This capability is poised to markedly boost developer productivity by enabling headless operations, automated code reviews, continuous integration pipelines, and the creation of bespoke tools that leverage Claude's advanced AI. It represents a strategic evolution towards making AI coding assistants more deeply programmable and embeddable within complex engineering ecosystems, moving beyond interactive UIs to support more automated and scalable development practices. This enhancement solidifies Claude Code's position as a versatile and powerful platform for AI-driven software engineering initiatives.

05

Show HN: Sgai – Goal-driven multi-agent software dev (GOAL.md → working code)

Sgai introduces a novel approach to AI-assisted software development, shifting from step-by-step prompting to a goal-driven methodology. Users define the desired outcome in a GOAL.md file, which Sgai then executes using a coordinated set of AI agents. The system intelligently decomposes the high-level goal into a Directed Acyclic Graph (DAG) of specialized roles, such as developer, reviewer, and safety analyst. It proactively engages by asking clarifying questions, iteratively writes and refines code, and validates its work through automated testing, utilizing completion gates like `make test` to ensure functional completion. Operating entirely locally within the user's repository, Sgai offers a real-time web dashboard to visualize agent execution without automatically pushing changes to external platforms like GitHub. While still in its early stages of development, Sgai has demonstrated its utility for rapid prototyping of small applications and internal tools, marking a step towards more autonomous software creation.

06

Show HN: A real-time strategy game that AI agents can play

The "LLM Skirmish" project introduces a real-time strategy (RTS) game environment specifically tailored to highlight the coding capabilities of advanced large language models (LLMs). This initiative stems from an observation that while frontier LLMs can autonomously complete complex coding tasks, they frequently struggle with navigation and decision-making in simpler game settings, such as Pokémon Red. Drawing inspiration from Screeps, an "MMO RTS sandbox for programmers" released a decade ago, LLM Skirmish adapts its open-source API to create a platform where LLM agents compete head-to-head in 1v1 RTS matches. The core mechanic involves LLMs generating and executing code in real-time to manage their in-game factions. Early evaluations revealed Claude Opus 4.5 as a particularly strong contender, though it demonstrated a tendency towards economic over-focus in initial rounds. This innovative game environment serves as a valuable tool for exploring and benchmarking the strategic and programming abilities of AI agents within dynamic, competitive scenarios.

huggingface

6 stories
01

On Data Engineering for Scaling LLM Terminal Capabilities

Despite rapid recent progress in the terminal capabilities of large language models, the training data strategies behind state-of-the-art terminal agents remain largely undisclosed. We address this gap through a systematic study of data engineering practices for terminal agents, making two key contributions: (1) Terminal-Task-Gen, a lightweight synthetic task generation pipeline that supports seed-based and skill-based task construction, and (2) a comprehensive analysis of data and training strategies, including filtering, curriculum learning, long context training, and scaling behavior. Our pipeline yields Terminal-Corpus, a large-scale open-source dataset for terminal tasks. Using this dataset, we train Nemotron-Terminal, a family of models initialized from Qwen3(8B, 14B, 32B) that achieve substantial gains on Terminal-Bench 2.0: Nemotron-Terminal-8B improves from 2.5% to 13.0% Nemotron-Terminal-14B improves from 4.0% to 20.2%, and Nemotron-Terminal-32B improves from 3.4% to 27.4%, matching the performance of significantly larger models. To accelerate research in this domain, we open-source our model checkpoints and most of our synthetic datasets at https://huggingface.co/collections/nvidia/nemotron-terminal.

02

PyVision-RL: Forging Open Agentic Vision Models via RL

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. Our approach combines an oversampling-filtering-ranking rollout strategy with an accumulative tool reward to prevent collapse and encourage multi-turn tool use. Using a unified training pipeline, we develop PyVision-Image and PyVision-Video for image and video understanding. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames during reasoning to significantly reduce visual token usage. Experiments show strong performance and improved efficiency, demonstrating that sustained interaction and on-demand visual processing are critical for scalable multimodal agents.

03

See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis

Despite recent advances in diffusion models, AI generated images still often contain visual artifacts that compromise realism. Although more thorough pre-training and bigger models might reduce artifacts, there is no assurance that they can be completely eliminated, which makes artifact mitigation a highly crucial area of study. Previous artifact-aware methodologies depend on human-labeled artifact datasets, which are costly and difficult to scale, underscoring the need for an automated approach to reliably acquire artifact-annotated datasets. In this paper, we propose ArtiAgent, which efficiently creates pairs of real and artifact-injected images. It comprises three agents: a perception agent that recognizes and grounds entities and subentities from real images, a synthesis agent that introduces artifacts via artifact injection tools through novel patch-wise embedding manipulation within a diffusion transformer, and a curation agent that filters the synthesized artifacts and generates both local and global explanations for each instance. Using ArtiAgent, we synthesize 100K images with rich artifact annotations and demonstrate both efficacy and versatility across diverse applications. Code is available at link.

04

Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs

Embodied LLMs endow robots with high-level task reasoning, but they cannot reflect on what went wrong or why, turning deployment into a sequence of independent trials where mistakes repeat rather than accumulate into experience. Drawing upon human reflective practitioners, we introduce Reflective Test-Time Planning, which integrates two modes of reflection: reflection-in-action, where the agent uses test-time scaling to generate and score multiple candidate actions using internal reflections before execution; and reflection-on-action, which uses test-time training to update both its internal reflection model and its action policy based on external reflections after execution. We also include retrospective reflection, allowing the agent to re-evaluate earlier decisions and perform model updates with hindsight for proper long-horizon credit assignment. Experiments on our newly-designed Long-Horizon Household benchmark and MuJoCo Cupboard Fitting benchmark show significant gains over baseline models, with ablative studies validating the complementary roles of reflection-in-action and reflection-on-action. Qualitative analyses, including real-robot trials, highlight behavioral correction through reflection.

05

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking

Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism. The dominant approaches in this family of methods, such as Ring Attention or DeepSpeed Ulysses, enable scaling over the context dimension but do not focus on memory efficiency, which limits the sequence lengths they can support. More advanced techniques, such as Fully Pipelined Distributed Transformer or activation offloading, can further extend the possible context length at the cost of training throughput. In this paper, we present UPipe, a simple yet effective context parallelism technique that performs fine-grained chunking at the attention head level. This technique significantly reduces the activation memory usage of self-attention, breaking the activation memory barrier and unlocking much longer context lengths. Our approach reduces intermediate tensor memory usage in the attention layer by as much as 87.5% for 32B Transformers, while matching previous context parallelism techniques in terms of training speed. UPipe can support the context length of 5M tokens when training Llama3-8B on a single 8timesH100 node, improving upon prior methods by over 25%.

06

One-step Language Modeling via Continuous Denoising

Language models based on discrete diffusion have attracted widespread interest for their potential to provide faster generation than autoregressive models. In practice, however, they exhibit a sharp degradation of sample quality in the few-step regime, failing to realize this promise. Here we show that language models leveraging flow-based continuous denoising can outperform discrete diffusion in both quality and speed. By revisiting the fundamentals of flows over discrete modalities, we build a flow-based language model (FLM) that performs Euclidean denoising over one-hot token encodings. We show that the model can be trained by predicting the clean data via a cross entropy objective, where we introduce a simple time reparameterization that greatly improves training stability and generation quality. By distilling FLM into its associated flow map, we obtain a distilled flow map language model (FMLM) capable of few-step generation. On the LM1B and OWT language datasets, FLM attains generation quality matching state-of-the-art discrete diffusion models. With FMLM, our approach outperforms recent few-step language models across the board, with one-step generation exceeding their 8-step quality. Our work calls into question the widely held hypothesis that discrete diffusion processes are necessary for generative modeling over discrete modalities, and paves the way toward accelerated flow-based language modeling at scale. Code is available at https://github.com/david3684/flm.