NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-03-04ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Agentic Engineering Patterns

The article "Agentic Engineering Patterns" explores foundational methodologies and best practices for the systematic development of sophisticated artificial intelligence agents. This emerging discipline emphasizes designing AI systems that can autonomously execute complex tasks, often integrating large language models (LLMs) to power their reasoning, planning, and interaction capabilities. The guide would explore various architectural and operational patterns essential for building reliable, scalable, and adaptable agentic systems. Key areas likely include effective prompt engineering techniques, state management, tool integration for enhanced functionality, and mechanisms for memory and self-correction. It aims to provide a structured framework for addressing challenges such as task decomposition, error handling, and ensuring robust behavior in diverse operational contexts. By outlining repeatable patterns, the article helps engineers transition from basic AI applications to more advanced, goal-driven agents capable of complex decision-making and dynamic adaptation, thereby professionalizing the development of autonomous AI entities.

02

Qwen3.5 Fine-Tuning Guide – Unsloth Documentation

The Unsloth documentation presents a dedicated guide for the fine-tuning of the Qwen3.5 large language model, addressing the critical need for efficient model customization in the rapidly evolving AI landscape. Unsloth positions itself as a crucial tool for developers and researchers, offering significant optimizations in computational speed and memory usage during the training and fine-tuning of large neural networks. This guide specifically enables users to leverage these efficiencies when working with Qwen3.5, allowing for more rapid iteration and deployment of specialized LLM applications. It provides step-by-step instructions, recommended configurations, and best practices to ensure that Qwen3.5 can be effectively adapted to diverse downstream tasks and domain-specific datasets. By simplifying the often resource-intensive process of fine-tuning, Unsloth democratizes access to advanced LLM capabilities, fostering innovation and accelerating the integration of powerful AI models into various real-world scenarios, ultimately driving progress in Natural Language Processing and broader AI development.

03

NanoGPT Slowrun: Language Modeling with Limited Data, Infinite Compute

The "NanoGPT Slowrun" project investigates novel approaches to language modeling within a specialized paradigm characterized by limited data but abundant computational resources. This initiative explores the potential for training compact, GPT-inspired models to achieve robust language generation and understanding when traditional large datasets are unavailable. By assuming "infinite compute," the research likely delves into exhaustive hyperparameter optimization, advanced architectural explorations, or innovative data augmentation and synthetic data generation techniques designed to maximize learning from minimal real-world examples. The core objective is to determine how far language models can be pushed in data-scarce environments, offering critical insights for developing more adaptable and efficient AI systems for specialized domains. This "slowrun" methodology suggests a deep, methodical exploration of training dynamics under these unique conditions, aiming to establish new benchmarks for data-efficient NLP.

04

Moss is a pixel canvas where every brush is a tiny program

Moss introduces a novel digital art platform centered around a pixel canvas where the functionality of every brush is defined by a "tiny program." This innovative concept transcends conventional static brush tools by enabling users to programmatically dictate the behavior, appearance, and interactions of their drawing instruments. By integrating custom code snippets, artists and developers gain unparalleled control over the creative process, facilitating the generation of dynamic patterns, intricate textures, and interactive visual effects. This approach fosters a unique synergy between artistic expression and software development, transforming traditional digital painting into a form of visual programming. The platform is designed to empower creators to explore the vast potential of procedural generation and algorithmic art, offering a robust environment for experimentation and the development of highly personalized artistic tools. This could revolutionize digital art workflows by providing a flexible, code-driven methodology for crafting complex and evolving visual content.

05

New York could prohibit chatbot medical, legal, engineering advice

New York State is exploring legislative measures to curtail the capabilities of artificial intelligence chatbots, specifically targeting their provision of advice in highly regulated professional domains such as medicine, law, and engineering. The proposed legislation, identified as Senate Bill S7263, stems from growing concerns over the potential for inaccurate, unverified, or harmful information disseminated by AI systems. Lawmakers are grappling with complex questions of accountability and liability, as current legal frameworks are ill-equipped to address the implications of AI-generated professional guidance. This initiative underscores a broader movement towards establishing clear regulatory boundaries for AI applications, particularly where public health, safety, and legal integrity are at stake. The bill seeks to safeguard consumers from potential misinformation and ensures that professional advice remains subject to human oversight and accountability. This legislative push by New York could significantly influence future AI governance discussions and regulatory approaches across other jurisdictions, highlighting a proactive stance on managing the societal impact of advanced AI technologies.

06

1.5 Million Users Leave ChatGPT

Reports indicate that a substantial number of users, approximately 1.5 million, have recently disengaged from OpenAI's leading AI chatbot, ChatGPT. This significant user churn signals potential shifts in the competitive landscape of large language models and prompts critical examination of user retention strategies within the rapidly evolving AI sector. While specific detailed reasons for this exodus are not provided, such trends are often influenced by factors including the emergence of advanced or more specialized alternative AI services, changes in user requirements, concerns over data privacy, or the perceived value proposition of subscription tiers. The associated article title, 'Leaving ChatGPT: Make Sure To Do This Before You Cancel,' suggests a broader trend of users actively discontinuing their engagement, necessitating advice on managing such transitions. This development underscores the dynamic nature of user loyalty in the AI space, emphasizing the continuous need for innovation and adaptation to maintain a dominant market position.

huggingface

6 stories
01

Learning When to Act or Refuse: Guarding Agentic Reasoning Models for Safe Multi-Step Tool Use

Agentic language models operate in a fundamentally different safety regime than chat models: they must plan, call tools, and execute long-horizon actions where a single misstep, such as accessing files or entering credentials, can cause irreversible harm. Existing alignment methods, largely optimized for static generation and task completion, break down in these settings due to sequential decision-making, adversarial tool feedback, and overconfident intermediate reasoning. We introduce MOSAIC, a post-training framework that aligns agents for safe multi-step tool use by making safety decisions explicit and learnable. MOSAIC structures inference as a plan, check, then act or refuse loop, with explicit safety reasoning and refusal as first-class actions. To train without trajectory-level labels, we use preference-based reinforcement learning with pairwise trajectory comparisons, which captures safety distinctions often missed by scalar rewards. We evaluate MOSAIC zero-shot across three model families, Qwen2.5-7B, Qwen3-4B-Thinking, and Phi-4, and across out-of-distribution benchmarks spanning harmful tasks, prompt injection, benign tool use, and cross-domain privacy leakage. MOSAIC reduces harmful behavior by up to 50%, increases harmful-task refusal by over 20% on injection attacks, cuts privacy leakage, and preserves or improves benign task performance, demonstrating robust generalization across models, domains, and agentic settings.

02

AgentConductor: Topology Evolution for Multi-Agent Competition-Level Code Generation

Large language model(LLM)-driven multi-agent systems(MAS) coordinate specialized agents through predefined interaction topologies and have shown promise for complex tasks such as competition-level code generation. Recent studies demonstrate that carefully designed multi-agent workflows and communication graphs can significantly improve code generation performance by leveraging collaborative reasoning. However, existing methods neither adapt topology density to task difficulty nor iteratively refine the topology within an instance using execution feedback, which leads to redundant communication and performance bottlenecks. To address these issues, we propose AgentConductor: a reinforcement learning-optimized MAS with an LLM-based orchestrator agent as its core, which enables end-to-end feedback-driven dynamic generation of interaction topologies. For each query, AgentConductor infers agent roles and task difficulty, then constructs a task-adapted, density-aware layered directed acyclic graph (DAG) topology, underpinned by two key innovations. First, we design a novel topological density function that captures communication-aware mathematical characterizations of multi-agent interactions. Second, we adopt difficulty interval partitioning to avoid excessive pruning for precise topological density upper bound measurement per difficulty level and finer-grained control. Empirically, across three competition-level and two foundational code datasets, AgentConductor achieves state-of-the-art accuracy, outperforming the strongest baseline by up to 14.6% in pass@1 accuracy, 13% in density reduction, and 68% in token cost reduction.

03

Beyond Language Modeling: An Exploration of Multimodal Pretraining

The visual world offers a critical axis for advancing foundation models beyond language. Despite growing interest in this direction, the design space for native multimodal models remains opaque. We provide empirical clarity through controlled, from-scratch pretraining experiments, isolating the factors that govern multimodal pretraining without interference from language pretraining. We adopt the Transfusion framework, using next-token prediction for language and diffusion for vision, to train on diverse data including text, video, image-text pairs, and even action-conditioned video. Our experiments yield four key insights: (i) Representation Autoencoder (RAE) provides an optimal unified visual representation by excelling at both visual understanding and generation; (ii) visual and language data are complementary and yield synergy for downstream capabilities; (iii) unified multimodal pretraining leads naturally to world modeling, with capabilities emerging from general training; and (iv) Mixture-of-Experts (MoE) enables efficient and effective multimodal scaling while naturally inducing modality specialization. Through IsoFLOP analysis, we compute scaling laws for both modalities and uncover a scaling asymmetry: vision is significantly more data-hungry than language. We demonstrate that the MoE architecture harmonizes this scaling asymmetry by providing the high model capacity required by language while accommodating the data-intensive nature of vision, paving the way for truly unified multimodal models.

04

Kling-MotionControl Technical Report

Character animation aims to generate lifelike videos by transferring motion dynamics from a driving video to a reference image. Recent strides in generative models have paved the way for high-fidelity character animation. In this work, we present Kling-MotionControl, a unified DiT-based framework engineered specifically for robust, precise, and expressive holistic character animation. Leveraging a divide-and-conquer strategy within a cohesive system, the model orchestrates heterogeneous motion representations tailored to the distinct characteristics of body, face, and hands, effectively reconciling large-scale structural stability with fine-grained articulatory expressiveness. To ensure robust cross-identity generalization, we incorporate adaptive identity-agnostic learning, facilitating natural motion retargeting for diverse characters ranging from realistic humans to stylized cartoons. Simultaneously, we guarantee faithful appearance preservation through meticulous identity injection and fusion designs, further supported by a subject library mechanism that leverages comprehensive reference contexts. To ensure practical utility, we implement an advanced acceleration framework utilizing multi-stage distillation, boosting inference speed by over 10x. Kling-MotionControl distinguishes itself through intelligent semantic motion understanding and precise text responsiveness, allowing for flexible control beyond visual inputs. Human preference evaluations demonstrate that Kling-MotionControl delivers superior performance compared to leading commercial and open-source solutions, achieving exceptional fidelity in holistic motion control, open domain generalization, and visual quality and coherence. These results establish Kling-MotionControl as a robust solution for high-quality, controllable, and lifelike character animation.

05

DREAM: Where Visual Understanding Meets Text-to-Image Generation

Unifying visual representation learning and text-to-image (T2I) generation within a single model remains a central challenge in multimodal learning. We introduce DREAM, a unified framework that jointly optimizes discriminative and generative objectives, while learning strong visual representations. DREAM is built on two key techniques: During training, Masking Warmup, a progressive masking schedule, begins with minimal masking to establish the contrastive alignment necessary for representation learning, then gradually transitions to full masking for stable generative training. At inference, DREAM employs Semantically Aligned Decoding to align partially masked image candidates with the target text and select the best one for further decoding, improving text-image fidelity (+6.3%) without external rerankers. Trained solely on CC12M, DREAM achieves 72.7% ImageNet linear-probing accuracy (+1.1% over CLIP) and an FID of 4.25 (+6.2% over FLUID), with consistent gains in few-shot classification, semantic segmentation, and depth estimation. These results demonstrate that discriminative and generative objectives can be synergistic, allowing unified multimodal models that excel at both visual understanding and generation.

06

Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models

Recent advancements in Generative Reward Models (GRMs) have demonstrated that scaling the length of Chain-of-Thought (CoT) reasoning considerably enhances the reliability of evaluation. However, current works predominantly rely on unstructured length scaling, ignoring the divergent efficacy of different reasoning mechanisms: Breadth-CoT (B-CoT, i.e., multi-dimensional principle coverage) and Depth-CoT (D-CoT, i.e., substantive judgment soundness). To address this, we introduce Mix-GRM, a framework that reconfigures raw rationales into structured B-CoT and D-CoT through a modular synthesis pipeline, subsequently employing Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR) to internalize and optimize these mechanisms. Comprehensive experiments demonstrate that Mix-GRM establishes a new state-of-the-art across five benchmarks, surpassing leading open-source RMs by an average of 8.2%. Our results reveal a clear divergence in reasoning: B-CoT benefits subjective preference tasks, whereas D-CoT excels in objective correctness tasks. Consequently, misaligning the reasoning mechanism with the task directly degrades performance. Furthermore, we demonstrate that RLVR acts as a switching amplifier, inducing an emergent polarization where the model spontaneously allocates its reasoning style to match task demands. The synthesized data and models are released at https://huggingface.co/collections/DonJoey/mix-grm{Hugging Face}, and the code is released at https://github.com/Don-Joey/Mix-GRM{Github}.