NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-02-03DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Qwen3-Coder-Next

Qwen3-Coder-Next is introduced as the latest iteration in the Qwen series of large language models, specifically engineered for advanced code generation and software development tasks. Developed by Qwen.ai, this model is anticipated to build upon the capabilities of its predecessors, offering enhanced performance in understanding complex programming languages, debugging, and generating high-quality, efficient code. The 'Next' in its name suggests significant improvements in architectural design or training methodologies, potentially leading to better contextual understanding, fewer hallucinations in code output, and broader language support. This release aims to further empower developers by streamlining coding workflows, automating repetitive tasks, and assisting in the creation of more robust and reliable software solutions across various platforms. Its introduction signifies a continued push towards more sophisticated AI assistants in the software engineering domain, promising to integrate seamlessly into existing development environments and boost productivity, thereby accelerating innovation in software development processes.

02

Agent Skills

Agent Skills refers to the specialized competencies and functionalities that empower artificial intelligence agents to perform complex tasks and interact effectively within dynamic environments. This concept encompasses the design, acquisition, and management of distinct capabilities, allowing AI agents to learn, adapt, and execute actions autonomously. These skills often involve integrating various AI techniques, such as natural language understanding, problem-solving, decision-making, and interaction protocols. The development of robust agent skills is crucial for advancing AI's practical applications, enabling agents to achieve higher levels of autonomy, efficiency, and intelligence in diverse domains, from task automation and virtual assistance to complex scientific research and human-computer collaboration. Platforms focusing on agent skills typically aim to provide frameworks or tools for developers to build, train, and deploy agents with a versatile set of abilities, thereby enhancing their overall performance and utility in real-world scenarios.

03

Xcode 26.3 unlocks the power of agentic coding

Apple has announced the release of Xcode 26.3, introducing groundbreaking capabilities that fundamentally transform the software development process through 'agentic coding.' This update integrates advanced artificial intelligence agents directly into the integrated development environment, empowering developers with intelligent automation and enhanced productivity features. Agentic coding in Xcode 26.3 is designed to assist with various development tasks, including intelligent code completion, automated debugging, proactive error detection, and streamlined refactoring suggestions, aiming to reduce repetitive work and accelerate development cycles. The new features leverage sophisticated AI models to understand code context, anticipate developer needs, and provide highly relevant, context-aware assistance. This marks a significant step towards a more autonomous and efficient coding paradigm, allowing developers to focus more on architectural design and complex problem-solving while AI agents handle routine tasks. The integration promises to improve code quality, enforce best practices, and significantly enhance the overall developer experience within the Apple ecosystem.

04

Show HN: I built "AI Wattpad" to eval LLMs on fiction

A new platform, Narrator, has been developed to evaluate Large Language Models (LLMs) on their ability to generate engaging serialized fiction. The creator, a long-time webfiction reader, identified a gap in current LLM evaluation methods, which often fail to assess the holistic nature of creative writing. Unlike fragmented benchmarks that test isolated capabilities like brainstorming, writing, or memory, Narrator focuses on real reader engagement to rank LLMs. This approach acknowledges that creative writing is a complex pipeline requiring consistent narrative, good prose, and strong memory across extended content. The platform aims to provide a more comprehensive and reader-centric evaluation of LLMs' performance in generating compelling long-form fictional content, addressing the limitations of existing, more academic evaluation landscapes.

05

How does misalignment scale with model intelligence and task complexity?

A recent inquiry delves into the intricate relationship between AI misalignment, model intelligence, and task complexity, addressing a fundamental concern within AI safety research. The central question revolves around how the challenge of ensuring AI systems act in accordance with human intent scales as these models become more capable and are deployed in increasingly complex environments. This research area is critical for understanding the potential for unintended consequences and control issues as AI advances. It explores whether misalignment issues intensify, transform, or become more tractable with greater intelligence. The investigation likely considers different facets of misalignment, such as goal misspecification, emergent behaviors, and value drift, examining how these manifest across varying levels of AI sophistication. The findings could offer crucial insights for developing more robust alignment strategies and safe AI architectures, ultimately influencing the trajectory of advanced artificial intelligence development.

06

LNAI \b Define AI coding tool configs once, sync to Claude, Cursor, Codex, etc.

LNAI is an innovative project designed to streamline the configuration management for various AI coding tools. Its primary function is to enable developers to define their preferred settings, preferences, and prompt engineering instructions once, and then automatically synchronize these configurations across multiple AI assistants and platforms. This centralized approach aims to reduce redundancy and ensure consistency in how developers interact with tools like Claude, Cursor, and Codex. By abstracting the configuration layer, LNAI enhances productivity, minimizes setup time, and helps maintain a uniform coding environment regardless of the specific AI coding tool being utilized. This project addresses the challenge of managing disparate configurations in an increasingly fragmented landscape of AI-powered development aids, offering a unified solution for workflow optimization and standard adherence.

huggingface

6 stories
01

Kimi K2.5: Visual Agentic Intelligence

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of techniques such as joint text-vision pre-training, zero-vision SFT, and joint text-vision reinforcement learning. Building on this multimodal foundation, K2.5 introduces Agent Swarm, a self-directed parallel agent orchestration framework that dynamically decomposes complex tasks into heterogeneous sub-problems and executes them concurrently. Extensive evaluations show that Kimi K2.5 achieves state-of-the-art results across various domains including coding, vision, reasoning, and agentic tasks. Agent Swarm also reduces latency by up to 4.5times over single-agent baselines. We release the post-trained Kimi K2.5 model checkpoint to facilitate future research and real-world applications of agentic intelligence.

02

RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System

We propose RLAnything, a reinforcement learning framework that dynamically forges environment, policy, and reward models through closed-loop optimization, amplifying learning signals and strengthening the overall RL system for any LLM or agentic scenarios. Specifically, the policy is trained with integrated feedback from step-wise and outcome signals, while the reward model is jointly optimized via consistency feedback, which in turn further improves policy training. Moreover, our theory-motivated automatic environment adaptation improves training for both the reward and policy models by leveraging critic feedback from each, enabling learning from experience. Empirically, each added component consistently improves the overall system, and RLAnything yields substantial gains across various representative LLM and agentic tasks, boosting Qwen3-VL-8B-Thinking by 9.1% on OSWorld and Qwen2.5-7B-Instruct by 18.7% and 11.9% on AlfWorld and LiveBench, respectively. We also that optimized reward-model signals outperform outcomes that rely on human labels. Code: https://github.com/Gen-Verse/Open-AgentRL

03

PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards

Text-to-video (T2V) generation aims to synthesize videos with high visual quality and temporal consistency that are semantically aligned with input text. Reward-based post-training has emerged as a promising direction to improve the quality and semantic alignment of generated videos. However, recent methods either rely on large-scale human preference annotations or operate on misaligned embeddings from pre-trained vision-language models, leading to limited scalability or suboptimal supervision. We present PISCES, an annotation-free post-training algorithm that addresses these limitations via a novel Dual Optimal Transport (OT)-aligned Rewards module. To align reward signals with human judgment, PISCES uses OT to bridge text and video embeddings at both distributional and discrete token levels, enabling reward supervision to fulfill two objectives: (i) a Distributional OT-aligned Quality Reward that captures overall visual quality and temporal coherence; and (ii) a Discrete Token-level OT-aligned Semantic Reward that enforces semantic, spatio-temporal correspondence between text and video tokens. To our knowledge, PISCES is the first to improve annotation-free reward supervision in generative post-training through the lens of OT. Experiments on both short- and long-video generation show that PISCES outperforms both annotation-based and annotation-free methods on VBench across Quality and Semantic scores, with human preference studies further validating its effectiveness. We show that the Dual OT-aligned Rewards module is compatible with multiple optimization paradigms, including direct backpropagation and reinforcement learning fine-tuning.

04

PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss

Pixel diffusion generates images directly in pixel space in an end-to-end manner, avoiding the artifacts and bottlenecks introduced by VAEs in two-stage latent diffusion. However, it is challenging to optimize high-dimensional pixel manifolds that contain many perceptually irrelevant signals, leaving existing pixel diffusion methods lagging behind latent diffusion models. We propose PixelGen, a simple pixel diffusion framework with perceptual supervision. Instead of modeling the full image manifold, PixelGen introduces two complementary perceptual losses to guide diffusion model towards learning a more meaningful perceptual manifold. An LPIPS loss facilitates learning better local patterns, while a DINO-based perceptual loss strengthens global semantics. With perceptual supervision, PixelGen surpasses strong latent diffusion baselines. It achieves an FID of 5.11 on ImageNet-256 without classifier-free guidance using only 80 training epochs, and demonstrates favorable scaling performance on large-scale text-to-image generation with a GenEval score of 0.79. PixelGen requires no VAEs, no latent representations, and no auxiliary stages, providing a simpler yet more powerful generative paradigm. Codes are publicly available at https://github.com/Zehong-Ma/PixelGen.

05

Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models

Multimodal large language models (MLLMs) have achieved remarkable success across a broad range of vision tasks. However, constrained by the capacity of their internal world knowledge, prior work has proposed augmenting MLLMs by "reasoning-then-tool-call" for visual and textual search engines to obtain substantial gains on tasks requiring extensive factual information. However, these approaches typically define multimodal search in a naive setting, assuming that a single full-level or entity-level image query and few text query suffices to retrieve the key evidence needed to answer the question, which is unrealistic in real-world scenarios with substantial visual noise. Moreover, they are often limited in the reasoning depth and search breadth, making it difficult to solve complex questions that require aggregating evidence from diverse visual and textual sources. Building on this, we propose Vision-DeepResearch, which proposes one new multimodal deep-research paradigm, i.e., performs multi-turn, multi-entity and multi-scale visual and textual search to robustly hit real-world search engines under heavy noise. Our Vision-DeepResearch supports dozens of reasoning steps and hundreds of engine interactions, while internalizing deep-research capabilities into the MLLM via cold-start supervision and RL training, resulting in a strong end-to-end multimodal deep-research MLLM. It substantially outperforming existing multimodal deep-research MLLMs, and workflows built on strong closed-source foundation model such as GPT-5, Gemini-2.5-pro and Claude-4-Sonnet. The code will be released in https://github.com/Osilly/Vision-DeepResearch.

06

Why Steering Works: Toward a Unified View of Language Model Parameter Dynamics

Methods for controlling large language models (LLMs), including local weight fine-tuning, LoRA-based adaptation, and activation-based interventions, are often studied in isolation, obscuring their connections and making comparison difficult. In this work, we present a unified view that frames these interventions as dynamic weight updates induced by a control signal, placing them within a single conceptual framework. Building on this view, we propose a unified preference-utility analysis that separates control effects into preference, defined as the tendency toward a target concept, and utility, defined as coherent and task-valid generation, and measures both on a shared log-odds scale using polarity-paired contrastive examples. Across methods, we observe a consistent trade-off between preference and utility: stronger control increases preference while predictably reducing utility. We further explain this behavior through an activation manifold perspective, in which control shifts representations along target-concept directions to enhance preference, while utility declines primarily when interventions push representations off the model's valid-generation manifold. Finally, we introduce a new steering approach SPLIT guided by this analysis that improves preference while better preserving utility. Code is available at https://github.com/zjunlp/EasyEdit/blob/main/examples/SPLIT.md.