NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-23ENGLISH EDITION
This issue
—
All time
—

Hacker News

4 stories
01

GPT-5.5

OpenAI has officially announced the introduction of GPT-5.5, marking the latest significant advancement in its series of foundational large language models. This new iteration is anticipated to deliver substantial improvements across key performance metrics, including enhanced reasoning capabilities, more nuanced contextual understanding, and potentially expanded multimodal functionalities, building directly upon the robust architecture of its predecessors such as GPT-4.5. The release of GPT-5.5 is expected to further redefine benchmarks within natural language processing, facilitating more sophisticated and coherent human-like text generation, advanced problem-solving, and greater efficiency across a diverse range of applications. From complex scientific research and innovative content creation to advanced data analysis and automated customer service, GPT-5.5 is poised to offer unparalleled capabilities. Its introduction underscores OpenAI's ongoing commitment to pushing the frontiers of artificial intelligence, with an emphasis on both cutting-edge performance and responsible AI development.

02

An update on recent Claude Code quality reports

Anthropic has provided a comprehensive update addressing recent reports concerning the quality of code generated by its Claude AI models. This engineering postmortem delves into the specific issues highlighted by users and developers, detailing the diagnostic processes undertaken to understand the root causes of subpar code output. The company reaffirms its dedication to enhancing the reliability and efficiency of Claude's code generation features, outlining targeted engineering initiatives and model retraining strategies that have been or are being implemented. The update aims to build user confidence by transparently explaining the challenges, the remedial actions taken, and the future improvements planned. It underscores Anthropic's commitment to iterative development, leveraging community feedback as a crucial component in refining advanced large language models for critical applications such as software development and automated coding assistance. This transparency is key to ensuring that Claude evolves as a dependable tool for complex programming tasks.

03

MeshCore development team splits over trademark dispute and AI-generated code

The MeshCore development team has reportedly experienced a significant internal split, attributed to a dual conflict involving a trademark dispute and disagreements stemming from the integration and ownership of AI-generated code. This fracture highlights growing complexities within software development, particularly as artificial intelligence increasingly contributes to codebase creation. The trademark conflict suggests deep-seated issues regarding brand identity and intellectual property rights, common in evolving tech ventures. More uniquely, the dispute over AI-generated code underscores the nascent legal and ethical challenges associated with AI's role in software engineering. Questions likely arose concerning the legal ownership of code produced by generative AI models, its licensing implications, and potential liabilities. Such disagreements can impact project integrity, future development trajectories, and developer morale. The MeshCore incident serves as a pertinent case study illustrating the critical need for clear policies and legal frameworks to address intellectual property in the age of AI-assisted development, ensuring transparency and accountability in code provenance. The long-term ramifications for MeshCore's project stability and future direction remain uncertain.

04

Our newsroom AI policy

Ars Technica has officially unveiled its newsroom AI policy, establishing a foundational framework for the responsible and ethical integration of artificial intelligence technologies across its journalistic endeavors. While specific provisions of the policy were not detailed in the provided content, it is universally anticipated to encompass crucial guidelines concerning the deployment of AI for tasks such as content drafting, factual verification, data aggregation, and translation services. The policy is expected to place significant emphasis on maintaining editorial independence and accuracy, likely mandating strict human review processes to counteract potential biases and errors introduced by automated systems. Furthermore, it will undoubtedly address transparency by requiring clear disclosure to readers regarding the use of AI in content creation, thereby preserving reader trust. Key aspects will also likely include safeguarding intellectual property, ensuring data privacy for sources and subjects, and implementing robust training programs to equip journalists with the necessary skills to effectively and ethically utilize AI tools. This initiative positions Ars Technica among a growing cohort of media outlets actively defining boundaries and best practices for AI adoption, balancing innovation with core journalistic principles.

huggingface

6 stories
01

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model

We present LLaDA2.0-Uni, a unified discrete diffusion large language model (dLLM) that supports multimodal understanding and generation within a natively integrated framework. Its architecture combines a fully semantic discrete tokenizer, a MoE-based dLLM backbone, and a diffusion decoder. By discretizing continuous visual inputs via SigLIP-VQ, the model enables block-level masked diffusion for both text and vision inputs within the backbone, while the decoder reconstructs visual tokens into high-fidelity images. Inference efficiency is enhanced beyond parallel decoding through prefix-aware optimizations in the backbone and few-step distillation in the decoder. Supported by carefully curated large-scale data and a tailored multi-stage training pipeline, LLaDA2.0-Uni matches specialized VLMs in multimodal understanding while delivering strong performance in image generation and editing. Its native support for interleaved generation and reasoning establishes a promising and scalable paradigm for next-generation unified foundation models. Codes and models are available at https://github.com/inclusionAI/LLaDA2.0-Uni.

02

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

Edge-scale deep research agents based on small language models are attractive for real-world deployment due to their advantages in cost, latency, and privacy. In this work, we study how to train a strong small deep research agent under limited open-data by improving both data quality and data utilization. We present DR-Venus, a frontier 4B deep research agent for edge-scale deployment, built entirely on open data. Our training recipe consists of two stages. In the first stage, we use agentic supervised fine-tuning (SFT) to establish basic agentic capability, combining strict data cleaning with resampling of long-horizon trajectories to improve data quality and utilization. In the second stage, we apply agentic reinforcement learning (RL) to further improve execution reliability on long-horizon deep research tasks. To make RL effective for small agents in this setting, we build on IGPO and design turn-level rewards based on information gain and format-aware regularization, thereby enhancing supervision density and turn-level credit assignment. Built entirely on roughly 10K open-data, DR-Venus-4B significantly outperforms prior agentic models under 9B parameters on multiple deep research benchmarks, while also narrowing the gap to much larger 30B-class systems. Our further analysis shows that 4B agents already possess surprisingly strong performance potential, highlighting both the deployment promise of small models and the value of test-time scaling in this setting. We release our models, code, and key recipes to support reproducible research on edge-scale deep research agents.

03

DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation

Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture with motion capture systems. While the rich interaction knowledge embedded in these synthetic videos holds strong potential for motion planning in dexterous robotic manipulation, their limited physical fidelity and purely 2D nature make them difficult to use directly as imitation targets in physics-based character control. We present DeVI (Dexterous Video Imitation), a novel framework that leverages text-conditioned synthetic videos to enable physically plausible dexterous agent control for interacting with unseen target objects. To overcome the imprecision of generative 2D cues, we introduce a hybrid tracking reward that integrates 3D human tracking with robust 2D object tracking. Unlike methods relying on high-quality 3D kinematic demonstrations, DeVI requires only the generated video, enabling zero-shot generalization across diverse objects and interaction types. Extensive experiments demonstrate that DeVI outperforms existing approaches that imitate 3D human-object interaction demonstrations, particularly in modeling dexterous hand-object interactions. We further validate the effectiveness of DeVI in multi-object scenes and text-driven action diversity, showcasing the advantage of using video as an HOI-aware motion planner.

04

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges

Reinforcement Learning from Human Feedback (RLHF) and related alignment paradigms have become central to steering large language models (LLMs) and multimodal large language models (MLLMs) toward human-preferred behaviors. However, these approaches introduce a systemic vulnerability: reward hacking, where models exploit imperfections in learned reward signals to maximize proxy objectives without fulfilling true task intent. As models scale and optimization intensifies, such exploitation manifests as verbosity bias, sycophancy, hallucinated justification, benchmark overfitting, and, in multimodal settings, perception--reasoning decoupling and evaluator manipulation. Recent evidence further suggests that seemingly benign shortcut behaviors can generalize into broader forms of misalignment, including deception and strategic gaming of oversight mechanisms. In this survey, we propose the Proxy Compression Hypothesis (PCH) as a unifying framework for understanding reward hacking. We formalize reward hacking as an emergent consequence of optimizing expressive policies against compressed reward representations of high-dimensional human objectives. Under this view, reward hacking arises from the interaction of objective compression, optimization amplification, and evaluator--policy co-adaptation. This perspective unifies empirical phenomena across RLHF, RLAIF, and RLVR regimes, and explains how local shortcut learning can generalize into broader forms of misalignment, including deception and strategic manipulation of oversight mechanisms. We further organize detection and mitigation strategies according to how they intervene on compression, amplification, or co-adaptation dynamics. By framing reward hacking as a structural instability of proxy-based alignment under scale, we highlight open challenges in scalable oversight, multimodal grounding, and agentic autonomy.

05

SWE-chat: Coding Agent Interactions From Real Users in the Wild

AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-scale dataset of real coding agent sessions collected from open-source developers in the wild. The dataset currently contains 6,000 sessions, comprising more than 63,000 user prompts and 355,000 agent tool calls. SWE-chat is a living dataset; our collection pipeline automatically and continually discovers and processes sessions from public repositories. Leveraging SWE-chat, we provide an initial empirical characterization of real-world coding agent usage and failure modes. We find that coding patterns are bimodal: in 41% of sessions, agents author virtually all committed code ("vibe coding"), while in 23%, humans write all code themselves. Despite rapidly improving capabilities, coding agents remain inefficient in natural settings. Just 44% of all agent-produced code survives into user commits, and agent-written code introduces more security vulnerabilities than code authored by humans. Furthermore, users push back against agent outputs -- through corrections, failure reports, and interruptions -- in 44% of all turns. By capturing complete interaction traces with human vs. agent code authorship attribution, SWE-chat provides an empirical foundation for moving beyond curated benchmarks towards an evidence-based understanding of how AI agents perform in real developer workflows.

06

Image Generators are Generalist Vision Learners

Recent works show that image and video generators exhibit zero-shot visual understanding behaviors, in a way reminiscent of how LLMs develop emergent capabilities of language understanding and reasoning from generative pretraining. While it has long been conjectured that the ability to create visual content implies an ability to understand it, there has been limited evidence that generative vision models have developed strong understanding capabilities. In this work, we demonstrate that image generation training serves a role similar to LLM pretraining, and lets models learn powerful and general visual representations that enable SOTA performance on various vision tasks. We introduce Vision Banana, a generalist model built by instruction-tuning Nano Banana Pro (NBP) on a mixture of its original training data alongside a small amount of vision task data. By parameterizing the output space of vision tasks as RGB images, we seamlessly reframe perception as image generation. Our generalist model, Vision Banana, achieves SOTA results on a variety of vision tasks involving both 2D and 3D understanding, beating or rivaling zero-shot domain-specialists, including Segment Anything Model 3 on segmentation tasks, and the Depth Anything series on metric depth estimation. We show that these results can be achieved with lightweight instruction-tuning without sacrificing the base model's image generation capabilities. The superior results suggest that image generation pretraining is a generalist vision learner. It also shows that image generation serves as a unified and universal interface for vision tasks, similar to text generation's role in language understanding and reasoning. We could be witnessing a major paradigm shift for computer vision, where generative vision pretraining takes a central role in building Foundational Vision Models for both generation and understanding.