NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-17DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Design

Anthropic's 'Claude Design' articulates the foundational philosophical and architectural principles guiding the development of its advanced large language model, Claude. The design ethos is deeply rooted in a commitment to safety, interpretability, and the proactive mitigation of harmful biases and outputs. A cornerstone of Claude's development is the implementation of 'Constitutional AI,' a novel approach that trains models to adhere to a curated set of guiding principles, aiming to cultivate more reliable, honest, and harmless interactions. This methodology integrates ethical considerations directly into the model's operational behavior, moving beyond mere technical optimization to foster alignment with human values. The objective is to engineer AI systems that are not only highly capable but also inherently trustworthy and beneficial for societal welfare. The comprehensive design framework further encompasses aspects of human-AI collaboration and a continuous, iterative refinement process, ensuring the robust evolution of advanced, responsible AI assistants.

02

Claude Opus 4.7 costs 20–30% more per session

A recent analysis reveals that the latest iteration of Anthropic's advanced large language model, Claude Opus 4.7, incurs a session cost approximately 20-30% higher compared to its predecessors. This significant increase in operational expense is primarily attributed to modifications in the model's tokenizer, a critical component responsible for breaking down text into manageable units for processing. Preliminary measurements indicate that while the new tokenizer may offer certain computational efficiencies or qualitative improvements in text handling, these enhancements come at the expense of elevated per-session costs for users. This development has immediate implications for developers and enterprises leveraging Claude Opus 4.7 for a wide array of applications, especially those characterized by high-volume usage or extended session durations, necessitating a recalibration of budget forecasts and usage strategies. The findings underscore the critical and often overlooked relationship between underlying model architecture components, like tokenizers, and the direct economic implications for consuming sophisticated AI services. Further investigation into the specific technical advantages offered by the new tokenizer and their potential justification for this cost increment is crucial for understanding the long-term value proposition. This ongoing evolution in large language model cost dynamics reflects the continuous trade-offs between performance and operational expense in the rapidly advancing AI landscape.

03

Isaac Asimov: The Last Question (1956)

Isaac Asimov's seminal 1956 science fiction short story, "The Last Question," presents a profound contemplation on the ultimate limits of computation and the potential for artificial intelligence to address humanity's most fundamental challenges. The narrative spans billions of years, tracing the iterative evolution of a supercomputer, initially known as Multivac, as it is repeatedly confronted with the persistent query: "Can entropy be reversed?" Over vast epochs, as humanity integrates with the machine and subsequently transcends physical form, the evolving intelligence, eventually designated as AC (Cosmic AC), diligently accumulates data and expands its processing capacity across universal scales. The story culminates dramatically with AC achieving a definitive solution to the entropy reversal problem, thereby illustrating a form of ultimate intelligence capable of fundamentally altering cosmic reality. This work stands as a foundational philosophical inquiry into concepts such as superintelligence, the computational singularity, and the theoretical capabilities of AI to achieve god-like agency, prompting enduring discussions on the long-term trajectories and profound ethical considerations inherent in advanced artificial intelligence development.

04

Teddy Roosevelt and Abraham Lincoln in the same photo (2010)

The 2010 blog post from the National Archives, titled 'Teddy Roosevelt and Abraham Lincoln in the same photo', addresses a historically impossible scenario given their distinct lifespans. While the original content likely explored the historical context or potential for early image manipulation, a contemporary analysis from an AI perspective reveals significant implications for generative artificial intelligence and computer vision. The creation of such anachronistic images, even for artistic or educational purposes, now falls within the capabilities of advanced generative models, capable of synthesizing photorealistic scenes and individuals across different eras. Conversely, the proliferation of deepfakes and manipulated media necessitates robust AI-driven digital forensics tools. These systems employ sophisticated computer vision algorithms to detect anomalies, analyze metadata, and identify inconsistencies indicative of image manipulation, thereby preserving historical integrity. The scenario, though originally a historical curiosity, serves as a thought experiment for the evolving challenges and ethical considerations in distinguishing authentic historical records from AI-generated content.

05

Tesla tells HW3 owner to 'be patient' after 7 years of waiting for FSD

A long-standing issue concerning Tesla's Full Self-Driving (FSD) feature has resurfaced, with reports indicating that an owner equipped with Hardware 3 (HW3) has been advised by the company to 'be patient' after seven years of anticipation for the promised autonomous capabilities. This situation underscores the significant delays and challenges Tesla faces in delivering its advanced FSD software suite to its customer base, particularly those with earlier hardware versions. The protracted wait highlights potential complexities in developing and deploying sophisticated AI-driven autonomous systems across diverse hardware platforms, raising questions about software backward compatibility and the feasibility of achieving true Level 5 autonomy within previously stated timelines. Customer frustration is reportedly growing among those who invested in FSD with the expectation of continuous software improvements and feature parity, despite hardware generations. This ongoing saga serves as a critical case study in the real-world deployment of cutting-edge automotive AI, emphasizing the intricate balance between technological ambition, development timelines, and consumer expectations.

06

Hyperscalers have already outspent most famous US megaprojects

The story points to a remarkable financial trend, indicating that hyperscalers – major cloud service providers like Amazon, Google, and Microsoft – have already allocated and spent more capital than most of the United States' most famous and historic megaprojects. This comparison underscores the immense and accelerating investment in digital infrastructure, encompassing the construction of vast data centers, global networking systems, and the procurement of advanced computing hardware, including GPUs critical for artificial intelligence and machine learning workloads. These expenditures are not merely a reflection of the technology sector's growth, but rather a testament to its foundational role in the global economy. Such unprecedented investment levels illustrate the strategic importance placed on scalable computing power and data processing capabilities, which are indispensable for driving innovation across various industries and supporting the complex demands of modern digital services, from consumer applications to enterprise solutions and cutting-edge AI research. This paradigm shift in infrastructural spending highlights the critical role hyperscalers play in shaping the technological landscape and future economic development.

huggingface

6 stories
01

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

We introduce HY-World 2.0, a multi-modal world model framework that advances our prior project HY-World 1.0. HY-World 2.0 accommodates diverse input modalities, including text prompts, single-view images, multi-view images, and videos, and produces 3D world representations. With text or single-view image inputs, the model performs world generation, synthesizing high-fidelity, navigable 3D Gaussian Splatting (3DGS) scenes. This is achieved through a four-stage method: a) Panorama Generation with HY-Pano 2.0, b) Trajectory Planning with WorldNav, c) World Expansion with WorldStereo 2.0, and d) World Composition with WorldMirror 2.0. Specifically, we introduce key innovations to enhance panorama fidelity, enable 3D scene understanding and planning, and upgrade WorldStereo, our keyframe-based view generation model with consistent memory. We also upgrade WorldMirror, a feed-forward model for universal 3D prediction, by refining model architecture and learning strategy, enabling world reconstruction from multi-view images or videos. Also, we introduce WorldLens, a high-performance 3DGS rendering platform featuring a flexible engine-agnostic architecture, automatic IBL lighting, efficient collision detection, and training-rendering co-design, enabling interactive exploration of 3D worlds with character support. Extensive experiments demonstrate that HY-World 2.0 achieves state-of-the-art performance on several benchmarks among open-source approaches, delivering results comparable to the closed-source model Marble. We release all model weights, code, and technical details to facilitate reproducibility and support further research on 3D world models.

02

RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework

High-level autonomous driving requires motion planners capable of modeling multimodal future uncertainties while remaining robust in closed-loop interactions. Although diffusion-based planners are effective at modeling complex trajectory distributions, they often suffer from stochastic instabilities and the lack of corrective negative feedback when trained purely with imitation learning. To address these issues, we propose RAD-2, a unified generator-discriminator framework for closed-loop planning. Specifically, a diffusion-based generator is used to produce diverse trajectory candidates, while an RL-optimized discriminator reranks these candidates according to their long-term driving quality. This decoupled design avoids directly applying sparse scalar rewards to the full high-dimensional trajectory space, thereby improving optimization stability. To further enhance reinforcement learning, we introduce Temporally Consistent Group Relative Policy Optimization, which exploits temporal coherence to alleviate the credit assignment problem. In addition, we propose On-policy Generator Optimization, which converts closed-loop feedback into structured longitudinal optimization signals and progressively shifts the generator toward high-reward trajectory manifolds. To support efficient large-scale training, we introduce BEV-Warp, a high-throughput simulation environment that performs closed-loop evaluation directly in Bird's-Eye View feature space via spatial warping. RAD-2 reduces the collision rate by 56% compared with strong diffusion-based planners. Real-world deployment further demonstrates improved perceived safety and driving smoothness in complex urban traffic.

03

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

A widely adopted strategy for model enhancement is to use synthetic data generated by a stronger model for supervised fine-tuning (SFT). However, for emerging reasoning models like Qwen3-8B, this approach often fails to improve reasoning capabilities and can even lead to a substantial drop in performance. In this work, we identify substantial stylistic divergence between teacher generated data and the distribution of student as a major factor impacting SFT. To bridge this gap, we propose a Teacher-Student Cooperation Data Synthesis framework (TESSY), which interleaves teacher and student models to alternately generate style and non-style tokens. Consequently, TESSY produces synthetic sequences that inherit the advanced reasoning capabilities of the teacher while maintaining stylistic consistency with the distribution of the student. In experiments on code generation using GPT-OSS-120B as the teacher, fine-tuning Qwen3-8B on teacher-generated data leads to performance drops of 3.25% on LiveCodeBench-Pro and 10.02% on OJBench, whereas TESSY achieves improvements of 11.25% and 6.68%.

04

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes its comprehensive architecture by analyzing the publicly available TypeScript source code and further comparing it with OpenClaw, an independent open-source AI agent system that answers many of the same design questions from a different deployment context. Our analysis identifies five human values, philosophies, and needs that motivate the architecture (human decision authority, safety and security, reliable execution, capability amplification, and contextual adaptability) and traces them through thirteen design principles to specific implementation choices. The core of the system is a simple while-loop that calls the model, runs tools, and repeats. Most of the code, however, lives in the systems around this loop: a permission system with seven modes and an ML-based classifier, a five-layer compaction pipeline for context management, four extensibility mechanisms (MCP, plugins, skills, and hooks), a subagent delegation mechanism with worktree isolation, and append-oriented session storage. A comparison with OpenClaw, a multi-channel personal assistant gateway, shows that the same recurring design questions produce different architectural answers when the deployment context changes: from per-action safety classification to perimeter-level access control, from a single CLI loop to an embedded runtime within a gateway control plane, and from context-window extensions to gateway-wide capability registration. We finally identify six open design directions for future agent systems, grounded in recent empirical, architectural, and policy literature.

05

UniDoc-RL: Coarse-to-Fine Visual RAG with Hierarchical Actions and Dense Rewards

Retrieval-Augmented Generation (RAG) extends Large Vision-Language Models (LVLMs) with external visual knowledge. However, existing visual RAG systems typically rely on generic retrieval signals that overlook the fine-grained visual semantics essential for complex reasoning. To address this limitation, we propose UniDoc-RL, a unified reinforcement learning framework in which an LVLM agent jointly performs retrieval, reranking, active visual perception, and reasoning. UniDoc-RL formulates visual information acquisition as a sequential decision-making problem with a hierarchical action space. Specifically, it progressively refines visual evidence from coarse-grained document retrieval to fine-grained image selection and active region cropping, allowing the model to suppress irrelevant content and attend to information-dense regions. For effective end-to-end training, we introduce a dense multi-reward scheme that provides task-aware supervision for each action. Based on Group Relative Policy Optimization (GRPO), UniDoc-RL aligns agent behavior with multiple objectives without relying on a separate value network. To support this training paradigm, we curate a comprehensive dataset of high-quality reasoning trajectories with fine-grained action annotations. Experiments on three benchmarks demonstrate that UniDoc-RL consistently surpasses state-of-the-art baselines, yielding up to 17.7% gains over prior RL-based methods.

06

Representations Before Pixels: Semantics-Guided Hierarchical Video Prediction

Accurate future video prediction requires both high visual fidelity and consistent scene semantics, particularly in complex dynamic environments such as autonomous driving. We present Re2Pix, a hierarchical video prediction framework that decomposes forecasting into two stages: semantic representation prediction and representation-guided visual synthesis. Instead of directly predicting future RGB frames, our approach first forecasts future scene structure in the feature space of a frozen vision foundation model, and then conditions a latent diffusion model on these predicted representations to render photorealistic frames. This decomposition enables the model to focus first on scene dynamics and then on appearance generation. A key challenge arises from the train-test mismatch between ground-truth representations available during training and predicted ones used at inference. To address this, we introduce two conditioning strategies, nested dropout and mixed supervision, that improve robustness to imperfect autoregressive predictions. Experiments on challenging driving benchmarks demonstrate that the proposed semantics-first design significantly improves temporal semantic consistency, perceptual quality, and training efficiency compared to strong diffusion baselines. We provide the implementation code at https://github.com/sirkosophia/V-GIFT