NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-24DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API

OpenAI has officially unveiled its latest iterations of large language models, GPT-5.5 and GPT-5.5 Pro, making them immediately accessible via its developer API. This significant announcement empowers developers and enterprises with direct access to what are anticipated to be more sophisticated and efficient AI capabilities. While specific details on the enhancements are often found in accompanying documentation, the introduction of a "Pro" variant typically implies a more powerful version, potentially offering increased context windows, improved reasoning accuracy, faster processing speeds, or specialized features tailored for advanced applications. This release highlights OpenAI's ongoing strategy to push the boundaries of artificial intelligence technology and integrate these advancements directly into the developer ecosystem. The availability through the API ensures seamless integration for building next-generation applications leveraging state-of-the-art natural language processing, content generation, and intelligent automation. This move is expected to catalyze innovation across various sectors, from customer service and content creation to complex data analysis and beyond.

02

DeepSeek v4

DeepSeek AI has officially unveiled DeepSeek V4, representing a significant advancement in their series of artificial intelligence models. Although the initial announcement is brief, referencing a technical paper accessible via Hugging Face titled "DeepSeek_V4.pdf", it signals a major new release within the AI community. This new iteration from DeepSeek AI is expected to showcase substantial improvements in its underlying architecture, training methodologies, and overall performance benchmarks. Given DeepSeek AI's focus, DeepSeek V4 is highly anticipated to be a large language model, offering enhanced capabilities in natural language understanding, generation, and complex reasoning tasks. The provision of a detailed technical document suggests that the release is aimed at providing in-depth insights for researchers, developers, and practitioners into the model's specifications, evaluation results, and potential applications. This launch reinforces DeepSeek AI's commitment to advancing the frontiers of deep learning and generative AI, contributing to the broader landscape of AI research and development. It highlights the continuous innovation driving the progression of sophisticated AI systems, offering new tools and possibilities for various applications.

03

Show HN: Browser Harness  Gives LLM freedom to complete any browser task

Browser Harness introduces a novel approach to browser automation for Large Language Models (LLMs) by eliminating traditional restrictive frameworks. The project focuses on granting LLMs maximum operational freedom, allowing them to leverage their pre-trained capabilities more effectively. It incorporates a comprehensive "Browser Use library" built upon tens of thousands of lines of deterministic heuristics, which interact with Chrome via the Chrome DevTools Protocol (CDP) websocket. This library handles complex browser interactions, including element extraction, click assistance, intricate target management, and robust crash handling through watchdogs. A key feature is the LLM's ability to self-correct and dynamically add new tools to manage challenging edge cases, such as native file popups, which typically stall agents in conventional setups. This design aims to overcome common limitations in browser automation for AI agents.

04

Tesla (TSLA) discloses $2B AI hardware company acquisition buried

Tesla (TSLA) has reportedly disclosed a significant acquisition, purchasing an unnamed AI hardware company for an estimated $2 billion. This crucial detail, noted as 'buried' within recent financial disclosures, points to Tesla's aggressive strategy in fortifying its artificial intelligence ecosystem. The acquisition is poised to substantially enhance Tesla's internal capabilities in developing advanced AI processing units, which are vital for powering its autonomous driving systems, humanoid robots like Optimus, and its burgeoning supercomputing infrastructure, including the Dojo platform. By integrating specialized AI hardware expertise, Tesla aims to accelerate the performance and efficiency of its neural network training and inference operations. This substantial investment of $2 billion in AI hardware signifies a strategic move to vertically integrate key technological components, reducing reliance on external suppliers and securing a competitive advantage in the rapidly evolving fields of AI and robotics. The development is expected to have long-term implications for Tesla's technological independence and innovation trajectory in AI.

05

Google investing up to $40B in Anthropic

Google is reportedly committing up to $40 billion in investment to Anthropic, an artificial intelligence startup known for its Claude large language models and a key competitor in the generative AI space. This substantial financial commitment underscores Google's aggressive strategy to bolster its position in the rapidly evolving generative AI market and compete more effectively with rivals such as Microsoft-backed OpenAI. The investment, which expands on Google's prior funding in Anthropic, signals a deepening strategic partnership that aims to secure critical talent, technology, and market share in the foundational model space. This move highlights the high stakes and intense competition within the AI industry, where major tech companies are pouring immense resources into developing and deploying cutting-edge AI capabilities. The significant capital injection reflects the burgeoning valuations within the AI sector and is expected to accelerate the development of new AI products and services from Anthropic, potentially influencing the future landscape of artificial intelligence innovation and adoption across various industries.

06

South Korea police arrest man for posting AI photo of runaway wolf

The incident in South Korea, where police apprehended an individual for circulating an AI-generated image of a supposedly runaway wolf, underscores critical contemporary challenges posed by advanced artificial intelligence technologies. This event illustrates the escalating capability of generative AI to produce highly realistic imagery, blurring the lines between authentic and fabricated content. Such misinformation, particularly concerning public safety or animal welfare, can provoke widespread panic, misallocate emergency resources, and erode public trust in official communications. The police intervention signals a growing global concern among authorities regarding the misuse of AI for deceptive purposes, necessitating robust legal and ethical responses. It highlights the imperative for developing effective digital forensics tools capable of identifying AI-generated content and for fostering widespread media literacy to equip the public with skills to critically evaluate digital information. This case serves as a poignant reminder of the dual nature of AI innovation, emphasizing the need for balanced development that considers both technological advancement and its profound societal, ethical, and legal ramifications in an increasingly digitized world. The incident compels stakeholders to address accountability for digital fabrications and reinforces the importance of discerning factual information from synthetic visuals to maintain societal order.

huggingface

6 stories
01

Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Long horizon interactive environments are a testbed for evaluating agents skill usage abilities. These environments demand multi step reasoning, the chaining of multiple skills over many timesteps, and robust decision making under delayed rewards and partial observability. Games are a good testbed for evaluating agent skill usage in environments. Large Language Models (LLMs) offer a promising alternative as game playing agents, but they often struggle with consistent long horizon decision making because they lack a mechanism to discover, retain, and reuse structured skills across episodes. We present COSPLAY, a co evolution framework in which an LLM decision agent retrieves skills from a learnable skill bank to guide action taking, while an agent managed skill pipeline discovers reusable skills from the agents unlabeled rollouts to form a skill bank. Our framework improves both the decision agent to learn better skill retrieval and action generation, while the skill bank agent continually extracts, refines, and updates skills together with their contracts. Experiments across six game environments show that COSPLAY with an 8B base model achieves over 25.1 percent average reward improvement against four frontier LLM baselines on single player game benchmarks while remaining competitive on multi player social reasoning games.

02

VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automation

Autonomous GUI agents face two fundamental challenges: early stopping, where agents prematurely declare success without verifiable evidence, and repetitive loops, where agents cycle through the same failing actions without recovery. We present VLAA-GUI, a modular GUI agentic framework built around three integrated components that guide the system on when to Stop, Recover, and Search. First, a mandatory Completeness Verifier enforces UI-observable success criteria and verification at every finish step -- with an agent-level verifier that cross-examines completion claims with decision rules, rejecting those lacking direct visual evidence. Second, a mandatory Loop Breaker provides multi-tier filtering: switching interaction mode after repeated failures, forcing strategy changes after persistent screen-state recurrence, and binding reflection signals to strategy shifts. Third, an on-demand Search Agent searches online for unfamiliar workflows by directly querying a capable LLM with search ability, returning results as plain text. We additionally integrate a Coding Agent for code-intensive actions and a Grounding Agent for precise action grounding, both invoked on demand when required. We evaluate VLAA-GUI across five top-tier backbones, including Opus 4.5, 4.6 and Gemini 3.1 Pro, on two benchmarks with Linux and Windows tasks, achieving top performance on both (77.5% on OSWorld and 61.0% on WindowsAgentArena). Notably, three of the five backbones surpass human performance (72.4%) on OSWorld in a single pass. Ablation studies show that all three proposed components consistently improve a strong backbone, while a weaker backbone benefits more from these tools when the step budget is sufficient. Further analysis also shows that the Loop Breaker nearly halves wasted steps for loop-prone models.

03

WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning

While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic multi-page websites remain highly challenging. Existing works are often limited to single-page static websites, while agentic frameworks typically rely on multi-turn execution with proprietary models, leading to substantial token costs, high latency, and brittle integration. Training a small LLM end-to-end with reinforcement learning (RL) is a promising alternative, yet it faces a critical bottleneck in designing reliable and computationally feasible rewards for website generation. Unlike single-file coding tasks that can be verified by unit tests, website generation requires evaluating inherently subjective aesthetics, cross-page interactions, and functional correctness. To this end, we propose WebGen-R1, an end-to-end RL framework tailored for project-level website generation. We first introduce a scaffold-driven structured generation paradigm that constrains the large open-ended action space and preserves architectural integrity. We then design a novel cascaded multimodal reward that seamlessly couples structural guarantees with execution-grounded functional feedback and vision-based aesthetic supervision. Extensive experiments demonstrate that our WebGen-R1 substantially transforms a 7B base model from generating nearly nonfunctional websites into producing deployable, aesthetically aligned multi-page websites. Remarkably, our WebGen-R1 not only consistently outperforms heavily scaled open-source models (up to 72B), but also rivals the state-of-the-art DeepSeek-R1 (671B) in functional success, while substantially exceeding it in valid rendering and aesthetic alignment. These results position WebGen-R1 as a viable path for scaling small open models from function-level code generation to project-level web application generation.

04

Context Unrolling in Omni Models

We present Omni, a unified multimodal model natively trained on diverse modalities, including text, images, videos, 3D geometry, and hidden representations. We find that such training enables Context Unrolling, where the model explicitly reasons across multiple modal representations before producing predictions. This process enables the model to aggregate complementary information across heterogeneous modalities, facilitating a more faithful approximation of the shared multimodal knowledge manifold and improving downstream reasoning fidelity. As a result, Omni achieves strong performance on both multimodal generation and understanding benchmarks, while demonstrating advanced multimodal reasoning capabilities, including in-context generation of text, image, video, and 3D geometry.

05

WorldMark: A Unified Benchmark Suite for Interactive Video World Models

Interactive video generation models such as Genie, YUME, HY-World, and Matrix-Game are advancing rapidly, yet every model is evaluated on its own benchmark with private scenes and trajectories, making fair cross-model comparison impossible. Existing public benchmarks offer useful metrics such as trajectory error, aesthetic scores, and VLM-based judgments, but none supplies the standardized test conditions -- identical scenes, identical action sequences, and a unified control interface -- needed to make those metrics comparable across models with heterogeneous inputs. We introduce WorldMark, the first benchmark that provides such a common playing field for interactive Image-to-Video world models. WorldMark contributes: (1) a unified action-mapping layer that translates a shared WASD-style action vocabulary into each model's native control format, enabling apples-to-apples comparison across six major models on identical scenes and trajectories; (2) a hierarchical test suite of 500 evaluation cases covering first- and third-person viewpoints, photorealistic and stylized scenes, and three difficulty tiers from Easy to Hard spanning 20-60s; and (3) a modular evaluation toolkit for Visual Quality, Control Alignment, and World Consistency, designed so that researchers can reuse our standardized inputs while plugging in their own metrics as the field evolves. We will release all data, evaluation code, and model outputs to facilitate future research. Beyond offline metrics, we launch World Model Arena (warena.ai), an online platform where anyone can pit leading world models against each other in side-by-side battles and watch the live leaderboard.

06

UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling

Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that establishes a unified physical language for human-to-humanoid transfer. Grounded in the philosophy that heterogeneous kinematics share universal visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch synergies these purified modalities into a shared discrete latent space of embodiment-agnostic physical intents. We validate UniT across two paradigms: 1) Policy Learning (VLA-UniT): By predicting these unified tokens, it effectively leverages diverse human data to achieve state-of-the-art data efficiency and robust out-of-distribution (OOD) generalization on both humanoid simulation benchmark and real-world deployments, notably demonstrating zero-shot task transfer. 2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it realizes direct human-to-humanoid action transfer. This alignment ensures that human data seamlessly translates into enhanced action controllability for humanoid video generation. Ultimately, by inducing a highly aligned cross-embodiment representation (empirically verified by t-SNE visualizations revealing the convergence of human and humanoid features into a shared manifold), UniT offers a scalable path to distill vast human knowledge into general-purpose humanoid capabilities.