NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-28ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Opus 4.8

Anthropic has announced the release of Claude Opus 4.8, representing the latest advancement in their flagship model family. This update introduces significant performance enhancements across complex reasoning, advanced coding, and multilingual understanding tasks. Designed for enterprise-level applications, Claude Opus 4.8 demonstrates superior accuracy in analyzing unstructured data, executing intricate instructions, and generating highly reliable source code. The release highlights Anthropic's commitment to advancing frontiers in safety-critical AI deployments by implementing rigorous evaluation frameworks. Additionally, the model shows marked improvements in context retention and logical coherence, making it highly suitable for multi-step workflow automation and strategic decision-making support in professional environments.

02

Anthropic raises $65B in Series H funding at $965B post-money valuation

Anthropic has successfully raised sixty-five billion dollars in its Series H funding round, achieving an unprecedented post-money valuation of nine hundred and sixty-five billion dollars. This monumental capital injection marks one of the largest private financing rounds in the technology sector to date, highlighting the intense investor confidence in foundational artificial intelligence research. The funding is expected to accelerate Anthropic's research and development roadmap, specifically targeting the training of next-generation large language models and frontier safety-oriented AI systems. With this capital, Anthropic plans to scale its computational infrastructure significantly and expand its world-class research and engineering teams to push the boundaries of safe, helpful, and honest artificial intelligence. The historic valuation places Anthropic among the most valuable private technology companies globally, signaling a massive scale-up phase for competitive generative AI developers as they race to achieve artificial general intelligence capabilities.

03

Dynamic Workflows in Claude Code

Anthropic has introduced dynamic workflows in Claude Code, bringing a highly adaptive, agentic approach to software development directly within the terminal interface. Unlike traditional, rigid automation scripts, dynamic workflows enable Claude to intelligently plan, execute, and adjust its programming tasks based on real-time feedback from the development environment. The tool dynamically handles complex tasks such as code editing, testing, debugging, and git operations by determining the optimal sequence of actions autonomously. By utilizing feedback loops, Claude Code can detect test failures, inspect logs, and modify its strategies on the fly. This release represents a significant advancement in AI-driven coding assistants, moving beyond simple code completion to fully-fledged autonomous engineering agents capable of managing sophisticated technical lifecycles.

04

Disagreement among frontier LLMs on real-world fact-checks

This research investigates the performance and consistency of frontier large language models when evaluating real-world fact-checks. By analyzing how different state-of-the-art models handle factual verification tasks, the study reveals significant levels of disagreement among top-tier LLMs on complex truth claims. The findings highlight the current limitations of generative AI models in acting as reliable, autonomous arbitrators of truth. Furthermore, the research discusses how subtle variations in prompting strategies, underlying training data distributions, and internal alignment methodologies contribute to these conflicting outputs, emphasizing the urgent need for standardized benchmarks and robust evaluation frameworks in the field of automated fact-checking and information verification.

05

Show HN: Continue? Y/N: A 60-second game about AI agent permission fatigue

This interactive web-based showcase introduces 'Continue? Y/N', a brief 60-second conceptual game designed to highlight the emerging user experience challenge of permission fatigue in the era of autonomous AI agents. As AI systems increasingly execute complex workflows, search tasks, and system-level actions on behalf of humans, they constantly require authorization prompts. The game simulates this high-frequency decision-making environment, illustrating how rapidly users become desensitized to safety warnings and security prompts when repeatedly asked to confirm agent actions. By gamifying this friction point, the developer prompts a critical discussion on the balance between user control and automation convenience, emphasizing the necessity for smarter, context-aware authorization frameworks in future AI agent deployments.

06

Show HN: Ktx – Open-source executable context layer for data agents

Ktx is an open-source executable context layer designed to solve the critical accuracy challenges faced by LLM-powered data agents working with complex enterprise data stacks. While modern AI agents are highly capable of generating syntactically valid SQL, they frequently fail to produce semantically correct queries due to a lack of awareness regarding deprecated columns, hidden business logic, complex table joins, and organizational definitions. Built on top of extensive real-world experience deploying enterprise data agents, Ktx bridges this operational gap by serving as a dedicated contextual repository. It equips AI agents with the necessary semantic layer, metadata, and execution rules required to interact safely and accurately with data warehouses. By resolving common pitfalls like join fanouts and stale schema references, Ktx significantly increases the reliability and business-readiness of self-service data exploration systems.

Twitter

6 stories
01

Kling AI at AI on the Lot

Kling AI has announced its participation in AI on the Lot’s Community Day, recognized as the world’s largest conference dedicated to the intersection of artificial intelligence, film, and media production. During this prestigious event, the platform will showcase a curated selection of 20 original short films produced by members of the Prompt Club. These cinematic projects serve as a testament to the evolving capabilities of generative AI in storytelling, specifically demonstrating the ability to render high-fidelity content at native 4K resolution. This initiative highlights Kling AI's commitment to empowering filmmakers and pushing the creative boundaries of digital cinema. By bridging advanced technical generation with professional media standards, the showcase emphasizes how modern AI tools are reshaping visual production workflows and expanding the artistic potential for creators in the rapidly advancing film industry.

02

LumaLabsAI_BTS BTS Creation

Luma Labs AI has released a behind-the-scenes look at the technical production process behind their latest generative video showcase. The project highlights a complete workflow where every visual component, including individual characters, complex background scenes, and cinematic camera shots, was constructed entirely from scratch using advanced image generation and video synthesis technologies. By detailing the integration of these AI-driven generative tools, the team demonstrates how creators can now achieve high-quality, fully custom audiovisual content without relying on traditional filming or manual rendering techniques. This reveal serves as a practical blueprint for developers and creative professionals interested in the current capabilities of generative media tools, emphasizing the precision and control achievable within current AI video production pipelines, while inviting users to start creating their own projects via the Luma Labs platform.

03

c_valenzuelab_AI Video Future

Crist3bal Valenzuela discusses the rapid evolution of generative media, suggesting that the current trajectory of artificial intelligence adoption indicates a future where nearly all video content will be AI-generated. The tweet reflects on the implications of this paradigm shift, specifically challenging the utility of current labeling practices. As the line between authentic and synthetic content blurs, the author posits that identifying AI-generated material may eventually become as conceptually difficult as discerning traditional content creation methods. This perspective highlights the inevitable dominance of generative models in visual media production and raises significant questions regarding content provenance, digital authenticity, and the long-term societal integration of synthetic media as the default standard for video consumption across the creative industries.

04

ylecun_LLM and RL Debate

This tweet captures an engaging interaction involving Yann LeCun, focusing on his ongoing discourse regarding the limitations and potential of Large Language Models (LLMs) and Reinforcement Learning (RL). The discussion touches upon LeCun’s widely recognized skepticism regarding the sufficiency of current LLM architectures for achieving true world models or human-level intelligence. By positioning LLMs and RL in the context of world models, the conversation highlights the broader academic and industry debate about whether sequence prediction alone can capture the physical and causal nuances required for advanced reasoning. This exchange underscores the professional ongoing debate within the AI research community concerning the path toward Artificial General Intelligence, emphasizing the distinct viewpoints that differentiate mainstream scaling approaches from alternative architectural paradigms favored by visionaries like LeCun.

05

GoogleDeepMind_Nano Banana Launch

Google has officially announced the general availability of its latest image generation models, Nano Banana 2 and Nano Banana Pro. This release marks a significant milestone for developers seeking to integrate state-of-the-art generative capabilities into their applications. These models represent Google's most advanced work in image synthesis, offering enhanced fidelity, improved structural integrity, and greater creative flexibility compared to previous iterations. By making these powerful tools accessible for widespread development, Google aims to empower creators and engineers to build sophisticated visual content more efficiently. The rollout underscores Google's ongoing commitment to advancing multimodal artificial intelligence and providing robust, high-performance infrastructure for the global developer community to experiment with next-generation generative AI solutions for various professional and creative use cases.

06

GaryMarcus_Polymarket Spend

The tweet highlights a surprising development involving Polymarket, which reportedly incurred 500 million dollars in accidental spending on Claude AI services within a single month. This substantial expenditure, originally highlighted by reporting from Madison Mills at Axios, serves as a significant case study regarding the costs associated with scaling and operating large language model infrastructures. The incident underscores the potential financial risks and technical management challenges faced by high-growth AI startups when deploying advanced models at scale. As Polymarket leverages Claude for its predictive market operations, this disclosure provides crucial industry insight into the operational costs and technical complexities involved in maintaining cutting-edge AI architectures, drawing widespread attention from industry analysts and researchers monitoring the commercial integration of generative artificial intelligence technologies.

huggingface

6 stories
01

Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players

World models for interactive video generation have largely focused on single-agent settings, where future observations are generated from a single control signal. However, many generated environments require multi-agent interaction: multiple players, robots, or embodied agents act simultaneously within a shared space. Scaling world models to such settings requires a principled multi-agent design: agents should remain independently controllable, permutation-symmetric, and support efficient inference while maintaining consistency across time and perspectives. In this paper, we present our generative multi-agent world model for interactive simulation. It introduces Simplex Rotary Agent Encoding, a parameter-free extension of 3D RoPE that represents agents as vertices of a regular simplex in rotary angle space. This gives each agent a distinct phase while making all agents permutation-equivalent, enabling scalable agent identity without learned per-slot identities or a fixed agent ordering. To avoid dense all-to-all attention across agents, we further propose Sparse Hub Attention, where learnable hub tokens mediate token interaction across agents, reducing cross-agent attention cost from quadratic to linear in the number of agents. For real-time rollout, we distill a full-context diffusion teacher into a causal student that generates temporal blocks sequentially with KV caching, enabling action-responsive generation at 24 FPS. Experiments in multiplayer virtual environments show that our model improves video fidelity, action controllability, and inter-agent consistency over slot-based and dense-attention baselines, while generalizing from two to four players without additional training.

02

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

Vision-language models with extended reasoning succeed on complex problems, but many real-world problems require external tools that internal reasoning alone often cannot resolve. Agentic reasoning therefore interleaves two behaviors with a structural asymmetry: thinking (the self-contained default) and tool use (a high-variance auxiliary acting). We refer to this asymmetry as the Thinking-Acting Gap. Under standard RL recipes like GRPO, the gap manifests as two diagnostic symptoms during training: tool use is attempted on only ~30% of rollouts, and when attempted, the tool-using rollouts within a group are all-wrong on ~40% of questions, suppressing the learning signal at the tool calls that needed it. We propose AXPO (Agent eXplorative Policy Optimization): for each all-wrong tool-using subgroup, AXPO fixes the thinking prefix and resamples the tool call and its continuation, paired with uncertainty-based prefix selection. Across nine multimodal benchmarks and three scales of Qwen3-VL-Thinking, SFT+AXPO outperforms SFT+GRPO at average (+1.8pp Pass@1 and +1.8pp Pass@4 at 8B on average) and 8B with SFT+AXPO surpasses the 32B Base on Pass@4 with 4 times fewer parameters.

03

Self-Improving Language Models with Bidirectional Evolutionary Search

Search has been proposed as an effective method for self-improving language models and agentic systems, both for post-training sample generation and for inference. However, widely used methods such as best-of-N sampling and tree search face two fundamental limitations: they are guided by sparse verification signals, and they construct candidates primarily through autoregressive expansion, restricting exploration to regions with substantial model probability mass. To address these, we propose Bidirectional Evolutionary Search (BES), a search framework that couples forward candidate evolution with backward goal decomposition. In the forward search, BES augments standard expansion with evolution operators that recombine partial trajectories to generate candidates that are difficult to obtain from a single model rollout. In the backward search, BES recursively decomposes the original task into checkable subgoals, producing dense intermediate feedback that guides forward search. We provide theoretical motivation showing that candidates generated by expansion-only search are confined to a narrow entropy shell while evolutionary operators can escape it, and that backward search can exponentially reduce the number of required samples to find a correct answer. Experiments show that on challenging post-training tasks where mainstream post-training algorithms fail to improve, BES enables consistent gains, and on three open problem solving benchmarks at inference time, BES outperforms existing open-source frameworks in both average and best-case performance.

04

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.

05

OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning

Diffusion Transformers achieve strong video generation quality, but the quadratic cost of full attention limits efficiency. We introduce OSP-Next, an efficient text-to-video generation model that integrates sparse attention, parallelism, quantization, and reinforcement learning. OSP-Next uses a hybrid full-sparse attention architecture, where the sparse component is implemented with Skiparse-2D Attention. This fixed-pattern mechanism applies token-wise and group-wise sparse attention along spatial dimensions, leveraging locality while maintaining native compatibility with FlashAttention kernels. Based on the local equivalence of rearrangement in Skiparse-2D Attention, we further propose Sparse Sequence Parallelism (SSP), which partitions subsequences across ranks and switches sparse patterns through a single All-to-All communication. Compared with Ulysses Sequence Parallelism (SP), SSP provides a native parallel strategy for sparse attention and reduces communication volume by 75%. OSP-Next also incorporates HiF8 quantization to enable stable joint training with 8-bit quantization and sparse fine-tuning, and applies Mix-GRPO post-training to improve the performance of the sparse model. Experiments show that OSP-Next achieves a VBench total score of 83.73%, surpassing the Wan2.1 baseline. Under the 5-second 720P and 5-second 768P settings, OSP-Next achieves up to 1.64times single-GPU speedup and over 1.52times eight-GPU speedup on NVIDIA H200 GPUs. In addition, with only a 0.4% drop in VBench total score, OSP-Next-HiF8 achieves 1.69times and 2.27times speedups under the two settings on a single Ascend 950PR, demonstrating the efficiency and performance of OSP-Next across hardware platforms.

06

Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization

Recent advances in online reinforcement learning (RL) for large language models (LLMs) have demonstrated promising performance in complex reasoning tasks. However, they often exhibit an imbalanced exploration-exploitation trade-off, resulting in unstable optimization and sub-optimal performance. We introduce IB-Score, a novel metric grounded in Information Bottleneck theory that evaluates policy's exploration-exploitation balance by quantifying the trade-off between step-level reasoning diversity and mutual information shared with the correct answer. Analysis based on IB-Score shows that popular online RL approaches (e.g., GRPO) with common regularizers fail to consistently maintain balance during training with suboptimal results. To address this, we propose Information Bottleneck-driven Tree-based Policy Optimization (IB-TPO), a principled framework that formulates IB-Score as a fine-grained optimization objective and utilizes a novel IB-guided tree sampling strategy that not only improves the efficiency of online sampling with 50% more trajectories under the same token budget, but also reuses the tree structure for effective IB-Score Monte Carlo estimation. Extensive experiments across standard benchmarks show that our method significantly outperforms GRPO baseline by 2.9% to 3.6% and also outperforms other state-of-the-art online RL approaches.