NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-05-27ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Code as a Daily Driver: Claude.md, Skills, Subagents, Plugins, and MCPs

This article explores the implementation of Claude Code as a daily development tool, focusing on its integration with system files like Claude.md, custom skills, subagents, plugins, and the Model Context Protocol (MCP). It details how developers can configure Claude Code to automate complex programming tasks, manage multi-agent workflows, and extend its functionality through specialized plugins. By leveraging the Model Context Protocol, the tool establishes a standardized method for AI models to safely access external data sources and developer tools, significantly enhancing productivity and precision in daily software engineering workflows.

02

Multi-Agent LLM System for Automated Vulnerability Discovery and Reproduction

This research introduces a novel multi-agent Large Language Model system designed to automate the complex process of software vulnerability discovery and reproduction. By leveraging collaborative AI agents with specialized roles, the system can systematically analyze source code, identify potential security weaknesses, and automatically generate working exploits to verify these vulnerabilities. This automation significantly reduces the manual effort traditionally required by security researchers for triage and verification. The framework coordinates different LLM agents to perform tasks such as static analysis, dynamic testing, and exploit generation, showing promising results in accelerating patch verification and enhancing overall software security. The findings highlight the growing potential of autonomous AI agents in cybersecurity defense and offensive security testing, offering a scalable solution to secure modern software supply chains.

03

I think Anthropic and OpenAI have found product-market fit

This article analyzes the strategic milestones achieved by leading artificial intelligence research labs, Anthropic and OpenAI, arguing that both organizations have successfully established definitive product-market fit. By transitioning from general-purpose experimental models to targeted, high-utility developer APIs and enterprise-ready assistant platforms, these companies have secured stable, high-volume recurring revenue streams. The analysis highlights how the evolution of natural language interfaces, robust API integrations, and specialized developer tooling has cemented their positions as core infrastructure providers in the modern software ecosystem. Ultimately, the piece underscores that their utility extends beyond novel demonstrations into indispensable, everyday operational assets for global enterprises and software engineers alike.

04

I'm Tired of Talking to AI

This article discusses the growing user fatigue and frustration surrounding interactions with artificial intelligence, particularly the deluge of AI-generated answers in daily digital life. The author reflects on the diminishing quality of online content and search experiences when heavily mediated by generative AI systems. By examining the shift from human-curated resources to automated conversational agents, the piece highlights how persistent exposure to synthetic responses can feel impersonal, repetitive, and ultimately less helpful than authentic human exchanges. It addresses the broader implications of this trend on information quality, user experience design, and consumer trust, arguing that the industry needs a more balanced approach that values genuine human input and limits the over-saturation of automated conversational interfaces across standard web platforms.

05

DuckDuckGo search saw 28% more visits after Google said people love AI mode

DuckDuckGo experienced a significant 28% surge in traffic following Google's public claims regarding user enthusiasm for integrated AI search features. As Google continues to aggressively push AI Overviews and automated summaries within its primary search engine layout, a growing segment of the user base is actively seeking alternative platforms to bypass these algorithmic interventions. DuckDuckGo's positioning as an AI-free, privacy-centric search engine has successfully attracted users who prefer traditional organic search results over generative AI answers. This shift highlights emerging consumer friction regarding the mandatory integration of generative AI into daily digital utilities. It underscores a growing demand for classical web indexing solutions amidst the rapid industry-wide transition toward Large Language Model-driven search architectures, demonstrating that a substantial market segment still prioritizes clean, non-synthesized search interfaces over AI-generated outputs.

06

The AI tech job slaughter gets real

This article explores the accelerating trend of workforce reductions within the technology sector, driven by the rapid adoption and integration of generative artificial intelligence. As tech giants and startups pivot heavily toward AI development, traditional software engineering and IT support roles are being downsized or phased out entirely. Industry analysts discuss how companies are restructuring their organizations to prioritize AI expertise, rendering certain legacy technical skill sets obsolete. The report emphasizes the stark contrast between the booming job market for machine learning specialists and the severe job cuts affecting general IT professionals, marking a profound shift in the employment landscape of the tech industry.

Twitter

6 stories
01

GoogleDeepMind_Gemini Emb

Google DeepMind has officially released the white paper for Gemini Embedding 2, a significant advancement in native multimodal embedding technology. This new model represents a major leap in how machines represent and relate diverse data types, including text, images, and other modalities, within a unified latent space. By leveraging the advanced capabilities of the Gemini architecture, this embedding model enhances performance across various downstream tasks, such as retrieval-augmented generation, semantic search, and complex reasoning. The research underscores Google's continued investment in developing versatile AI models capable of seamless multimodal understanding. This release is expected to influence how developers and researchers build integrated AI systems, providing more accurate and context-aware representations that bridge the gap between different data modalities in modern enterprise and research AI applications.

02

Thom_Wolf_CARBON DNA Model

Thom Wolf shared insights regarding the launch of CARBON, a groundbreaking 8 billion parameter open-source DNA model designed for genomic research. Developed to push the boundaries of biological data processing, the model features a massive 65,000-token context window, allowing for the analysis of the entire human genome on a single GPU in less than two days. This significant advancement in computational biology represents a milestone for open-source AI, offering researchers high-performance tools for complex DNA sequencing tasks without the need for extensive hardware clusters. The tweet highlights the integration of large-scale language modeling techniques into scientific domains, signaling a transformative shift toward efficient, specialized AI for genomics and life sciences, ultimately accelerating discovery in clinical language applications and beyond.

03

c_valenzuelab_AI Video Leap

Cristobal Valenzuela, the CEO of Runway, recently highlighted a significant breakthrough in the field of AI-generated video. The tweet emphasizes that current technological advancements have effectively navigated past the uncanny valley, a common hurdle where synthetic content appears disturbingly unnatural or inconsistent. By overcoming these aesthetic and structural limitations, AI video models are now capable of producing hyper-realistic, high-fidelity visual outputs that closely mirror authentic human cinematography. This development marks a pivotal shift in the creative industries, enabling creators to generate complex, motion-heavy content with unprecedented realism. As these tools continue to evolve, the distinction between computer-generated imagery and real-world footage becomes increasingly blurred, signaling a transformative era for digital storytelling, visual production workflows, and the broader creative economy fueled by sophisticated generative architectures.

04

LumaLabsAI_Dream Machine AI

Luma Labs AI has released a creative promotional campaign highlighting the versatile generative capabilities of their AI platform. The tweet showcases a series of imaginative scenarios featuring anthropomorphized animals—a fox, a walrus, and an otter—performing complex human tasks like soldiering, seafaring, and medical care. This content serves as a high-quality demonstration of the Luma Dream Machine's ability to interpret creative prompts and generate consistent, detailed visual content. By emphasizing that everyone has a unique calling, the company encourages users to explore their own creative potential through their video generation tools. This initiative reflects the broader industry trend of making sophisticated generative video technology accessible to creative professionals and casual users alike, marking a significant step in the development of tools that bridge the gap between abstract human imagination and high-fidelity digital visual output.

05

natolambert_Continual Learning

Nathan Lambert discusses the future trajectory of continual learning, positing that its primary application in the coming years will be within products designed for knowledge work. The tweet highlights a strategic shift where AI platforms—similar to how Cursor utilizes real-world data and reinforcement learning to iteratively improve its models—will likely integrate continuous training mechanisms into their workflows. By analyzing tools such as Claude and Copilot, Lambert suggests that the evolution of knowledge work will be defined by an AI's ability to learn from ongoing user interactions and real-time data feedback. This transition marks a critical trend in how large-scale models adapt to specific professional contexts, ultimately bridging the gap between static training datasets and the dynamic, iterative requirements of complex, knowledge-based productivity tasks in the modern digital enterprise ecosystem.

06

rasbt_MiniMax M2 LLM Insight

The tweet highlights the release of a new technical report concerning the MiniMax M2 series, which was previously recognized as a prominent open-weight Large Language Model. The author notes that the report provides valuable technical insights, particularly regarding the architecture's approach to attention mechanisms. A key technical takeaway mentioned is the model's experimental use of hybrid sliding-window attention, which the authors seemingly evaluated against the prevailing trend of utilizing full attention mechanisms. This analysis offers developers and researchers an interesting look into the design choices behind a significant LLM series, shedding light on the architectural trade-offs made by the team to balance performance and efficiency. These findings contribute to the ongoing industry discussion on optimizing attention models for better scalability and performance in high-compute natural language processing tasks.

huggingface

6 stories
01

EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation

The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community transitions towards Reinforcement Learning (RL) and agentic workflows. However, reliable evaluation has emerged as a critical bottleneck. Existing benchmarks predominantly evaluate ''whether it is right'' (basic prompt-following) while fundamentally neglecting ''whether it is good'' (cinematic quality, acting, and aesthetics). Furthermore, current automated metrics lack the domain-specific rigor required to provide trustworthy signals, creating a severe credibility gap between human aesthetic perception and machine scoring. To bridge this gap, we introduce EvalVerse, a comprehensive, pipeline-aware, and expert-calibrated evaluation framework. We treat video generation assessment not merely as an engineering task, but as a core scientific problem: the systematic digitization of subjective cinematic expertise. First, we organize domain knowledge into an evaluation taxonomy aligned with the professional filmmaking workflow (pre-production, production, and post-production). Second, we distill human expert judgments into a curated dataset with large-scale human annotations. Third, we inject this knowledge into Vision-Language Models (VLMs) through an expert-calibrated fine-tuning strategy, enabling the VLM to perform explicit Chain-of-Thought reasoning. Compared to previous works, EvalVerse not only retains compatibility with foundational ''rightness'' metrics, but also significantly expands the criteria to ''goodness'' and broaden the task coverage to complex multi-shot sequencing and audio-visual integration. Consequently, by providing granular diagnostic signals, EvalVerse transcends a static leaderboard and establishes a fundamental infrastructure for future work, such as reward models and evaluator agent.

02

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

We present MobileGym, a browser-hosted, lightweight, fully controllable environment for everyday mobile use, targeting interaction fidelity without replicating proprietary backends. It enables two capabilities previously out of reach for everyday apps: verifiable outcome signals through deterministic state-based judging over structured JSON state, and scalable online RL through low-cost parallel rollouts. The full environment state is captured, configured, forked, and compared as structured JSON, and a single server can host hundreds of parallel instances, with about 400 MB memory per instance and about 3 s cold start. A layered state model and a declarative task-definition framework keep state programmability and task creation practical at scale, and a single programmatic judging mechanism delivers both deterministic evaluation verdicts and dense RL rewards. The accompanying MobileGym-Bench provides 416 parameterized task templates, including 256 test and 160 train templates, over 28 apps, with deterministic judges and a structured AnswerSheet protocol that avoids free-text matching failures. In a Sim-to-Real case study, GRPO on Qwen3-VL-4B-Instruct gains +12.8 percentage points on the 256-task test set, and on a 59-task real-device signal subset, real-device execution retains 95.1% of the simulation-side training gain. Project page: https://mobilegym.github.io.

03

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.

04

Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling

Test-Time Scaling (TTS) enhances the reasoning capabilities of large language models by allocating additional inference compute to explore the solution space. However, existing parallel TTS methods typically keep branches isolated during search: intermediate discoveries remain branch-private and cannot guide other branches in time. This information isolation causes substantial redundant exploration, as branches repeatedly rediscover information already found elsewhere and require more search steps to collect complete decision information needed to reach correct answers. To bridge this gap, we propose Collaborative Parallel Thinking (CPT), a training-free inference framework that enables search-time information sharing across parallel branches. CPT extracts compact intermediate information from ongoing branches, maintains a deduplicated query-level information pool, and broadcasts pool entries through the input context, allowing each branch in subsequent search steps to reuse discoveries made by other branches rather than rediscover the same information. Empirically, experiments on HMMT and AIME benchmarks show that CPT establishes a stronger accuracy--latency Pareto frontier than strong baselines across rollout budgets and model scales, highlighting search-time collaboration as an effective direction for efficient parallel TTS.

05

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

We study series-level cinematic remaking, a long-horizon video-to-video generation problem that localizes full episodes or films via stylization or actor replacement while strictly preserving narrative structure, motion choreography, and character identity across hundreds of shots. Existing video generation and editing pipelines often break down in this regime due to compounding identity drift, background mutation, and semantic erosion under large camera motions and viewpoint changes. We propose Soap2Soap, a multi-agent framework that enforces long-term language-visual consistency through a Dual-Bridge Consistency mechanism: a scene-aware JSON screenplay serving as a persistent semantic backbone, and dynamically allocated visual reference anchors at both scene and shot levels. To suppress drift before video synthesis, we introduce batch keyframe consistency, jointly generating multiple keyframes in a shared latent context via a grid-based formulation. A closed-loop verification agent further audits identity, stability, and alignment to trigger selective regeneration. Experiments on SoapBench demonstrate strong improvements over commercial video generation APIs in long-term consistency and narrative fidelity.

06

Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancement

Agentic reinforcement learning (RL) has proven effective for training LLM-based agents with external tool-use capabilities. However, we identify that agentic RL training induces increasing redundant tool calls and blurs the model's intrinsic knowledge boundary, where the model fails to distinguish when tools are needed versus when parametric knowledge suffices. Existing solutions based on reward shaping create coarse-grained optimization targets that tend to incentivize indiscriminate tool-call suppression, leading to reward hacking. In this paper, we propose AKBE (Agentic Knowledge Boundary Enhancement), an on-policy method that dynamically probes the model's intrinsic knowledge boundary through dual-path (with-tool and no-tool) rollouts during training. We define the knowledge boundary as the per-instance determination of whether tools are required and the minimum tool calls necessary. By comparing correctness across paths, AKBE categorizes trajectories and constructs targeted supervisory signals that guide efficient tool-use patterns for each question. These signals are integrated seamlessly into the agentic RL training loop. Experiments on seven QA benchmarks demonstrate that AKBE improves task accuracy by +1.85 on average and reduces tool calls by 18% over standard agentic RL, yielding 25% higher tool productivity without any accuracy-efficiency trade-off. Further analysis suggests its plug-and-play compatibility across different RL algorithms and the mechanism of each signal category. Our code is available at https://github.com/CuSO4-Chen/AKBE.