NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-06-01DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

CS336: Language Modeling from Scratch

CS336: Language Modeling from Scratch is a comprehensive academic course from Stanford University that focuses on building and understanding large language models from the ground up. The curriculum guides students through the fundamental mechanics of modern language modeling, detailing the mathematical foundations and architectural components such as the Transformer. Students gain hands-on experience by implementing key steps including tokenization, self-attention mechanisms, and optimization techniques. The course also explores training dynamics, scaling laws, and efficient distributed training methods necessary for handling massive datasets. Additionally, the curriculum addresses practical aspects of model evaluation and fine-tuning strategies like supervised fine-tuning and reinforcement learning from human feedback. By demystifying the entire pipeline from data preprocessing to deployment, CS336 provides a rigorous technical foundation for students and researchers aiming to master the underlying engineering principles of contemporary generative AI systems.

02

A 10 year old Xeon is all you need

This article explores the feasibility and performance of running modern large language models, specifically Google's Gemma, on highly accessible, decade-old hardware. By utilizing a 2016-era Intel Xeon processor, the author demonstrates that consumer-grade legacy server hardware remains surprisingly capable of executing local AI inference tasks. This approach challenges the prevailing industry narrative that modern generative artificial intelligence workloads strictly require expensive, cutting-edge GPU clusters. Through efficient software optimizations, quantized model weights, and optimized inference engines like llama.cpp, older CPU architectures can achieve usable token generation speeds. The findings highlight a highly cost-effective and sustainable alternative for hobbyists, researchers, and developers who wish to experiment with local AI models without investing in dedicated modern accelerators. This self-hosting setup underscores the democratization of AI technology, proving that software-level efficiency gains can breathe new life into older enterprise hardware for modern natural language processing applications.

03

Qwen3.7-Plus: Multimodal Agent Intelligence

Alibaba's Qwen team has introduced Qwen3.7-Plus, a highly advanced model optimized for multimodal agent intelligence. This model marks a significant leap forward by combining powerful visual, auditory, and textual comprehension with active decision-making capabilities. Designed specifically to power complex AI agents, Qwen3.7-Plus excels in dynamic environments where it must perceive rich multimodal inputs, reason over sophisticated tasks, and execute precise actions. By bridging the gap between perception and action, the model supports seamless tool integration, complex logical planning, and real-time environment interaction. Its launch highlights a growing industry shift toward agentic workflows and interactive systems that operate autonomously across diverse software interfaces, setting a new benchmark for multimodal LLM performance in practical, real-world agent applications.

04

AI Agent Guidelines for CS336 at Stanford

This document outlines the operational guidelines for integrating AI agents within the CS336 course curriculum at Stanford University. Focused on the practical implementation of AI assistants, the guidelines establish best practices for students and developers working on assignments, with a focus on code structure, testing, and formatting standards. The guidelines define explicit instructions for code compilation, linting, and systematic execution of test suites to ensure agent-generated code complies with academic rigor. By formalizing these interactions, the material provides a structured framework for using Large Language Model agents like Claude as co-pilots in deep learning educational settings. It highlights how educational institutions are adapting to automated coding workflows, emphasizing standardized environment configurations, strict testing protocols, and robust error-handling mechanisms in student-driven machine learning projects.

05

Build a Basic AI Agent from Scratch: Tools

This technical guide provides a comprehensive, step-by-step walkthrough for building a basic artificial intelligence agent from scratch, with a specific focus on implementing tool integration. It details how to equip a large language model with external capabilities, allowing the agent to interact with APIs, perform calculations, and fetch real-time data to solve complex tasks. By explaining the underlying mechanisms of tool definition, function calling, and execution loops, the article demystifies how modern AI systems transition from passive text generators to active problem-solving entities. Developers will learn how to design, write, and integrate custom tools into an agentic workflow, enhancing the model's overall utility and reasoning capabilities. Ultimately, the guide serves as a practical foundation for understanding agentic architectures, showing that robust AI agents can be constructed with minimal dependencies and a clear understanding of prompt engineering and structured JSON outputs.

06

Visa invests in Replit to power agentic payments for developers

Visa has announced a strategic investment in Replit, a prominent cloud-based software development platform, to accelerate the integration of transactional capabilities into AI development workflows. This partnership is aimed at pioneering the concept of 'agentic payments,' enabling autonomous AI agents designed by developers to execute financial transactions securely and seamlessly. By embedding Visa's massive payment network infrastructure directly into Replit's coding environment, the collaboration lowers the barrier to entry for software engineers creating financial applications. Developers will gain access to specialized tools and APIs to program intelligent agents capable of managing budgets, processing payments, and executing complex financial microtransactions autonomously. This milestone investment highlights the growing intersections of fintech and artificial intelligence, showcasing a future where AI systems are not only capable of generating code but are also legally and technically equipped to handle real-world economic exchanges safely.

Twitter

6 stories
01

fchollet_ARC-AGI-3 SOTA

The tweet highlights a significant development in the field of artificial intelligence, specifically regarding the ARC-AGI-3 benchmark. It reports that the Anthropic Opus 4.8 model has achieved a new state-of-the-art (SOTA) performance status on the ARC-AGI-3 task. With a benchmark score of 1.5% and a associated prize value of approximately $10,000, this milestone underscores the ongoing evolution of models capable of abstract reasoning and problem-solving in complex environments. Analysis notes indicate that Opus 4.8 demonstrated unique capabilities in reading and interpreting environmental parameters during the testing process. This achievement serves as a technical benchmark for researchers aiming to push the boundaries of general intelligence and algorithmic efficiency, reflecting the rigorous standards currently applied to evaluate advanced large-scale language and reasoning models in specialized academic and competitive AI environments.

02

ylecun_LLM vs JEPA Models

Yann LeCun highlights a fundamental architectural distinction in current artificial intelligence research, contrasting Large Language Models (LLMs) with Joint-Embedding Predictive Architectures (JEPA). The core argument centers on the mechanism of learning: whereas LLMs derive their intelligence through the statistical prediction of discrete tokens, models such as JEPA and data2vec operate by predicting abstract representations. This shift towards predicting latent space abstractions is proposed as a potential pathway toward achieving more robust world models. By bypassing token-level predictions in favor of conceptual abstraction, these architectures aim to build a more nuanced understanding of the environment. This technical comparison underscores the ongoing industry debate regarding the limitations of autoregressive scaling and the potential advantages of alternative, abstraction-based learning paradigms in the pursuit of advanced machine intelligence.

03

c_valenzuelab_Runway UK HQ

Runway, the leading generative AI video research and product company, has officially announced the establishment of its new European headquarters in the United Kingdom. This expansion is supported by a significant $200 million investment directed towards advanced world model artificial intelligence research and development. The initiative aims to bolster the regional AI ecosystem while creating numerous high-skilled technology jobs within the British market. By scaling its operational presence in Europe, Runway intends to accelerate the deployment of its innovative multimodal creative tools and advance its long-term strategic mission of building sophisticated world models. This announcement marks a pivotal moment for international AI infrastructure development, signaling Runway's commitment to global growth and the continued professionalization of AI research initiatives through collaborative efforts with regional partners and policymakers.

04

natolambert_Nvidia Open AI

Nvidia is playing a central role in the advancement of US-based open model initiatives, acting as a key driver for the technology's rapid evolution. The tweet highlights that the release of a massive 550B parameter model serves as a significant inflection point, compelling a broader audience to take notice of the industry's shift toward open-source accessibility. Furthermore, the author emphasizes that beyond just the computational power of the models, the release of high-quality, valuable training datasets is an often-overlooked but essential contribution to the ecosystem. This development marks a maturation of the open AI landscape, where infrastructure providers like Nvidia and model contributors are collaborating to push the boundaries of what is possible in the current open-source research and development cycle for large-scale generative architectures.

05

AndrewYNg_AI Job Trends

Andrew Ng discusses the rise of the Forward Deployed Engineer (FDE) role in AI, an embedded position popularized by Palantir that helps clients integrate LLMs into custom agentic workflows. While FDEs are gaining traction at companies like OpenAI and Anthropic, Ng argues that the broader demand remains with generalist AI Engineers who build software applications using LLM prompting and agentic frameworks. He notes that companies often prioritize vendor-neutral optionality over deep integration by single-vendor FDEs. Ng anticipates that the AI engineering field will eventually fragment into specialized roles such as LLMOps, Evals, and AI Data Engineering, similar to the historical evolution of traditional software engineering. Ultimately, he highlights the robust, growing demand for skilled professionals capable of navigating the maturing AI landscape, fostering long-term job growth across the technology sector.

06

LumaLabsAI_AI VFX Shift

LumaLabsAI recently highlighted a transformative shift in the visual effects industry through the integration of artificial intelligence. By retweeting DreamLabLA, the company emphasizes the evolution from manual pixel-based editing to a more intuitive process of directing desired outcomes. The technical demonstration showcases how modern AI models can automate complex compositing and rendering tasks, significantly streamlining production workflows. This transition suggests a profound impact on creative industries, where AI-driven tools perform intricate technical operations, allowing artists to focus on high-level creative vision. As generative AI continues to mature, its role in video production is becoming increasingly sophisticated, promising a future where traditional rendering bottlenecks are mitigated by intelligent automation. This advancement signifies a major milestone in the fusion of deep learning with professional-grade media creation, setting a new benchmark for efficiency and creative potential within visual content production.

huggingface

6 stories
01

Mellum2 Technical Report

We present Mellum 2, an open-weight 12B-parameter Mixture-of-Experts (MoE) language model with 2.5B active parameters per token. Mellum 2 is a general-purpose language model specialized in software engineering, spanning code generation and editing, debugging, multi-step reasoning, tool use and function calling, agentic coding, and conversational programming assistance, and it is the successor to the completion-focused 4B dense Mellum model. The architecture builds on the Mixture-of-Experts (64 experts, 8 active) and combines Grouped-Query Attention with 4 KV heads, Sliding Window Attention on three of every four layers, and a single Multi-Token Prediction head that doubles as both an auxiliary pre-training objective and a built-in draft model for speculative decoding; each choice was validated by ablation with inference efficiency on commodity GPUs as a design constraint. Pre-training spans approximately 10.6 trillion tokens through a three-phase curriculum that progressively shifts the mixture from diverse web data toward curated code and mathematical content, optimized with Muon under FP8 hybrid precision and a Warmup-Hold-Decay schedule with linear decay to zero. The pre-trained base is extended to a 128K context window via a layer-selective YaRN and then post-trained in two stages (supervised fine-tuning followed by RLVR), yielding two released variants: an Instruct model that answers directly and a Thinking model that emits an explicit reasoning trace before its final answer. Across code generation, math and reasoning, tool use, knowledge, and safety benchmarks, Mellum 2 is competitive with open-weight baselines in the 4B-14B range while running at the per-token compute of a 2.5B dense model. We release the base, instruct, and thinking checkpoints, together with this report on the architecture decisions, data pipeline, and training recipe behind them, under the Apache 2.0 license.

02

SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or frontier-model judges. We introduce SCOPE, a data-free self-play framework for open-ended tasks that co-evolves two policies: a Challenger that generates document-grounded tasks, and a Solver that answers them through multi-turn retrieval. A frozen copy of the initial model serves as the self-judge, which writes task-specific rubrics from the source document and grades Solver responses against them. Across three 7-8B instruction-tuned models (Qwen2.5, Qwen3, OLMo-3), SCOPE improves open-ended performance by up to +10.4 points on eight benchmarks and matches or exceeds GRPO_data trained on ~9K curated prompts. Although trained only on open-ended tasks, SCOPE also improves held-out short-form QA by up to +13.8 points on seven held-out benchmarks, surpassing GRPO_data on all three models. Ablations show that co-evolving the Challenger is necessary to keep tasks near the Solver's frontier, that gains arise from improvements in both retrieval and synthesis with the relative contribution varying by task, and that rubric generation quality is the bottleneck for self-judging.

03

DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory

Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains a key challenge. In this work, we move beyond explicit 3D memory and coarse frame-level implicit modeling, and propose a fine-grained, learnable, and scalable memory for consistent world generation. We first identify two fundamental limitations of na'fve learnable memory architectures in long-horizon extrapolation, namely computational inefficiency and attention dispersion. Through a systematic analysis of attention dispersion, we propose DecMem, a decoupled memory architecture that employs Sparse Global Memory for efficient fine-grained access to global history and Anchored Local Memory for stable and high-quality extrapolation. Extensive experiments demonstrate that DecMem significantly outperforms current state-of-the-art methods. By ensuring precise and efficient long-term memory and achieving superior extrapolation capabilities, DecMem enables minute-level controllable long video generation with high fidelity and consistency.

04

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these systems often suffer from a critical limitation in practice: agents fail to recognize their own knowledge boundaries, blindly triggering searches when internal knowledge suffices and failing to terminate search even when adequate evidence has been collected. The lack of self-awareness leads to severe over-search, incurring substantial inference latency and prohibitive computational cost. To this end, we propose SAAS, a novel RL framework designed to cultivate dynamic self-awareness that precisely regulates search behavior without compromising accuracy. SAAS introduces three key components: (i) a search boundary modeling mechanism, which identifies the search boundary under the evolving policy by contrasting search-disabled and search-enabled rollouts; (ii) a boundary-aware reward module, which translates this boundary awareness into trajectory-level penalties, suppressing unnecessary and redundant searches; and (iii) a stage-wise optimization strategy, which leverages a sequential curriculum to prioritize reasoning over search regularization, thereby avoiding reward hacking. Extensive experiments demonstrate that SAAS substantially reduces over-search, while maintaining accuracy. Our code is anonymously released at https://github.com/XMUDeepLIT/SAAS.

05

From Model Scaling to System Scaling: Scaling the Harness in Agentic AI

This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures around foundation models. We refer to this shift as scaling the harness: treating the structured execution layer around a foundation model as a first-class object of design, evaluation, and optimization. Although recent large language models enable agents to use tools, retrieve information, maintain memory, and execute long-horizon workflows, evaluation remains largely model-centric, often reducing agents to final-task success while treating memory, retrieval, tool use, orchestration, verification, and governance as secondary implementation details. This framing is increasingly inadequate because agent performance emerges from the interaction among the foundation model, memory substrate, context constructor, skill-routing layer, orchestration loop, and verification-and-governance layer. Together, these components form the agent harness, which translates model capability into long-horizon agent behavior. We study scaling the harness through three core bottlenecks: context governance, trustworthy memory, and dynamic skill routing, together with the orchestration and governance mechanisms that coordinate and constrain them. We further outline a research agenda for harness-level benchmarks that go beyond one-shot task success to measure trajectory quality, memory hygiene, context efficiency, communication fidelity, verification cost, and safe evolution over time. To make the discussion concrete, we develop CheetahClaws: https://github.com/SafeRL-Lab/cheetahclaws, a Python-native reference harness, and compare it with Claude Code and OpenClaw. Our main claim is that future progress in agentic AI will depend as much on system design as on stronger foundation models.

06

The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement

Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation or judge models. It is dramatically worse as the policy evolves beyond the static RM training. Therefore, we propose SAVE (Self-supervised reward model improvement via Value-Anchored On-policy feedback), a framework that grades on-policy responses as feedback by using the value function for on-policy RM training. SAVE naturally converts the reward-graded on-policy responses into supervision with a prompt-specific value head as an adaptive anchor. It computes RM advantages and filters ambiguous samples to update the RM via a contrastive objective. The effectiveness of SAVE for enhancing RM training is strongly validated through rigorous empirical evaluation across six diverse benchmarks. It achieves outperforming results across all datasets while maintaining consistent improvements across three RL algorithms (GRPO, RLOO, GSPO) and different policy backbones.