NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-02-04ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Code for Infrastructure

The Hacker News story, "Claude Code for Infrastructure," introduces Fluid.sh, a specialized platform designed to provide developer-friendly infrastructure explicitly for AI-first applications. Although the direct content provided is concise, the title strongly implies an innovative approach: utilizing advanced AI models, specifically Claude, to generate, manage, or optimize infrastructure code. This paradigm aims to significantly streamline the deployment and operational lifecycle of complex AI models by automating the typically manual and intricate processes of infrastructure provisioning and configuration. Fluid.sh positions itself as a crucial enabler for AI engineers, offering scalable, secure, and cost-effective infrastructure solutions, including essential computational resources such as GPUs, CPUs, and storage. The platform's objective is to mitigate the complexities associated with underlying operational challenges, thereby allowing AI developers to allocate more focus towards model development and less on infrastructure overhead. This initiative underscores a broader industry movement towards integrating AI capabilities into infrastructure management to enhance efficiency and accessibility for demanding AI workloads.

02

Claude Is a Space to Think

Anthropic's latest announcement positions its AI assistant, Claude, as 'a space to think,' emphasizing its role beyond simple query responses to foster deep engagement and complex cognitive tasks. This initiative highlights Claude's advanced capabilities in sustained reasoning, enabling users to explore intricate problems, develop ideas, and participate in reflective processes. The concept suggests an AI that supports iterative thought, critical analysis, and creative problem-solving, offering a collaborative platform where users can extend their intellectual capacity. By focusing on the quality of interaction and depth of processing, Anthropic aims to differentiate Claude as a sophisticated AI partner. It transcends basic conversational functions, providing a robust platform for intellectual exploration and innovation, making it an invaluable asset for professionals and researchers engaged in analytical and creative workflows.

03

Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation

A recent research paper titled 'Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation' introduces an innovative method to enhance the efficiency of attention mechanisms, a fundamental component in many advanced deep learning architectures. The study proposes a technique that ensures a constant computational cost per token, independent of the input sequence length. This is achieved by employing a novel symmetry-aware Taylor approximation. This approach directly addresses a critical limitation of traditional attention models, where computational complexity typically scales quadratically with sequence length, thereby restricting their scalability for processing very long sequences. By providing a more efficient and scalable alternative, this research holds the potential to significantly improve the performance and applicability of deep learning models in scenarios demanding extensive contextual understanding, making advanced AI models more resource-friendly and practical for real-world applications.

04

Intel will start making GPUs

Intel, a long-standing titan in the microprocessor industry, has officially declared its intention to commence manufacturing Graphics Processing Units (GPUs). This pivotal strategic decision marks a significant re-entry or deepened commitment by Intel into a hardware segment largely dominated by NVIDIA and AMD. The move is set to intensify competition within the semiconductor landscape, particularly impacting markets crucial for high-performance computing, data centers, and the burgeoning fields of artificial intelligence and machine learning. GPUs are indispensable accelerators for complex computational tasks, including deep learning model training, scientific simulations, and advanced graphics rendering. Intel's venture into this domain suggests a broader ambition to provide comprehensive computing solutions, leveraging its extensive manufacturing capabilities and architectural expertise. Industry observers anticipate this development could stimulate innovation across the board, potentially leading to more competitive pricing and diverse architectural choices for consumers and enterprises reliant on GPU-powered systems. This strategic pivot highlights the increasing importance of integrated and specialized computing hardware in the evolving technological ecosystem.

05

Voxtral Transcribe 2

Mistral AI has announced the launch of "Voxtral Transcribe 2," marking an evolution in its sophisticated speech-to-text technology offerings. This release, while concisely presented, signifies a likely advancement in the capabilities of Mistral AI's transcription service. Industry trends for such updates typically include notable improvements in transcription accuracy, particularly in noisy environments or with diverse accents, enhanced processing speed for real-time applications, and expanded support for a wider array of languages and dialects. Leveraging Mistral AI's foundational expertise in advanced AI models, Voxtral Transcribe 2 is expected to integrate cutting-edge deep learning techniques to deliver superior performance in converting spoken language into text. This upgrade is poised to benefit various sectors, from media and customer service to legal and healthcare, by providing highly precise and efficient automated transcription solutions crucial for data analysis, accessibility, and operational efficiency. The second iteration implies a refined and optimized solution, building upon previous successes to meet growing demands for high-quality audio processing.

06

Show HN: Ghidra MCP Server – 110 tools for AI-assisted reverse engineering

The Ghidra MCP Server, showcased as a new project, introduces a comprehensive suite of 110 specialized tools aimed at revolutionizing reverse engineering workflows with artificial intelligence assistance. Built to complement or integrate with the established open-source Ghidra reverse engineering framework, this server infrastructure is designed to automate and significantly enhance the analysis of complex binaries. By incorporating advanced AI and machine learning techniques, it promises to elevate the efficiency and precision of critical tasks such as vulnerability identification, intricate software behavior understanding, and sophisticated malware detection. The extensive array of tools within the Multi-Compiler Platform (MCP) Server is expected to span diverse functionalities, ranging from intelligent pattern recognition and anomaly detection within assembly code to automated deobfuscation and refined function signature identification. This robust platform empowers cybersecurity researchers, software analysts, and forensic experts with cutting-edge capabilities, enabling them to more effectively dissect and comprehend elaborate software systems. This project underscores the accelerating integration of AI into cybersecurity, signaling a transformative shift in the methodologies and tools available for addressing contemporary reverse engineering challenges.

huggingface

6 stories
01

No Global Plan in Chain-of-Thought: Uncover the Latent Planning Horizon of LLMs

This work stems from prior complementary observations on the dynamics of Chain-of-Thought (CoT): Large Language Models (LLMs) is shown latent planning of subsequent reasoning prior to CoT emergence, thereby diminishing the significance of explicit CoT; whereas CoT remains critical for tasks requiring multi-step reasoning. To deepen the understanding between LLM's internal states and its verbalized reasoning trajectories, we investigate the latent planning strength of LLMs, through our probing method, Tele-Lens, applying to hidden states across diverse task domains. Our empirical results indicate that LLMs exhibit a myopic horizon, primarily conducting incremental transitions without precise global planning. Leveraging this characteristic, we propose a hypothesis on enhancing uncertainty estimation of CoT, which we validate that a small subset of CoT positions can effectively represent the uncertainty of the entire path. We further underscore the significance of exploiting CoT dynamics, and demonstrate that automatic recognition of CoT bypass can be achieved without performance degradation. Our code, data and models are released at https://github.com/lxucs/tele-lens.

02

Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection

The quadratic complexity of attention remains the central bottleneck in long-context inference for large language models. Prior acceleration methods either sparsify the attention map with structured patterns or permanently evict tokens at specific layers, which can retain irrelevant tokens or rely on irreversible early decisions despite the layer-/head-wise dynamics of token importance. In this paper, we propose Token Sparse Attention, a lightweight and dynamic token-level sparsification mechanism that compresses per-head Q, K, V to a reduced token set during attention and then decompresses the output back to the original sequence, enabling token information to be reconsidered in subsequent layers. Furthermore, Token Sparse Attention exposes a new design point at the intersection of token selection and sparse attention. Our approach is fully compatible with dense attention implementations, including Flash Attention, and can be seamlessly composed with existing sparse attention kernels. Experimental results show that Token Sparse Attention consistently improves accuracy-latency trade-off, achieving up to times3.23 attention speedup at 128K context with less than 1% accuracy degradation. These results demonstrate that dynamic and interleaved token-level sparsification is a complementary and effective strategy for scalable long-context inference.

03

Decouple Searching from Training: Scaling Data Mixing via Model Merging for Large Language Model Pre-training

Determining an effective data mixture is a key factor in Large Language Model (LLM) pre-training, where models must balance general competence with proficiency on hard tasks such as math and code. However, identifying an optimal mixture remains an open challenge, as existing approaches either rely on unreliable tiny-scale proxy experiments or require prohibitively expensive large-scale exploration. To address this, we propose Decouple Searching from Training Mix (DeMix), a novel framework that leverages model merging to predict optimal data ratios. This paradigm decouples search from training costs, enabling evaluation of unlimited sampled mixtures without extra training burden and thus facilitating better mixture discovery through more search trials. Extensive experiments demonstrate that DeMix breaks the trade-off between sufficiency, accuracy and efficiency, obtaining the optimal mixture with higher benchmark performance at lower search cost. Additionally, we release the DeMix Corpora, a comprehensive 22T-token dataset comprising high-quality pre-training data with validated mixtures to facilitate open research. Our code and DeMix Corpora is available at https://github.com/Lucius-lsr/DeMix.

04

AOrchestra: Automating Sub-Agent Creation for Agentic Orchestration

Language agents have shown strong promise for task automation. Realizing this promise for increasingly complex, long-horizon tasks has driven the rise of a sub-agent-as-tools paradigm for multi-turn task solving. However, existing designs still lack a dynamic abstraction view of sub-agents, thereby hurting adaptability. We address this challenge with a unified, framework-agnostic agent abstraction that models any agent as a tuple Instruction, Context, Tools, Model. This tuple acts as a compositional recipe for capabilities, enabling the system to spawn specialized executors for each task on demand. Building on this abstraction, we introduce an agentic system AOrchestra, where the central orchestrator concretizes the tuple at each step: it curates task-relevant context, selects tools and models, and delegates execution via on-the-fly automatic agent creation. Such designs enable reducing human engineering efforts, and remain framework-agnostic with plug-and-play support for diverse agents as task executors. It also enables a controllable performance-cost trade-off, allowing the system to approach Pareto-efficient. Across three challenging benchmarks (GAIA, SWE-Bench, Terminal-Bench), AOrchestra achieves 16.28% relative improvement against the strongest baseline when paired with Gemini-3-Flash. The code is available at: https://github.com/FoundationAgents/AOrchestra

05

daVinci-Agency: Unlocking Long-Horizon Agency Data-Efficiently

While Large Language Models (LLMs) excel at short-term tasks, scaling them to long-horizon agentic workflows remains challenging. The core bottleneck lies in the scarcity of training data that captures authentic long-dependency structures and cross-stage evolutionary dynamics--existing synthesis methods either confine to single-feature scenarios constrained by model distribution, or incur prohibitive human annotation costs, failing to provide scalable, high-quality supervision. We address this by reconceptualizing data synthesis through the lens of real-world software evolution. Our key insight: Pull Request (PR) sequences naturally embody the supervision signals for long-horizon learning. They decompose complex objectives into verifiable submission units, maintain functional coherence across iterations, and encode authentic refinement patterns through bug-fix histories. Building on this, we propose daVinci-Agency, which systematically mines structured supervision from chain-of-PRs through three interlocking mechanisms: (1) progressive task decomposition via continuous commits, (2) long-term consistency enforcement through unified functional objectives, and (3) verifiable refinement from authentic bug-fix trajectories. Unlike synthetic trajectories that treat each step independently, daVinci-Agency's PR-grounded structure inherently preserves the causal dependencies and iterative refinements essential for teaching persistent goal-directed behavior and enables natural alignment with project-level, full-cycle task modeling. The resulting trajectories are substantial--averaging 85k tokens and 116 tool calls--yet remarkably data-efficient: fine-tuning GLM-4.6 on 239 daVinci-Agency samples yields broad improvements across benchmarks, notably achieving a 47% relative gain on Toolathlon. Beyond benchmark performance, our analysis confirms...

06

3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation

Existing methods for human motion control in video generation typically rely on either 2D poses or explicit 3D parametric models (e.g., SMPL) as control signals. However, 2D poses rigidly bind motion to the driving viewpoint, precluding novel-view synthesis. Explicit 3D models, though structurally informative, suffer from inherent inaccuracies (e.g., depth ambiguity and inaccurate dynamics) which, when used as a strong constraint, override the powerful intrinsic 3D awareness of large-scale video generators. In this work, we revisit motion control from a 3D-aware perspective, advocating for an implicit, view-agnostic motion representation that naturally aligns with the generator's spatial priors rather than depending on externally reconstructed constraints. We introduce 3DiMo, which jointly trains a motion encoder with a pretrained video generator to distill driving frames into compact, view-agnostic motion tokens, injected semantically via cross-attention. To foster 3D awareness, we train with view-rich supervision (i.e., single-view, multi-view, and moving-camera videos), forcing motion consistency across diverse viewpoints. Additionally, we use auxiliary geometric supervision that leverages SMPL only for early initialization and is annealed to zero, enabling the model to transition from external 3D guidance to learning genuine 3D spatial motion understanding from the data and the generator's priors. Experiments confirm that 3DiMo faithfully reproduces driving motions with flexible, text-driven camera control, significantly surpassing existing methods in both motion fidelity and visual quality.