NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-01-13ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Signal leaders warn agentic AI is an insecure, unreliable surveillance risk

Signal's leadership, including its President and VP, has issued a significant warning regarding the inherent risks associated with agentic AI systems, explicitly labeling them as a "surveillance nightmare." They contend that these advanced AI agents are fundamentally insecure and unreliable, posing substantial threats to user privacy and data integrity. The primary concern stems from the autonomous nature of agentic AI, which can operate with limited human oversight, creating potential vulnerabilities ripe for exploitation. This could lead to extensive surveillance capabilities and data breaches, compromising sensitive information on a massive scale. Furthermore, Signal leaders highlight the unreliability of these systems, noting that their complex and often opaque decision-making processes render them unpredictable and challenging to control effectively in real-world scenarios. This cautionary stance from a prominent privacy-focused organization underscores a growing apprehension within the technology community concerning the ethical, security, and practical implications of rapidly deploying highly autonomous AI technologies without implementing comprehensive privacy-by-design principles and robust security safeguards.

02

AI Generated Music Barred from Bandcamp

A significant policy shift has been observed on Bandcamp, the popular online music store and community platform, which has reportedly moved to prohibit music generated by artificial intelligence. This development, though succinctly stated, signals a critical juncture for both generative AI artists and digital music platforms. The underlying reasons for Bandcamp's decision are likely multifaceted, encompassing concerns over copyright ownership, the potential for market saturation with AI-produced content, and the desire to uphold the value and authenticity of human-created art. Such a stance reflects a broader industry-wide struggle to define the role and limitations of AI in creative fields, particularly regarding fair compensation for human artists and ethical considerations of authorship. This move by Bandcamp is poised to spark further debate within the music community, influencing how other platforms might approach AI-generated content and challenging the future trajectory of independent music distribution in the age of artificial intelligence. It underscores the ongoing need for clear guidelines and discussions surrounding intellectual property in the rapidly evolving landscape of AI creativity.

03

Confer – End to end encrypted AI chat

Confer, a new initiative spearheaded by Signal creator Moxie Marlinspike, aims to introduce end-to-end encryption to AI chat services, drawing parallels to his pioneering work in secure messaging. The project's core proposition is to ensure user privacy in AI interactions, tackling the inherent data privacy challenges associated with artificial intelligence models. Marlinspike's vision extends to developing a paradigm where AI processes, particularly inference, can occur without compromising the confidentiality of user inputs. This approach is elaborated upon in their "Private Inference" blog post, suggesting a technical methodology for secure AI computations. By focusing on end-to-end encryption, Confer seeks to establish a new standard for privacy in the evolving landscape of AI applications, allowing users to engage with AI systems confidently, knowing their conversations and data remain private and protected from unauthorized access, effectively extending the security principles of secure messaging to the realm of artificial intelligence.

04

Instagram AI Influencers Are Defaming Celebrities with Sex Scandals

The emergence of AI-powered virtual influencers on platforms like Instagram has introduced novel ethical and legal challenges, particularly concerning the potential for defamation against real-world public figures. This report highlights instances where these AI entities are allegedly being used to spread false narratives, including fabricated sex scandals, targeting celebrities. The proliferation of such AI-generated content raises significant concerns about digital integrity, the spread of misinformation, and the protection of individuals' reputations in the age of advanced synthetic media. The technology enabling realistic deepfakes and AI-driven content creation makes it increasingly difficult for audiences to discern authentic information from fabricated narratives, posing a serious threat to celebrity image and public trust. This trend underscores an urgent need for robust platform policies, advanced detection mechanisms for AI-generated defamation, and legal frameworks to address the misuse of AI in digital spaces. The issue necessitates a broader discussion on the accountability of AI creators and platform operators in mitigating harm caused by synthetic content.

05

Apple Creator Studio

Apple has unveiled plans for the forthcoming "Apple Creator Studio," an innovative and comprehensive collection of creative applications designed to empower a wide spectrum of creators, including artists, designers, musicians, and filmmakers. This new suite, as suggested by initial reports, is anticipated to offer deep integration across Apple's expansive ecosystem, spanning macOS, iOS, and iPadOS devices. While granular details regarding specific features are yet to be fully disclosed, the "Creator Studio" is expected to incorporate advanced functionalities for various creative disciplines. It will likely leverage cutting-edge artificial intelligence and machine learning technologies to streamline workflows, enhance creative output, and unlock new possibilities for digital content creation. This strategic initiative underscores Apple's sustained commitment to delivering industry-leading software solutions, catering to both professional and aspiring creators by fostering a more intuitive, powerful, and inspiring creative environment. The platform aims to provide a unified hub for accessing sophisticated tools for video editing, graphic design, audio production, and potentially immersive media development, thereby solidifying Apple's standing as a premier platform for digital creativity.

06

FOSS in times of war, scarcity and (adversarial) AI [video]

This FOSDEM talk explores the critical intersection of Free and Open Source Software (FOSS) with contemporary global challenges, specifically focusing on the impacts of war, economic scarcity, and the rise of adversarial artificial intelligence. The presentation likely delves into how FOSS ecosystems can maintain resilience and foster innovation amidst geopolitical instability and resource limitations. Key discussion points would include the vulnerabilities and strengths of open-source projects when confronted with state-sponsored attacks or malicious AI applications designed to exploit software weaknesses or manipulate information. The talk is expected to provide insights into ethical considerations for AI development within FOSS frameworks, strategies for enhancing software supply chain security, and the role of the open-source community in building robust, trustworthy digital infrastructure. Furthermore, it aims to examine the strategic importance of FOSS in ensuring technological sovereignty and access to essential tools during crises, emphasizing the need for proactive measures to safeguard open-source principles against emerging threats, including those posed by sophisticated AI systems.

huggingface

6 stories
01

MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era

The rapid development of interactive and autonomous AI systems signals our entry into the agentic era. Training and evaluating agents on complex agentic tasks such as software engineering and computer use requires not only efficient model computation but also sophisticated infrastructure capable of coordinating vast agent-environment interactions. However, no open-source infrastructure can effectively support large-scale training and evaluation on such complex agentic tasks. To address this challenge, we present MegaFlow, a large-scale distributed orchestration system that enables efficient scheduling, resource allocation, and fine-grained task management for agent-environment workloads. MegaFlow abstracts agent training infrastructure into three independent services (Model Service, Agent Service, and Environment Service) that interact through unified interfaces, enabling independent scaling and flexible resource allocation across diverse agent-environment configurations. In our agent training deployments, MegaFlow successfully orchestrates tens of thousands of concurrent agent tasks while maintaining high system stability and achieving efficient resource utilization. By enabling such large-scale agent training, MegaFlow addresses a critical infrastructure gap in the emerging agentic AI landscape.

02

Lost in the Noise: How Reasoning Models Fail with Contextual Distractors

Recent advances in reasoning models and agentic AI systems have led to an increased reliance on diverse external information. However, this shift introduces input contexts that are inherently noisy, a reality that current sanitized benchmarks fail to capture. We introduce NoisyBench, a comprehensive benchmark that systematically evaluates model robustness across 11 datasets in RAG, reasoning, alignment, and tool-use tasks against diverse noise types, including random documents, irrelevant chat histories, and hard negative distractors. Our evaluation reveals a catastrophic performance drop of up to 80% in state-of-the-art models when faced with contextual distractors. Crucially, we find that agentic workflows often amplify these errors by over-trusting noisy tool outputs, and distractors can trigger emergent misalignment even without adversarial intent. We find that prompting, context engineering, SFT, and outcome-reward only RL fail to ensure robustness; in contrast, our proposed Rationale-Aware Reward (RARE) significantly strengthens resilience by incentivizing the identification of helpful information within noise. Finally, we uncover an inverse scaling trend where increased test-time computation leads to worse performance in noisy settings and demonstrate via attention visualization that models disproportionately focus on distractor tokens, providing vital insights for building the next generation of robust, reasoning-capable agents.

03

DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving, this vision gives rise to driving world models: generative simulators that imagine ego and agent futures, enabling scalable simulation, safe testing of corner cases, and rich synthetic data generation. Yet, despite fast-growing research activity, the field lacks a rigorous benchmark to measure progress and guide priorities. Existing evaluations remain limited: generic video metrics overlook safety-critical imaging factors; trajectory plausibility is rarely quantified; temporal and agent-level consistency is neglected; and controllability with respect to ego conditioning is ignored. Moreover, current datasets fail to cover the diversity of conditions required for real-world deployment. To address these gaps, we present DrivingGen, the first comprehensive benchmark for generative driving world models. DrivingGen combines a diverse evaluation dataset curated from both driving datasets and internet-scale video sources, spanning varied weather, time of day, geographic regions, and complex maneuvers, with a suite of new metrics that jointly assess visual realism, trajectory plausibility, temporal coherence, and controllability. Benchmarking 14 state-of-the-art models reveals clear trade-offs: general models look better but break physics, while driving-specific ones capture motion realistically but lag in visual quality. DrivingGen offers a unified evaluation framework to foster reliable, controllable, and deployable driving world models, enabling scalable simulation, planning, and data-driven decision-making.

04

PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning

We introduce Parallel Coordinated Reasoning (PaCoRe), a training-and-inference framework designed to overcome a central limitation of contemporary language models: their inability to scale test-time compute (TTC) far beyond sequential reasoning under a fixed context window. PaCoRe departs from the traditional sequential paradigm by driving TTC through massive parallel exploration coordinated via a message-passing architecture in multiple rounds. Each round launches many parallel reasoning trajectories, compacts their findings into context-bounded messages, and synthesizes these messages to guide the next round and ultimately produce the final answer. Trained end-to-end with large-scale, outcome-based reinforcement learning, the model masters the synthesis abilities required by PaCoRe and scales to multi-million-token effective TTC without exceeding context limits. The approach yields strong improvements across diverse domains, and notably pushes reasoning beyond frontier systems in mathematics: an 8B model reaches 94.5% on HMMT 2025, surpassing GPT-5's 93.2% by scaling effective TTC to roughly two million tokens. We open-source model checkpoints, training data, and the full inference pipeline to accelerate follow-up work.

05

X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests

Competitive programming presents great challenges for Code LLMs due to its intensive reasoning demands and high logical complexity. However, current Code LLMs still rely heavily on real-world data, which limits their scalability. In this paper, we explore a fully synthetic approach: training Code LLMs with entirely generated tasks, solutions, and test cases, to empower code reasoning models without relying on real-world data. To support this, we leverage feature-based synthesis to propose a novel data synthesis pipeline called SynthSmith. SynthSmith shows strong potential in producing diverse and challenging tasks, along with verified solutions and tests, supporting both supervised fine-tuning and reinforcement learning. Based on the proposed synthetic SFT and RL datasets, we introduce the X-Coder model series, which achieves a notable pass rate of 62.9 avg@8 on LiveCodeBench v5 and 55.8 on v6, outperforming DeepCoder-14B-Preview and AReal-boba2-14B despite having only 7B parameters. In-depth analysis reveals that scaling laws hold on our synthetic dataset, and we explore which dimensions are more effective to scale. We further provide insights into code-centric reinforcement learning and highlight the key factors that shape performance through detailed ablations and analysis. Our findings demonstrate that scaling high-quality synthetic data and adopting staged training can greatly advance code reasoning, while mitigating reliance on real-world coding data.

06

OS-Symphony: A Holistic Framework for Robust and Generalist Computer-Using Agent

While Vision-Language Models (VLMs) have significantly advanced Computer-Using Agents (CUAs), current frameworks struggle with robustness in long-horizon workflows and generalization in novel domains. These limitations stem from a lack of granular control over historical visual context curation and the absence of visual-aware tutorial retrieval. To bridge these gaps, we introduce OS-Symphony, a holistic framework that comprises an Orchestrator coordinating two key innovations for robust automation: (1) a Reflection-Memory Agent that utilizes milestone-driven long-term memory to enable trajectory-level self-correction, effectively mitigating visual context loss in long-horizon tasks; (2) Versatile Tool Agents featuring a Multimodal Searcher that adopts a SeeAct paradigm to navigate a browser-based sandbox to synthesize live, visually aligned tutorials, thereby resolving fidelity issues in unseen scenarios. Experimental results demonstrate that OS-Symphony delivers substantial performance gains across varying model scales, establishing new state-of-the-art results on three online benchmarks, notably achieving 65.84% on OSWorld.