NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-11-19DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Building more with GPT-5.1-Codex-Max

OpenAI has introduced GPT-5.1-Codex-Max, a new, highly advanced iteration within its large language model series, specifically tailored to revolutionize code generation and software development workflows. This model, a significant leap from its predecessors, is engineered to empower developers by offering substantially enhanced capabilities in software creation, sophisticated debugging, and intricate code optimization. GPT-5.1-Codex-Max is anticipated to deliver superior contextual understanding, advanced problem-solving capacities for highly complex programming challenges, and exceptional efficiency in translating natural language prompts into high-quality, functional code across a diverse array of programming languages. Its deployment is set to profoundly impact the software development lifecycle by accelerating innovation, streamlining complex tasks, and reducing development cycles across various applications. The "Max" designation underlines its presumed superior performance, extensive scale, and potential to redefine the paradigm of AI-assisted code production and developer productivity.

02

Launch HN: Mosaic (YC W25) – Agentic Video Editing

Mosaic, a YC W25 startup, has launched its agentic video editing platform designed to revolutionize video production. Co-founded by Adish & Kyle, the platform distinguishes itself from traditional tools like DaVinci Resolve and Adobe Premiere Pro through its unique user interface and integrated visual intelligence. Mosaic allows users to create and deploy multimodal video editing agents within a node-based canvas, automating tedious tasks such as sifting through extensive raw footage. The founders' experience, stemming from the frustration of manually identifying specific objects like Cybertrucks in hours of video, highlights the platform's core value proposition: simplifying complex editing processes. By offering an agent-driven approach, Mosaic aims to make advanced video editing more efficient and accessible, overcoming the challenges of hidden features and cumbersome workflows common in existing software.

03

Multimodal Diffusion Language Models for Thinking-Aware Editing and Generation

The research presented in the GitHub repository 'MMaDA-Parallel' introduces a novel approach utilizing Multimodal Diffusion Language Models (MDLMs) for advanced 'Thinking-Aware Editing and Generation'. This initiative aims to bridge the gap between human cognitive processes and AI-driven content creation across various modalities. By integrating the generative power of diffusion models with the semantic understanding of language models, MDLMs are designed to comprehend and respond to intricate user intentions, moving beyond simple prompt-response mechanisms. The 'thinking-aware' aspect suggests an enhanced capacity for contextual reasoning, enabling more coherent, creative, and controllable output in editing existing content or generating new material. This technology promises to revolutionize fields requiring sophisticated multimodal content manipulation, offering tools that can interpret complex instructions and produce results that align closely with nuanced human thought processes, thereby advancing the capabilities of generative AI systems.

04

Show HN: Vibe Prolog

A developer has launched "Vibe Prolog," an experimental Prolog interpreter, after leveraging a $250 Claude Code credit. The project stands out due to its unique development process: it was predominantly "vibe coded" over a single weekend, largely on a mobile phone. This unconventional approach was motivated by the desire to fully utilize the Claude credit and simultaneously explore the boundaries of rapid, AI-assisted development in non-traditional environments. The creator views "Vibe Prolog" as an ongoing experiment to test the limits of what can be achieved with modern large language models and agile, mobile-first coding practices, even for complex system-level software like a programming language interpreter. This initiative offers an intriguing perspective on modern software engineering, showcasing the increasing utility of AI for quick prototyping and innovative development workflows beyond the conventional desktop setup.

05

The "Learned Helplessness" of AI

The concept of "Learned Helplessness" in AI explores how artificial intelligence systems might develop a state of passivity or failure to act, even when solutions are feasible, due to repeated negative experiences or an inability to generalize from their training. This phenomenon could manifest when AI models are consistently exposed to scenarios where their actions yield no positive outcomes or when they are trained on biased datasets that limit their problem-solving capabilities in novel situations. The article likely delves into the psychological parallels of learned helplessness, examining its implications for AI robustness, adaptability, and ethical development. It suggests a critical need to re-evaluate training methodologies and reward structures to foster more resilient and proactive AI systems, capable of overcoming perceived limitations and actively seeking solutions rather than defaulting to inaction.

06

Sam 3D: Powerful 3D Reconstruction for Physical World Images

Meta has introduced Sam 3D, a significant advancement in the field of 3D reconstruction tailored for real-world physical images. Building upon foundational work like the Segment Anything Model (SAM), Sam 3D aims to deliver robust and highly accurate 3D representations from diverse visual inputs. This technology addresses persistent challenges in generating precise three-dimensional models from complex, unconstrained environments, which often suffer from issues like occlusions, varied lighting conditions, and intricate object geometries. By potentially integrating advanced segmentation capabilities, Sam 3D could revolutionize how digital twins are created, enhancing applications across augmented and virtual reality, robotics, and content generation. The system's power lies in its ability to reconstruct detailed and coherent 3D structures, making it a crucial step towards more immersive digital experiences and advanced environmental understanding for AI systems. This development underscores Meta's ongoing commitment to pushing the boundaries of AI-driven computer vision and spatial computing.

GitHub

3 stories
01

TrendRadar

TrendRadar is an open-source, lightweight, and easily deployable hot topic assistant designed to help users efficiently consume news and information, avoiding information overload. It aggregates real-time hotspots from over 11 mainstream platforms including Zhihu, Douyin, Weibo, and Baidu, with support for custom platform additions. Key functionalities encompass intelligent push strategies like daily summaries, current榜单, and incremental monitoring, precise content filtering using custom keywords, and in-depth hotspot trend analysis. The system employs a personalized algorithm to re-sort global hot searches based on rank, frequency, and quality. It offers multi-channel real-time notifications via WeChat Work, Feishu, DingTalk, Telegram, Email, and ntfy. A major update introduced AI intelligent analysis based on the Model Context Protocol (MCP), enabling natural language queries and deep data insights through 13 analysis tools. With zero technical barrier deployment via GitHub Fork and Docker, TrendRadar is ideal for investors, self-media professionals, public relations, and general users seeking targeted information.

02

Agent Development Kit (ADK) for Go

The Agent Development Kit (ADK) for Go is an open-source, code-first toolkit designed to simplify the building, evaluation, and deployment of sophisticated AI agents, applying robust software development principles to agent creation. This flexible and modular framework facilitates the orchestration of agent workflows, from simple tasks to complex multi-agent systems. While optimized for Google's Gemini, ADK ensures model-agnosticism and deployment flexibility, allowing integration with various frameworks and cloud-native environments like Google Cloud Run. Leveraging Go's strengths in concurrency and performance, it is particularly suited for developers creating high-performance, scalable agent applications. Core features include an idiomatic Go design, a rich tool ecosystem for expanding agent capabilities, a code-first approach for superior flexibility and testability, and strong support for developing and deploying modular multi-agent systems efficiently.

03

➤ Cursor Free VIP

The "Cursor Free VIP" project is a utility tool designed to enhance the Cursor AI-first IDE experience by providing

huggingface

6 stories
01

Agent-R1: Training Powerful LLM Agents with End-to-End Reinforcement Learning

Large Language Models (LLMs) are increasingly being explored for building Agents capable of active environmental interaction (e.g., via tool use) to solve complex problems. Reinforcement Learning (RL) is considered a key technology with significant potential for training such Agents; however, the effective application of RL to LLM Agents is still in its nascent stages and faces considerable challenges. Currently, this emerging field lacks in-depth exploration into RL approaches specifically tailored for the LLM Agent context, alongside a scarcity of flexible and easily extensible training frameworks designed for this purpose. To help advance this area, this paper first revisits and clarifies Reinforcement Learning methodologies for LLM Agents by systematically extending the Markov Decision Process (MDP) framework to comprehensively define the key components of an LLM Agent. Secondly, we introduce Agent-R1, a modular, flexible, and user-friendly training framework for RL-based LLM Agents, designed for straightforward adaptation across diverse task scenarios and interactive environments. We conducted experiments on Multihop QA benchmark tasks, providing initial validation for the effectiveness of our proposed methods and framework.

02

Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark

While Chain-of-Thought (CoT) prompting enables sophisticated symbolic reasoning in LLMs, it remains confined to discrete text and cannot simulate the continuous, physics-governed dynamics of the real world. Recent video generation models have emerged as potential world simulators through Chain-of-Frames (CoF) reasoning -- materializing thought as frame-by-frame visual sequences, with each frame representing a physically-grounded reasoning step. Despite compelling demonstrations, a challenge persists: existing benchmarks, focusing on fidelity or alignment, do not assess CoF reasoning and thus cannot measure core cognitive abilities in multi-step planning, algorithmic logic, or abstract pattern extrapolation. This evaluation void prevents systematic understanding of model capabilities and principled guidance for improvement. We introduce Gen-ViRe (Generative Visual Reasoning Benchmark), a framework grounded in cognitive science and real-world AI applications, which decomposes CoF reasoning into six cognitive dimensions -- from perceptual logic to abstract planning -- and 24 subtasks. Through multi-source data curation, minimal prompting protocols, and hybrid VLM-assisted evaluation with detailed criteria, Gen-ViRe delivers the first quantitative assessment of video models as reasoners. Our experiments on SOTA systems reveal substantial discrepancies between impressive visual quality and actual reasoning depth, establishing baselines and diagnostic tools to advance genuine world simulators.

03

A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space

Innovative visual stylization is a cornerstone of artistic creation, yet generating novel and consistent visual styles remains a significant challenge. Existing generative approaches typically rely on lengthy textual prompts, reference images, or parameter-efficient fine-tuning to guide style-aware image generation, but often struggle with style consistency, limited creativity, and complex style representations. In this paper, we affirm that a style is worth one numerical code by introducing the novel task, code-to-style image generation, which produces images with novel, consistent visual styles conditioned solely on a numerical style code. To date, this field has only been primarily explored by the industry (e.g., Midjourney), with no open-source research from the academic community. To fill this gap, we propose CoTyle, the first open-source method for this task. Specifically, we first train a discrete style codebook from a collection of images to extract style embeddings. These embeddings serve as conditions for a text-to-image diffusion model (T2I-DM) to generate stylistic images. Subsequently, we train an autoregressive style generator on the discrete style embeddings to model their distribution, allowing the synthesis of novel style embeddings. During inference, a numerical style code is mapped to a unique style embedding by the style generator, and this embedding guides the T2I-DM to generate images in the corresponding style. Unlike existing methods, our method offers unparalleled simplicity and diversity, unlocking a vast space of reproducible styles from minimal input. Extensive experiments validate that CoTyle effectively turns a numerical code into a style controller, demonstrating a style is worth one code.

04

Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution

We introduce Orion, a visual agent framework that can take in any modality and generate any modality. Using an agentic framework with multiple tool-calling capabilities, Orion is designed for visual AI tasks and achieves state-of-the-art results. Unlike traditional vision-language models that produce descriptive outputs, Orion orchestrates a suite of specialized computer vision tools, including object detection, keypoint localization, panoptic segmentation, Optical Character Recognition, and geometric analysis, to execute complex multi-step visual workflows. The system achieves competitive performance on MMMU, MMBench, DocVQA, and MMLongBench while extending monolithic vision-language models to production-grade visual intelligence. By combining neural perception with symbolic execution, Orion enables autonomous visual reasoning, marking a transition from passive visual understanding to active, tool-driven visual intelligence.

05

LLM-Powered Fully Automated Chaos Engineering: Towards Enabling Anyone to Build Resilient Software Systems at Low Cost

Chaos Engineering (CE) is an engineering technique aimed at improving the resilience of distributed systems. It involves intentionally injecting faults into a system to test its resilience, uncover weaknesses, and address them before they cause failures in production. Recent CE tools automate the execution of predefined CE experiments. However, planning such experiments and improving the system based on the experimental results still remain manual. These processes are labor-intensive and require multi-domain expertise. To address these challenges and enable anyone to build resilient systems at low cost, this paper proposes ChaosEater, a system that automates the entire CE cycle with Large Language Models (LLMs). It predefines an agentic workflow according to a systematic CE cycle and assigns subdivided processes within the workflow to LLMs. ChaosEater targets CE for software systems built on Kubernetes. Therefore, the LLMs in ChaosEater complete CE cycles through software engineering tasks, including requirement definition, code generation, testing, and debugging. We evaluate ChaosEater through case studies on small- and large-scale Kubernetes systems. The results demonstrate that it consistently completes reasonable CE cycles with significantly low time and monetary costs. Its cycles are also qualitatively validated by human engineers and LLMs.

06

ATLAS: A High-Difficulty, Multidisciplinary Benchmark for Frontier Scientific Reasoning

The rapid advancement of Large Language Models (LLMs) has led to performance saturation on many established benchmarks, questioning their ability to distinguish frontier models. Concurrently, existing high-difficulty benchmarks often suffer from narrow disciplinary focus, oversimplified answer formats, and vulnerability to data contamination, creating a fidelity gap with real-world scientific inquiry. To address these challenges, we introduce ATLAS (AGI-Oriented Testbed for Logical Application in Science), a large-scale, high-difficulty, and cross-disciplinary evaluation suite composed of approximately 800 original problems. Developed by domain experts (PhD-level and above), ATLAS spans seven core scientific fields: mathematics, physics, chemistry, biology, computer science, earth science, and materials science. Its key features include: (1) High Originality and Contamination Resistance, with all questions newly created or substantially adapted to prevent test data leakage; (2) Cross-Disciplinary Focus, designed to assess models' ability to integrate knowledge and reason across scientific domains; (3) High-Fidelity Answers, prioritizing complex, open-ended answers involving multi-step reasoning and LaTeX-formatted expressions over simple multiple-choice questions; and (4) Rigorous Quality Control, employing a multi-stage process of expert peer review and adversarial testing to ensure question difficulty, scientific value, and correctness. We also propose a robust evaluation paradigm using a panel of LLM judges for automated, nuanced assessment of complex answers. Preliminary results on leading models demonstrate ATLAS's effectiveness in differentiating their advanced scientific reasoning capabilities. We plan to develop ATLAS into a long-term, open, community-driven platform to provide a reliable "ruler" for progress toward Artificial General Intelligence.