NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-10-27ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude for Excel

Anthropic has introduced 'Claude for Excel,' an integration designed to bring advanced artificial intelligence capabilities directly into Microsoft Excel. This new offering aims to significantly enhance productivity and streamline data management tasks for users by leveraging Claude's powerful natural language processing and sophisticated reasoning abilities. The integration allows users to automate complex spreadsheet operations, generate accurate formulas, analyze intricate data patterns, and interpret information using intuitive natural language prompts, effectively transforming how data is handled and processed within Excel. This development signifies a strategic move towards embedding sophisticated AI agents into ubiquitous business applications, making AI-driven data analysis more accessible and intuitive for a broader audience. By substantially reducing manual effort and providing intelligent, context-aware insights, Claude for Excel seeks to empower users to derive greater value from their datasets, accelerate informed decision-making processes, and improve overall efficiency in data-intensive workflows.

02

Artificial Writing and Automated Detection

This research paper, titled 'Artificial Writing and Automated Detection,' critically examines the dual emergence of sophisticated AI-driven text generation and the imperative for effective automated detection systems. The proliferation of large language models has enabled the creation of highly coherent and contextually relevant artificial writing, blurring the lines between human and machine authorship. This development introduces considerable challenges across sectors such as education, journalism, and scientific research, where the authenticity of content is paramount. The paper likely investigates the technical mechanisms behind both the generation of such artificial text and the various methodologies employed for its identification, including statistical analysis, linguistic patterns, and potentially watermarking techniques. It highlights the dynamic and often adversarial relationship between advancements in generative AI and the development of robust detection tools. The work provides insights into the societal implications of widespread AI-generated content and the ongoing efforts to maintain transparency and integrity in written communication.

03

Creating an All-Weather Driver

Waymo's blog post, "Creating an All-Weather Driver," details the complex engineering challenges and innovative solutions involved in developing a robust autonomous driving system capable of operating safely and reliably across diverse and challenging weather conditions. The article likely explores how Waymo addresses critical issues such as obscured sensor data from rain, snow, or fog, and the impact of reduced visibility on perception and prediction systems. Key technical strategies could include advanced sensor fusion techniques, which integrate data from lidar, radar, cameras, and other sensors to create a comprehensive environmental model less susceptible to individual sensor limitations. Furthermore, the discussion would encompass sophisticated machine learning algorithms for object detection and tracking, specifically trained on vast datasets encompassing various meteorological scenarios. The ultimate goal is to ensure that Waymo's autonomous vehicles can maintain consistent performance and safety standards, irrespective of environmental factors, thereby expanding the operational design domain and bringing fully autonomous mobility closer to widespread reality. This initiative underscores the significant progress in AI and robotics required to achieve truly all-weather self-driving capabilities.

04

It's insulting to read AI-generated blog posts

The sentiment expressed in the Hacker News story title, 'It's insulting to read AI-generated blog posts,' reflects a growing critical stance among readers towards content perceived as originating from artificial intelligence. This perspective suggests that such material often lacks the depth, nuance, and genuine human insight expected from quality writing, leading to a negative user experience. Readers may interpret AI-generated posts as a form of intellectual laziness or a disrespect for their time and intelligence, fostering a sense of insult. The proliferation of easily identifiable AI-written content could erode trust in online publications and content creators, prompting a reevaluation of content generation strategies. This highlights a crucial challenge for Generative AI applications in writing: balancing efficiency with maintaining authenticity, originality, and a human connection to avoid alienating audiences. The critique underscores the importance of human oversight and unique value proposition in an increasingly automated content landscape, advocating for AI as an augmentative tool rather than a wholesale replacement for human creativity.

05

WorldGrow: Generating Infinite 3D World

WorldGrow introduces a novel approach to generating infinite 3D worlds, offering a robust framework for creating expansive and diverse virtual environments. The project focuses on procedural content generation techniques, enabling the dynamic creation of landscapes, structures, and interactive elements without manual intervention. This system aims to address challenges in developing vast virtual spaces for applications such as gaming, simulations, and virtual reality, where traditional asset creation can be time-consuming and resource-intensive. By leveraging algorithmic methods, WorldGrow facilitates the on-the-fly generation of unique and evolving environments, promising significant advancements in scalability and immersion. Its core idea revolves around automating the world-building process, providing developers with tools to rapidly prototype and deploy complex 3D scenes. The potential impact spans various digital domains requiring dynamic and endless spatial experiences.

06

Microsoft needs to open up more about its OpenAI dealings

The Wall Street Journal article highlights a critical need for Microsoft to disclose more details about its substantial dealings and strategic partnership with OpenAI. This demand for increased transparency emerges amid growing public and regulatory scrutiny over the extensive influence Microsoft exerts within the rapidly expanding artificial intelligence industry. Concerns are particularly focused on the implications of Microsoft's significant financial investment and operational integration with OpenAI, a frontrunner in AI research and deployment. Industry observers and critics are increasingly vocal about potential impacts on competitive fairness, the risk of monopolistic tendencies, and the broader challenges associated with governing a highly integrated relationship between two pivotal players in the AI domain. Enhanced transparency regarding their specific agreements, operational protocols, and joint strategic planning is considered vital for building public confidence, ensuring robust accountability, and fostering an equitable and innovative environment for AI development and application. This openness is essential to address issues of market concentration and to guide the responsible evolution of advanced AI technologies.

GitHub

2 stories
01

Agent Lightning⚡

Agent Lightning is a Microsoft-developed training framework designed to optimize AI agents with minimal code changes. It supports a wide array of agent frameworks, including LangChain, OpenAI Agent SDK, AutoGen, and CrewAI, or can be used with agents built without specific frameworks. A key feature is its ability to selectively optimize agents within complex multi-agent systems, leveraging algorithms such as Reinforcement Learning, Automatic Prompt Optimization, and Supervised Fine-tuning. The framework operates by tracing agent interactions —prompts, tool calls, and rewards —into a central LightningStore. This data fuels algorithms that refine agent behavior, producing updated resources like improved prompt templates or policy weights. The Trainer component orchestrates the learning loop, streaming data and applying improvements to the inference engine. This streamlined architecture enables a clear path from initial deployment to continuous performance enhancement, facilitating the development of robust and adaptable AI agents, as demonstrated in projects tackling long-horizon, sparse-reward tasks.

02

AFFiNE.ProWrite, Draw and Plan All at Once

AFFiNE is an open-source, local-first, and privacy-focused all-in-one workspace designed as a robust alternative to Notion and Miro. It uniquely hyper-merges writing, drawing, and planning functionalities into a single platform, enabling users to manage docs, canvases, and tables seamlessly. A core feature is its "true canvas for blocks," allowing the integration of diverse elements like rich text, sticky notes, embedded web pages, and multi-view databases onto an edgeless canvas. The platform is augmented by AFFiNE AI, a multimodal partner capable of assisting with tasks from generating professional reports and slides to summarizing articles into mindmaps and prototyping applications from prompts, leveraging advanced Canvas AI capabilities. Emphasizing data ownership, AFFiNE supports local storage alongside real-time collaboration and cross-platform synchronization. Users have the freedom to self-host, fork, and customize their AFFiNE instance, with an upcoming plugin community. Built upon open-source projects like Blocksuite and OctoBase, AFFiNE aims to provide a comprehensive and extensible knowledge management system.

huggingface

6 stories
01

DeepAgent: A General Reasoning Agent with Scalable Toolsets

Large reasoning models have demonstrated strong problem-solving abilities, yet real-world tasks often require external tools and long-horizon interactions. Existing agent frameworks typically follow predefined workflows, which limit autonomous and global task completion. In this paper, we introduce DeepAgent, an end-to-end deep reasoning agent that performs autonomous thinking, tool discovery, and action execution within a single, coherent reasoning process. To address the challenges of long-horizon interactions, particularly the context length explosion from multiple tool calls and the accumulation of interaction history, we introduce an autonomous memory folding mechanism that compresses past interactions into structured episodic, working, and tool memories, reducing error accumulation while preserving critical information. To teach general-purpose tool use efficiently and stably, we develop an end-to-end reinforcement learning strategy, namely ToolPO, that leverages LLM-simulated APIs and applies tool-call advantage attribution to assign fine-grained credit to the tool invocation tokens. Extensive experiments on eight benchmarks, including general tool-use tasks (ToolBench, API-Bank, TMDB, Spotify, ToolHop) and downstream applications (ALFWorld, WebShop, GAIA, HLE), demonstrate that DeepAgent consistently outperforms baselines across both labeled-tool and open-set tool retrieval scenarios. This work takes a step toward more general and capable agents for real-world applications. The code and demo are available at https://github.com/RUC-NLPIR/DeepAgent.

02

Video-As-Prompt: Unified Semantic Control for Video Generation

Unified, generalizable semantic control in video generation remains a critical open challenge. Existing methods either introduce artifacts by enforcing inappropriate pixel-wise priors from structure-based controls, or rely on non-generalizable, condition-specific finetuning or task-specific architectures. We introduce Video-As-Prompt (VAP), a new paradigm that reframes this problem as in-context generation. VAP leverages a reference video as a direct semantic prompt, guiding a frozen Video Diffusion Transformer (DiT) via a plug-and-play Mixture-of-Transformers (MoT) expert. This architecture prevents catastrophic forgetting and is guided by a temporally biased position embedding that eliminates spurious mapping priors for robust context retrieval. To power this approach and catalyze future research, we built VAP-Data, the largest dataset for semantic-controlled video generation with over 100K paired videos across 100 semantic conditions. As a single unified model, VAP sets a new state-of-the-art for open-source methods, achieving a 38.7% user preference rate that rivals leading condition-specific commercial models. VAP's strong zero-shot generalization and support for various downstream applications mark a significant advance toward general-purpose, controllable video generation.

03

From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model

Discrete diffusion models have emerged as a promising direction for vision-language tasks, offering bidirectional context modeling and theoretical parallelization. However, their practical application is severely hindered by a train-inference discrepancy, which leads to catastrophic error cascades: initial token errors during parallel decoding pollute the generation context, triggering a chain reaction of compounding errors and leading to syntactic errors and semantic hallucinations. To address this fundamental challenge, we reframe the generation process from passive denoising to active refining. We introduce ReDiff, a refining-enhanced diffusion framework that teaches the model to identify and correct its own errors. Our approach features a two-stage training process: first, we instill a foundational revision capability by training the model to revise synthetic errors; second, we implement a novel online self-correction loop where the model is explicitly trained to revise its own flawed drafts by learning from an expert's corrections. This mistake-driven learning endows the model with the crucial ability to revisit and refine its already generated output, effectively breaking the error cascade. Extensive experiments demonstrate that ReDiff significantly improves the coherence and factual accuracy of generated content, enabling stable and efficient parallel generation far superior to traditional denoising methods. Our codes and models are available at https://rediff-hku.github.io/.

04

A Definition of AGI

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quantifiable framework to address this, defining AGI as matching the cognitive versatility and proficiency of a well-educated adult. To operationalize this, we ground our methodology in Cattell-Horn-Carroll theory, the most empirically validated model of human cognition. The framework dissects general intelligence into ten core cognitive domains-including reasoning, memory, and perception-and adapts established human psychometric batteries to evaluate AI systems. Application of this framework reveals a highly "jagged" cognitive profile in contemporary models. While proficient in knowledge-intensive domains, current AI systems have critical deficits in foundational cognitive machinery, particularly long-term memory storage. The resulting AGI scores (e.g., GPT-4 at 27%, GPT-5 at 58%) concretely quantify both rapid progress and the substantial gap remaining before AGI.

05

UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning

GUI grounding, which maps natural-language instructions to actionable UI elements, is a core capability of GUI agents. Prior works largely treats instructions as a static proxy for user intent, overlooking the impact of instruction diversity and quality on grounding performance. Through a careful investigation of existing grounding datasets, we find a 23.3% flaw rate in their instructions and show that inference-time exploitation of instruction diversity yields up to a substantial 76% relative performance improvement. In this paper, we introduce the Instruction-as-Reasoning paradigm, treating instructions as dynamic analytical pathways that offer distinct perspectives and enabling the model to select the most effective pathway during reasoning. To achieve this, we propose a two-stage training framework: supervised fine-tuning (SFT) on synthesized, diverse instructions to instill multi-perspective reasoning, followed by reinforcement learning (RL) to optimize pathway selection and composition. Our resulting models, UI-Ins-7B and UI-Ins-32B, achieve state-of-the-art results on five challenging grounding benchmarks and exhibit emergent reasoning, selectively composing and synthesizing novel instruction pathways at inference. In particular, UI-Ins-32B attains the best grounding accuracy, scoring 87.3% on UI-I2E-Bench, 57.0% on ScreenSpot-Pro, and 84.9% on MMBench-GUI L2. Furthermore, our model demonstrates strong agentic potential, achieving a 74.1% success rate on AndroidWorld using UI-Ins-7B as the executor. Our in-depth analysis reveals additional insights such as how reasoning can be formulated to enhance rather than hinder grounding performance, and how our method mitigates policy collapse in the SFT+RL framework. All code and model checkpoints will be publicly released in https://github.com/alibaba/UI-Ins.

06

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling

Prompt design plays a crucial role in text-to-video (T2V) generation, yet user-provided prompts are often short, unstructured, and misaligned with training data, limiting the generative potential of diffusion-based T2V models. We present RAPO++, a cross-stage prompt optimization framework that unifies training-data--aligned refinement, test-time iterative scaling, and large language model (LLM) fine-tuning to substantially improve T2V generation without modifying the underlying generative backbone. In Stage 1, Retrieval-Augmented Prompt Optimization (RAPO) enriches user prompts with semantically relevant modifiers retrieved from a relation graph and refactors them to match training distributions, enhancing compositionality and multi-object fidelity. Stage 2 introduces Sample-Specific Prompt Optimization (SSPO), a closed-loop mechanism that iteratively refines prompts using multi-source feedback -- including semantic alignment, spatial fidelity, temporal coherence, and task-specific signals such as optical flow -- yielding progressively improved video generation quality. Stage 3 leverages optimized prompt pairs from SSPO to fine-tune the rewriter LLM, internalizing task-specific optimization patterns and enabling efficient, high-quality prompt generation even before inference. Extensive experiments across five state-of-the-art T2V models and five benchmarks demonstrate that RAPO++ achieves significant gains in semantic alignment, compositional reasoning, temporal stability, and physical plausibility, outperforming existing methods by large margins. Our results highlight RAPO++ as a model-agnostic, cost-efficient, and scalable solution that sets a new standard for prompt optimization in T2V generation. The code is available at https://github.com/Vchitect/RAPO.