NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-11-24ENGLISH EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Claude Opus 4.5

Anthropic is reportedly introducing "Claude Opus 4.5," representing a significant advancement in its series of large language models. This latest iteration is anticipated to deliver substantial improvements in core AI functionalities, including enhanced reasoning capabilities, more sophisticated contextual understanding, and increased efficiency for processing intricate prompts and tasks. Building upon the strong foundation of its predecessors, Claude Opus 4.5 aims to further elevate performance in critical areas such as complex problem-solving, nuanced conversational interactions, and high-quality content generation across various domains. The update underscores Anthropic's ongoing commitment to pushing the frontiers of responsible AI development, providing tools that meet the evolving demands of both enterprise applications and cutting-edge research. Detailed information regarding its specific new features, performance benchmarks, and potential use cases is expected to be outlined in official documentation, signaling a notable progression in the AI landscape.

02

The Bitter Lesson of LLM Extensions

The article delves into 'The Bitter Lesson' within the domain of Large Language Model (LLM) extensions, positing that the pursuit of complex, hand-engineered systems built atop LLMs often proves less effective than scaling up the core foundation models. It draws a historical parallel from artificial intelligence research, where general methods leveraging increased computation and data have consistently outperformed intricate, human-designed heuristics. Applied to LLMs, this suggests that while techniques such as Retrieval-Augmented Generation (RAG), external tool integration, and advanced prompting strategies aim to augment model capabilities or address limitations, the more potent long-term strategy might involve dedicating resources to developing larger, more robust base models. The core argument advocates for prioritizing fundamental LLM improvements over constructing elaborate external frameworks, reinforcing the idea that simplicity and computational scaling of the underlying model often yield superior results.

03

Claude Advanced Tool Use

Anthropic has unveiled significant advancements in the tool use capabilities of its Claude large language model, marking a crucial step towards enhancing AI agents' ability to interact with the real world. This development allows Claude to more effectively interpret complex instructions, engage with various external systems, and leverage a diverse set of tools and APIs to accomplish sophisticated tasks. The new functionalities empower Claude to autonomously plan and execute multi-step processes, breaking down intricate problems into manageable sub-tasks and orchestrating sequences of tool calls for information retrieval, data manipulation, or real-world interactions. These improvements are designed to increase Claude's agency and problem-solving prowess, enabling it to perform actions beyond its inherent linguistic generation abilities. This advancement is pivotal for developing more versatile and robust AI applications, further bridging the gap between language understanding and practical execution in diverse domains. It underscores a growing trend in AI research towards creating more capable and adaptable intelligent systems.

04

Launch HN: Karumi (YC F25) – Personalized, agentic product demos

Karumi, a YC F25 startup, introduces an innovative system for personalized, agentic product demonstrations, aiming to automate the sales and onboarding process. Developed by co-founders Toni and Pablo, Karumi enables users to receive instant, scalable, and guided product tours without any human intervention, supporting multiple languages. The platform utilizes an advanced AI agent that operates directly within a real web application in a shared browser session. This agent is designed to mimic human interaction by navigating the product interface, performing actions such as clicking buttons and filling out forms, while simultaneously providing real-time explanations of its actions to the user. This fully automated solution seeks to revolutionize how companies conduct product demos, offering a consistent, efficient, and on-demand method for showcasing product functionalities. It promises to significantly reduce operational overhead, enhance user engagement, and provide immediate, tailored insights into a product's capabilities, replacing traditional, resource-intensive human-led demonstrations.

05

Shopping research in ChatGPT

OpenAI has rolled out new functionalities within ChatGPT specifically tailored to streamline and enhance consumer shopping research. This innovative feature, as detailed in a ZDNet report, positions ChatGPT as a rapid, enjoyable, and no-cost resource for individuals aiming to compare products, synthesize user reviews, and accumulate thorough information prior to committing to a purchase. The objective is to simplify the often arduous and time-consuming process of online product investigation, enabling users to efficiently evaluate diverse options and pinpoint optimal solutions for their requirements. While this integration promises considerable convenience, it also sparks discussion regarding its capacity to truly rival human discerning skills and intuition in complex buying scenarios. The tool's accuracy, depth of analysis, and ability to handle nuanced consumer preferences are key areas of scrutiny. This expansion represents a notable advancement in applying large language models to practical, everyday consumer tasks, moving beyond mere conversational AI to become a functional research assistant.

06

What OpenAI did when ChatGPT users lost touch with reality

This report details OpenAI's proactive measures and strategic responses to instances where ChatGPT users reported experiencing AI hallucinations or a perceived disconnect from reality, a critical challenge as large language models become increasingly integrated into daily life. Addressing the complexities of model reliability and user perception is paramount for responsible AI development. OpenAI reportedly implemented a comprehensive series of interventions, including significant enhancements to model safety features, refining response generation algorithms to minimize misleading or unfactual outputs, and providing clearer user guidance on the inherent limitations of AI. These multifaceted efforts aim to bolster user trust and ensure the responsible deployment of generative AI. The strategies focus on improving the consistency, factual grounding, and contextual awareness of ChatGPT's interactions, thereby mitigating potential psychological impacts and reinforcing a transparent user experience. The article likely explores both the technical advancements and ethical considerations guiding these critical decisions, underscoring OpenAI's ongoing commitment to developing robust, safe, and trustworthy generative AI technologies.

GitHub

3 stories
01

TrendRadar

TrendRadar is a lightweight, easily deployable hotspot assistant designed to filter news and information, delivering relevant content to users rapidly. It aggregates trending topics from over 11 mainstream platforms, including Zhihu, Douyin, Weibo, and Wallstreetcn. Key features encompass intelligent push strategies with daily summaries, current topic rankings, and incremental monitoring modes, precise content filtering using custom keywords and advanced sorting, and real-time trend analysis. The system supports multi-channel real-time notifications via WeChat Work, Feishu, DingTalk, DingTalk, Telegram, Email, ntfy, and Bark. A significant addition in v3.0.0 is AI intelligent analysis, powered by the Model Context Protocol (MCP), which enables natural language querying and provides 13 analysis tools for deep data insights, topic trend tracking, and cross-platform comparisons. Deployable via GitHub Fork or Docker with zero technical barrier, TrendRadar helps users actively acquire desired information, suitable for investors, self-media creators, and public relations professionals.

02

Agent Development Kit (ADK) for Go

The Agent Development Kit (ADK) for Go is an open-source, code-first toolkit designed to streamline the building, evaluating, and deploying of sophisticated AI agents. Applying robust software development principles, ADK offers a flexible and modular framework for creating agent workflows, from simple automation to complex multi-agent systems. While optimized for Google's Gemini, the kit is model-agnostic and deployment-agnostic, ensuring compatibility across various AI models and environments. This Go version of ADK particularly leverages Go's strengths in concurrency and performance, making it ideal for cloud-native agent applications. Key features include an idiomatic Go design, a rich ecosystem for integrating pre-built or custom tools, code-first development for ultimate flexibility and testability, and strong support for containerization and deployment in cloud environments like Google Cloud Run. This empowers developers to create scalable and performant AI solutions with precise control over agent logic and orchestration.

03

➤ Cursor Free VIP

Cursor Free VIP is a comprehensive, cross-platform utility designed to enhance the experience of the Cursor AI-first integrated development environment (IDE). Supporting the latest Cursor versions, including 0.49.x, it offers compatibility across Windows (x64, x86), macOS (Intel, Apple Silicon), and Linux (x64, x86, ARM64) systems. Key features include the ability to reset Cursor's configuration and robust multi-language support, covering English, Simplified Chinese, Traditional Chinese, and Vietnamese. The project is explicitly positioned for educational and research purposes, with a strong disclaimer against generating fake email accounts or OAuth access and adherence to legal standards. It provides streamlined automated installation scripts for various operating systems (Linux/macOS with curl, Archlinux via AUR, Windows with PowerShell) and allows for detailed configuration of parameters such as browser paths, interaction timings, and update checks. This tool aims to optimize the use of Cursor IDE for development and learning, recommending administrator privileges for best performance and regular updates.

huggingface

6 stories
01

WorldGen: From Text to Traversable and Interactive 3D Worlds

We introduce WorldGen, a system that enables the automatic creation of large-scale, interactive 3D worlds directly from text prompts. Our approach transforms natural language descriptions into traversable, fully textured environments that can be immediately explored or edited within standard game engines. By combining LLM-driven scene layout reasoning, procedural generation, diffusion-based 3D generation, and object-aware scene decomposition, WorldGen bridges the gap between creative intent and functional virtual spaces, allowing creators to design coherent, navigable worlds without manual modeling or specialized 3D expertise. The system is fully modular and supports fine-grained control over layout, scale, and style, producing worlds that are geometrically consistent, visually rich, and efficient to render in real time. This work represents a step towards accessible, generative world-building at scale, advancing the frontier of 3D generative AI for applications in gaming, simulation, and immersive social environments.

02

O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents

Recent advancements in LLM-powered agents have demonstrated significant potential in generating human-like responses; however, they continue to face challenges in maintaining long-term interactions within complex environments, primarily due to limitations in contextual consistency and dynamic personalization. Existing memory systems often depend on semantic grouping prior to retrieval, which can overlook semantically irrelevant yet critical user information and introduce retrieval noise. In this report, we propose the initial design of O-Mem, a novel memory framework based on active user profiling that dynamically extracts and updates user characteristics and event records from their proactive interactions with agents. O-Mem supports hierarchical retrieval of persona attributes and topic-related context, enabling more adaptive and coherent personalized responses. O-Mem achieves 51.67% on the public LoCoMo benchmark, a nearly 3% improvement upon LangMem,the previous state-of-the-art, and it achieves 62.99% on PERSONAMEM, a 3.5% improvement upon A-Mem,the previous state-of-the-art. O-Mem also boosts token and interaction response time efficiency compared to previous memory frameworks. Our work opens up promising directions for developing efficient and human-like personalized AI assistants in the future.

03

Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models

The growing misuse of Vision-Language Models (VLMs) has led providers to deploy multiple safeguards, including alignment tuning, system prompts, and content moderation. However, the real-world robustness of these defenses against adversarial attacks remains underexplored. We introduce Multi-Faceted Attack (MFA), a framework that systematically exposes general safety vulnerabilities in leading defense-equipped VLMs such as GPT-4o, Gemini-Pro, and Llama-4. The core component of MFA is the Attention-Transfer Attack (ATA), which hides harmful instructions inside a meta task with competing objectives. We provide a theoretical perspective based on reward hacking to explain why this attack succeeds. To improve cross-model transferability, we further introduce a lightweight transfer-enhancement algorithm combined with a simple repetition strategy that jointly bypasses both input-level and output-level filters without model-specific fine-tuning. Empirically, we show that adversarial images optimized for one vision encoder transfer broadly to unseen VLMs, indicating that shared visual representations create a cross-model safety vulnerability. Overall, MFA achieves a 58.5% success rate and consistently outperforms existing methods. On state-of-the-art commercial models, MFA reaches a 52.8% success rate, surpassing the second-best attack by 34%. These results challenge the perceived robustness of current defense mechanisms and highlight persistent safety weaknesses in modern VLMs. Code: https://github.com/cure-lab/MultiFacetedAttack

04

Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story

Intrinsic dimension (ID) is an important tool in modern LLM analysis, informing studies of training dynamics, scaling behavior, and dataset structure, yet its textual determinants remain underexplored. We provide the first comprehensive study grounding ID in interpretable text properties through cross-encoder analysis, linguistic features, and sparse autoencoders (SAEs). In this work, we establish three key findings. First, ID is complementary to entropy-based metrics: after controlling for length, the two are uncorrelated, with ID capturing geometric complexity orthogonal to prediction quality. Second, ID exhibits robust genre stratification: scientific prose shows low ID (~8), encyclopedic content medium ID (~9), and creative/opinion writing high ID (~10.5) across all models tested. This reveals that contemporary LLMs find scientific text "representationally simple" while fiction requires additional degrees of freedom. Third, using SAEs, we identify causal features: scientific signals (formal tone, report templates, statistics) reduce ID; humanized signals (personalization, emotion, narrative) increase it. Steering experiments confirm these effects are causal. Thus, for contemporary models, scientific writing appears comparatively "easy", whereas fiction, opinion, and affect add representational degrees of freedom. Our multi-faceted analysis provides practical guidance for the proper use of ID and the sound interpretation of ID-based results.

05

OmniScientist: Toward a Co-evolving Ecosystem of Human and AI Scientists

With the rapid development of Large Language Models (LLMs), AI agents have demonstrated increasing proficiency in scientific tasks, ranging from hypothesis generation and experimental design to manuscript writing. Such agent systems are commonly referred to as "AI Scientists." However, existing AI Scientists predominantly formulate scientific discovery as a standalone search or optimization problem, overlooking the fact that scientific research is inherently a social and collaborative endeavor. Real-world science relies on a complex scientific infrastructure composed of collaborative mechanisms, contribution attribution, peer review, and structured scientific knowledge networks. Due to the lack of modeling for these critical dimensions, current systems struggle to establish a genuine research ecosystem or interact deeply with the human scientific community. To bridge this gap, we introduce OmniScientist, a framework that explicitly encodes the underlying mechanisms of human research into the AI scientific workflow. OmniScientist not only achieves end-to-end automation across data foundation, literature review, research ideation, experiment automation, scientific writing, and peer review, but also provides comprehensive infrastructural support by simulating the human scientific system, comprising: (1) a structured knowledge system built upon citation networks and conceptual correlations; (2) a collaborative research protocol (OSP), which enables seamless multi-agent collaboration and human researcher participation; and (3) an open evaluation platform (ScienceArena) based on blind pairwise user voting and Elo rankings. This infrastructure empowers agents to not only comprehend and leverage human knowledge systems but also to collaborate and co-evolve, fostering a sustainable and scalable innovation ecosystem.

06

Planning with Sketch-Guided Verification for Physics-Aware Video Generation

Recent video generation approaches increasingly rely on planning intermediate control signals such as object trajectories to improve temporal coherence and motion fidelity. However, these methods mostly employ single-shot plans that are typically limited to simple motions, or iterative refinement which requires multiple calls to the video generator, incuring high computational cost. To overcome these limitations, we propose SketchVerify, a training-free, sketch-verification-based planning framework that improves motion planning quality with more dynamically coherent trajectories (i.e., physically plausible and instruction-consistent motions) prior to full video generation by introducing a test-time sampling and verification loop. Given a prompt and a reference image, our method predicts multiple candidate motion plans and ranks them using a vision-language verifier that jointly evaluates semantic alignment with the instruction and physical plausibility. To efficiently score candidate motion plans, we render each trajectory as a lightweight video sketch by compositing objects over a static background, which bypasses the need for expensive, repeated diffusion-based synthesis while achieving comparable performance. We iteratively refine the motion plan until a satisfactory one is identified, which is then passed to the trajectory-conditioned generator for final synthesis. Experiments on WorldModelBench and PhyWorldBench demonstrate that our method significantly improves motion quality, physical realism, and long-term consistency compared to competitive baselines while being substantially more efficient. Our ablation study further shows that scaling up the number of trajectory candidates consistently enhances overall performance.