NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2025-12-16DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

GPT Image 1.5

OpenAI has unveiled "GPT Image 1.5," signaling a notable progression in its multimodal artificial intelligence offerings. This new release, though detailed specifications are anticipated, is strongly indicative of an enhanced model designed to improve capabilities in image generation, comprehensive image understanding, or a more deeply integrated vision-language processing framework within the overarching GPT architecture. The introduction of "GPT Image 1.5" suggests a strategic evolution from prior models like DALL-E, aiming to deliver more sophisticated visual processing directly within a unified generative pre-trained transformer. This advancement is poised to enable more nuanced and contextually aware interactions across both textual and visual data modalities. The initiative underscores OpenAI's ongoing commitment to pushing the frontiers of generative AI, developing more versatile and robust tools for a wide array of creative, analytical, and practical applications that demand high-fidelity understanding and generation across different data types. This iteration holds the potential to significantly streamline and enrich user experiences in various domains, from advanced content creation to complex data interpretation.

02

SHARP, an approach to photorealistic view synthesis from a single image

SHARP introduces a novel approach to photorealistic view synthesis, enabling the generation of high-fidelity, novel views from just a single input image. This method significantly advances the field of computer vision by tackling the complex challenge of reconstructing detailed 3D scene information and rendering it from arbitrary viewpoints with remarkable realism. The technique likely leverages advanced neural rendering or implicit neural representation architectures to infer geometric and appearance properties that allow for consistent and convincing viewpoint changes. Its capability to produce photorealistic results from minimal input has broad implications for applications in virtual reality, augmented reality, content creation, and 3D modeling, reducing the need for extensive multi-view datasets or intricate 3D scans. This breakthrough demonstrates progress in synthesizing complex visual data, pushing the boundaries of what's achievable with single-image inference.

03

Show HN: Solving the ~95% legislative coverage gap using LLM's

Jacek, the solo founder of Lustra, has developed a digital public infrastructure to address the significant legislative coverage gap, where approximately 95% of legislation goes unnoticed due to the complexity of raw legal texts. The platform aims to provide insightful, unbiased information by utilizing Large Language Models (LLMs), specifically Vertex AI, to ingest and sterilize raw legislative bills (PDF/XML) from US and Polish APIs, stripping away political spin and presenting data in a strict JSON format. Lustra incorporates a "Civic Algorithm" that sorts content based on user votes, fostering community-driven engagement. Additionally, it features "Civic Projects," an incubator for citizen legislation where user-submitted drafts are vetted using AI scoring and displayed alongside government bills. The technical foundation of Lustra includes Flutter for a unified web and mobile frontend, backed by Firebase and Google Cloud Run for its backend infrastructure.

04

AI is wiping out entry-level tech jobs, leaving graduates stranded

The proliferation of artificial intelligence technologies is profoundly transforming the entry-level job market across the tech sector, creating significant hurdles for recent university graduates. A growing body of evidence suggests a marked decline in positions traditionally open to individuals new to the industry, as AI-driven automation solutions increasingly take over tasks that previously required human input. This paradigm shift is fostering a highly competitive environment, leaving many qualified graduates struggling to secure their initial professional roles despite their academic achievements and specialized training. The situation has prompted widespread concern regarding AI's enduring impact on employment dynamics, particularly for those entering the workforce. Industry analysts and educators are increasingly calling for a fundamental reassessment of current educational frameworks and professional development programs. The objective is to foster skills that empower future graduates to effectively collaborate with, rather than be displaced by, advanced AI systems, thereby adapting to the rapidly evolving demands of an AI-centric economy.

05

A2UI: A Protocol for Agent-Driven Interfaces

A2UI introduces a novel protocol designed to facilitate the development and deployment of agent-driven interfaces, aiming to standardize how autonomous agents interact with and control user interfaces. This initiative addresses the growing need for robust and interoperable communication mechanisms between AI agents and diverse front-end systems. The proposed protocol seeks to define a structured framework for agents to perceive, interpret, and manipulate interface elements, enabling more sophisticated and dynamic user experiences. By establishing common data formats and interaction patterns, A2UI intends to reduce integration complexities, foster innovation in agent-based applications, and promote a more unified ecosystem for AI-powered interactions. The core objective is to move beyond simple command-response systems to enable agents to actively manage and adapt interfaces, leading to more intuitive and context-aware human-computer interaction paradigms. This standardization effort is crucial for scaling the capabilities of AI agents across various platforms and applications, from smart assistants to complex operational dashboards, ensuring consistency and reliability in agent-orchestrated environments.

06

CC, a new AI productivity agent that connects your Gmail, Calendar and Drive

Google has unveiled CC, an innovative AI productivity agent designed to deeply integrate with and enhance user experiences across key Google Workspace applications: Gmail, Calendar, and Drive. This new agent aims to optimize daily workflows by intelligently managing communications, scheduling events, and organizing digital assets. Leveraging advanced artificial intelligence capabilities, CC is poised to automate routine administrative tasks, offer intelligent recommendations, and empower users to maintain greater command over their digital environments. The introduction of CC signifies a notable advancement in the development of more intuitive and proactive personal AI assistants within the Google ecosystem, prioritizing enhanced efficiency and reduced cognitive burden for individuals navigating extensive information and task loads. The agent's fundamental purpose centers on interpreting context and user intent to proactively facilitate various productivity-centric activities, representing an evolution in personalized AI support for both professional and personal spheres.

huggingface

6 stories
01

Memory in the Age of AI Agents

Memory has emerged, and will continue to remain, a core capability of foundation model-based agents. As research on agent memory rapidly expands and attracts unprecedented attention, the field has also become increasingly fragmented. Existing works that fall under the umbrella of agent memory often differ substantially in their motivations, implementations, and evaluation protocols, while the proliferation of loosely defined memory terminologies has further obscured conceptual clarity. Traditional taxonomies such as long/short-term memory have proven insufficient to capture the diversity of contemporary agent memory systems. This work aims to provide an up-to-date landscape of current agent memory research. We begin by clearly delineating the scope of agent memory and distinguishing it from related concepts such as LLM memory, retrieval augmented generation (RAG), and context engineering. We then examine agent memory through the unified lenses of forms, functions, and dynamics. From the perspective of forms, we identify three dominant realizations of agent memory, namely token-level, parametric, and latent memory. From the perspective of functions, we propose a finer-grained taxonomy that distinguishes factual, experiential, and working memory. From the perspective of dynamics, we analyze how memory is formed, evolved, and retrieved over time. To support practical development, we compile a comprehensive summary of memory benchmarks and open-source frameworks. Beyond consolidation, we articulate a forward-looking perspective on emerging research frontiers, including memory automation, reinforcement learning integration, multimodal memory, multi-agent memory, and trustworthiness issues. We hope this survey serves not only as a reference for existing work, but also as a conceptual foundation for rethinking memory as a first-class primitive in the design of future agentic intelligence.

02

QwenLong-L1.5: Post-Training Recipe for Long-Context Reasoning and Memory Management

We introduce QwenLong-L1.5, a model that achieves superior long-context reasoning capabilities through systematic post-training innovations. The key technical breakthroughs of QwenLong-L1.5 are as follows: (1) Long-Context Data Synthesis Pipeline: We develop a systematic synthesis framework that generates challenging reasoning tasks requiring multi-hop grounding over globally distributed evidence. By deconstructing documents into atomic facts and their underlying relationships, and then programmatically composing verifiable reasoning questions, our approach creates high-quality training data at scale, moving substantially beyond simple retrieval tasks to enable genuine long-range reasoning capabilities. (2) Stabilized Reinforcement Learning for Long-Context Training: To overcome the critical instability in long-context RL, we introduce task-balanced sampling with task-specific advantage estimation to mitigate reward bias, and propose Adaptive Entropy-Controlled Policy Optimization (AEPO) that dynamically regulates exploration-exploitation trade-offs. (3) Memory-Augmented Architecture for Ultra-Long Contexts: Recognizing that even extended context windows cannot accommodate arbitrarily long sequences, we develop a memory management framework with multi-stage fusion RL training that seamlessly integrates single-pass reasoning with iterative memory-based processing for tasks exceeding 4M tokens. Based on Qwen3-30B-A3B-Thinking, QwenLong-L1.5 achieves performance comparable to GPT-5 and Gemini-2.5-Pro on long-context reasoning benchmarks, surpassing its baseline by 9.90 points on average. On ultra-long tasks (1M~4M tokens), QwenLong-L1.5's memory-agent framework yields a 9.48-point gain over the agent baseline. Additionally, the acquired long-context reasoning ability translates to enhanced performance in general domains like scientific reasoning, memory tool using, and extended dialogue.

03

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability, long-term visual quality, and temporal consistency. To this end, we take a progressive approach-first enhancing controllability and then extending toward long-term, high-quality generation. We present LongVie 2, an end-to-end autoregressive framework trained in three stages: (1) Multi-modal guidance, which integrates dense and sparse control signals to provide implicit world-level supervision and improve controllability; (2) Degradation-aware training on the input frame, bridging the gap between training and long-term inference to maintain high visual quality; and (3) History-context guidance, which aligns contextual information across adjacent clips to ensure temporal consistency. We further introduce LongVGenBench, a comprehensive benchmark comprising 100 high-resolution one-minute videos covering diverse real-world and synthetic environments. Extensive experiments demonstrate that LongVie 2 achieves state-of-the-art performance in long-range controllability, temporal coherence, and visual fidelity, and supports continuous video generation lasting up to five minutes, marking a significant step toward unified video world modeling.

04

WebOperator: Action-Aware Tree Search for Autonomous Agents in Web Environment

LLM-based agents often operate in a greedy, step-by-step manner, selecting actions solely based on the current observation without considering long-term consequences or alternative paths. This lack of foresight is particularly problematic in web environments, which are only partially observable-limited to browser-visible content (e.g., DOM and UI elements)-where a single misstep often requires complex and brittle navigation to undo. Without an explicit backtracking mechanism, agents struggle to correct errors or systematically explore alternative paths. Tree-search methods provide a principled framework for such structured exploration, but existing approaches lack mechanisms for safe backtracking, making them prone to unintended side effects. They also assume that all actions are reversible, ignoring the presence of irreversible actions-limitations that reduce their effectiveness in realistic web tasks. To address these challenges, we introduce WebOperator, a tree-search framework that enables reliable backtracking and strategic exploration. Our method incorporates a best-first search strategy that ranks actions by both reward estimates and safety considerations, along with a robust backtracking mechanism that verifies the feasibility of previously visited paths before replaying them, preventing unintended side effects. To further guide exploration, WebOperator generates action candidates from multiple, varied reasoning contexts to ensure diverse and robust exploration, and subsequently curates a high-quality action set by filtering out invalid actions pre-execution and merging semantically equivalent ones. Experimental results on WebArena and WebVoyager demonstrate the effectiveness of WebOperator. On WebArena, WebOperator achieves a state-of-the-art 54.6% success rate with gpt-4o, underscoring the critical advantage of integrating strategic foresight with safe execution.

05

Aesthetic Alignment Risks Assimilation: How Image Generation and Reward Models Reinforce Beauty Bias and Ideological "Censorship"

Over-aligning image generation models to a generalized aesthetic preference conflicts with user intent, particularly when "anti-aesthetic" outputs are requested for artistic or critical purposes. This adherence prioritizes developer-centered values, compromising user autonomy and aesthetic pluralism. We test this bias by constructing a wide-spectrum aesthetics dataset and evaluating state-of-the-art generation and reward models. We find that aesthetic-aligned generation models frequently default to conventionally beautiful outputs, failing to respect instructions for low-quality or negative imagery. Crucially, reward models penalize anti-aesthetic images even when they perfectly match the explicit user prompt. We confirm this systemic bias through image-to-image editing and evaluation against real abstract artworks.

06

Rethinking Expert Trajectory Utilization in LLM Post-training

While effective post-training integrates Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL), the optimal mechanism for utilizing expert trajectories remains unresolved. We propose the Plasticity-Ceiling Framework to theoretically ground this landscape, decomposing performance into foundational SFT performance and the subsequent RL plasticity. Through extensive benchmarking, we establish the Sequential SFT-then-RL pipeline as the superior standard, overcoming the stability deficits of synchronized approaches. Furthermore, we derive precise scaling guidelines: (1) Transitioning to RL at the SFT Stable or Mild Overfitting Sub-phase maximizes the final ceiling by securing foundational SFT performance without compromising RL plasticity; (2) Refuting ``Less is More'' in the context of SFT-then-RL scaling, we demonstrate that Data Scale determines the primary post-training potential, while Trajectory Difficulty acts as a performance multiplier; and (3) Identifying that the Minimum SFT Validation Loss serves as a robust indicator for selecting the expert trajectories that maximize the final performance ceiling. Our findings provide actionable guidelines for maximizing the value extracted from expert trajectories.