NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-03-03中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

GPT‑5.3 Instant

OpenAI has reportedly introduced a new iteration in its foundational language model series, dubbed 'GPT-5.3 Instant'. While specific technical details remain sparse from the initial announcement, the nomenclature 'Instant' strongly suggests a strategic focus on enhancing the model's speed, efficiency, and real-time processing capabilities. This potential development indicates OpenAI's ongoing efforts to push the boundaries of AI performance, aiming for quicker response times and more seamless integration into latency-sensitive applications. Such advancements would be critical for improving user experience in conversational AI, automated content generation, and various interactive AI services where rapid output is paramount. The launch of GPT-5.3 Instant could signify a significant step towards more agile and responsive AI systems, further solidifying OpenAI's position in the competitive landscape of large language model development and deployment. This release underscores the continuous innovation within the field, prioritizing not just scale and capability, but also practical operational efficiency for broader application across industries.

02

Launch HN: Cekura (YC F24) Testing and monitoring for voice and chat AI agents

Cekura (YC F24), founded by Tarush, Sidhant, and Shashij, introduces a new platform for testing and monitoring voice and chat AI agents. The company addresses the significant challenge of manually quality assuring AI agents, which is impractical given the myriad ways users can interact with them and the frequent updates to prompts, models, or tools. Cekura's solution involves simulation, leveraging synthetic users to interact with AI agents in a manner akin to real-world conversations. This infrastructure, initially developed for voice agent simulation over 1.5 years, has now been extended to chat AI. The platform helps teams stress-test prompts and LLM behavior, ensuring agents respond correctly and identifying regressions before they impact production environments. By using LLM-based judges to evaluate agent responses, Cekura provides a scalable and robust alternative to traditional manual spot-checking, which is often inefficient or too late. This approach helps maintain agent performance and reliability as they evolve.

03

India's top court angry after junior judge cites fake AI-generated orders

India's Supreme Court has expressed strong disapproval following an incident where a junior judge reportedly cited fabricated AI-generated orders in judicial proceedings. This event highlights growing concerns within the legal community regarding the responsible and ethical integration of artificial intelligence tools, particularly large language models, into the judiciary. The reliance on AI for legal research and document preparation without robust verification mechanisms poses significant risks to the integrity of court processes and the accuracy of legal judgments. The incident underscores the critical need for comprehensive guidelines, training, and strict oversight for legal professionals utilizing AI, to prevent the propagation of erroneous or entirely fictitious information. It also brings into focus the challenges posed by AI's potential for 'hallucinations' and the imperative for human cross-verification in high-stakes environments like the legal system. The Supreme Court's reaction signals a potential shift towards more stringent regulations and a cautious approach to AI adoption within India's judicial framework, emphasizing accountability and the preservation of public trust in legal institutions.

04

AI-generated art can’t be copyrighted after Supreme Court declines review

The Supreme Court has declined to review a pivotal case concerning the copyrightability of art exclusively generated by artificial intelligence, thereby solidifying the legal stance that such creations cannot be copyrighted under existing U.S. law. This decision upholds lower court rulings which consistently emphasized the requirement of human authorship for any work to qualify for copyright protection. The legal battle originated from a computer scientist's attempt to register an AI-created image, challenging the traditional interpretation of intellectual property rights in an era of advanced generative AI. The Supreme Court's refusal to intervene sends a clear signal to the burgeoning AI art community and the broader tech industry: current copyright statutes are designed for human creators. This development is crucial for artists utilizing AI tools, developers of generative models, and legal experts, as it defines the boundaries of ownership and protection for works where human creative input is deemed insufficient or absent. It underscores the ongoing challenge of adapting established legal frameworks to rapid technological advancements, particularly in the realm of automated content creation.

05

When AI writes the software, who verifies it?

The increasing adoption of artificial intelligence, particularly advanced large language models, in automating software development presents a significant and pressing challenge: ensuring the verification and reliability of AI-generated code. As AI systems evolve to craft complex software solutions, traditional human-centric quality assurance and testing methodologies are rapidly becoming insufficient. This paradigm shift demands a profound reconsideration of how software integrity, security vulnerabilities, and functional correctness are guaranteed throughout the development lifecycle. The core dilemma focuses on establishing accountability for the quality of code produced by autonomous AI, considering its potential to introduce nuanced errors, inherent biases, or even security exploits. The future may see AI systems themselves contributing to the verification process, creating a sophisticated interdependency that requires new frameworks. This evolving landscape necessitates the development of novel AI-powered testing tools, updated regulatory standards, and a reimagined role for human developers, shifting focus towards strategic oversight, architectural validation, and ethical compliance rather than granular code review.

06

Ars Technica fires reporter after AI controversy involving fabricated quotes

Ars Technica, a prominent technology news publication, has terminated one of its reporters following a significant controversy centered on the use of artificial intelligence to generate fabricated quotes. This incident highlights critical challenges at the intersection of journalism and advanced AI technologies, particularly concerning the ethical implications of content creation. The event underscores the growing concerns regarding journalistic integrity and the necessity for rigorous ethical practices in an era where AI tools possess the capability to easily create believable but false information. The decision to fire the reporter sends a strong message about accountability and the imperative for strict editorial oversight and robust verification processes, especially when artificial intelligence is integrated into journalistic workflows. This controversy serves as a stark reminder of the potential for AI misuse in professional environments and the crucial need for responsible AI deployment to prevent the propagation of misinformation and maintain public trust in reporting. The incident is likely to prompt broader discussions within the media industry about establishing clear guidelines for AI tool usage and fostering a culture of transparency regarding AI-assisted content.

huggingface

6 stories
01

OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens

OmniLottie is a versatile framework that generates high quality vector animations from multi-modal instructions. For flexible motion and visual content control, we focus on Lottie, a light weight JSON formatting for both shapes and animation behaviors representation. However, the raw Lottie JSON files contain extensive invariant structural metadata and formatting tokens, posing significant challenges for learning vector animation generation. Therefore, we introduce a well designed Lottie tokenizer that transforms JSON files into structured sequences of commands and parameters representing shapes, animation functions and control parameters. Such tokenizer enables us to build OmniLottie upon pretrained vision language models to follow multi-modal interleaved instructions and generate high quality vector animations. To further advance research in vector animation generation, we curate MMLottie-2M, a large scale dataset of professionally designed vector animations paired with textual and visual annotations. With extensive experiments, we validate that OmniLottie can produce vivid and semantically aligned vector animations that adhere closely to multi modal human instructions.

02

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

Software engineering agents (SWE) are improving rapidly, with recent gains largely driven by reinforcement learning (RL). However, RL training is constrained by the scarcity of large-scale task collections with reproducible execution environments and reliable test suites. Although a growing number of benchmarks have emerged, datasets suitable for training remain limited in scale and diversity or often target a limited set of high-resource language ecosystems. We introduce SWE-rebench V2, a language-agnostic automated pipeline for harvesting executable real-world SWE tasks and constructing RL training environments at scale. The pipeline synthesizes repository-specific installation and test procedures via an interactive setup agent, and filters unsound instances using an ensemble of LLM judges, validated against human-verified SWE-bench annotations. Using this pipeline, we construct a dataset of 32,000+ tasks spanning 20 languages and 3,600+ repositories, with pre-built images for reproducible execution. To further scale training data, we additionally release 120,000+ tasks with installation instructions, fail-to-pass tests and rich metadata, where the problem statement is generated based on the original pull request description. We validate the collected instances through a diagnostic study that covers a subset of tasks in five programming languages across seven popular models, and provide instance-level metadata that flags common confounders such as overly restrictive tests and underspecified descriptions. We release the datasets, the collection and execution code, and associated artifacts to enable large-scale training of SWE agents across diverse languages and repositories.

03

WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories

Recent advances in foundational Video Diffusion Models (VDMs) have yielded significant progress. Yet, despite the remarkable visual quality of generated videos, reconstructing consistent 3D scenes from these outputs remains challenging, due to limited camera controllability and inconsistent generated content when viewed from distinct camera trajectories. In this paper, we propose WorldStereo, a novel framework that bridges camera-guided video generation and 3D reconstruction via two dedicated geometric memory modules. Formally, the global-geometric memory enables precise camera control while injecting coarse structural priors through incrementally updated point clouds. Moreover, the spatial-stereo memory constrains the model's attention receptive fields with 3D correspondence to focus on fine-grained details from the memory bank. These components enable WorldStereo to generate multi-view-consistent videos under precise camera control, facilitating high-quality 3D reconstruction. Furthermore, the flexible control branch-based WorldStereo shows impressive efficiency, benefiting from the distribution matching distilled VDM backbone without joint training. Extensive experiments across both camera-guided video generation and 3D reconstruction benchmarks demonstrate the effectiveness of our approach. Notably, we show that WorldStereo acts as a powerful world model, tackling diverse scene generation tasks (whether starting from perspective or panoramic images) with high-fidelity 3D results. Models will be released.

04

CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning

Large Language Models (LLMs) have recently exhibited remarkable reasoning capabilities, largely enabled by supervised fine-tuning (SFT)- and reinforcement learning (RL)-based post-training on high-quality reasoning data. However, reproducing and extending these capabilities in open and scalable settings is hindered by three fundamental data-centric challenges: (1) the cold-start problem, arising from the lack of seed datasets with detailed, long Chain-of-Thought (CoT) trajectories needed to initialize reasoning policies; (2) limited domain coverage, as most existing open-source reasoning datasets are concentrated in mathematics, with limited coverage of broader scientific disciplines; and (3) the annotation bottleneck, where the difficulty of frontier-level reasoning tasks makes reliable human annotation prohibitively expensive or infeasible. To address these challenges, we introduce CHIMERA, a compact synthetic reasoning dataset comprising 9K samples for generalizable cross-domain reasoning. CHIMERA is constructed with three key properties: (1) it provides rich, long CoT reasoning trajectories synthesized by state-of-the-art reasoning models; (2) it has broad and structured coverage, spanning 8 major scientific disciplines and over 1K fine-grained topics organized via a model-generated hierarchical taxonomy; and (3) it employs a fully automated, scalable evaluation pipeline that uses strong reasoning models to cross-validate both problem validity and answer correctness. We use CHIMERA to post-train a 4B Qwen3 model. Despite the dataset's modest size, the resulting model achieves strong performance on a suite of challenging reasoning benchmarks, including GPQA-Diamond, AIME 24/25/26, HMMT 25, and Humanity's Last Exam, approaching or matching the reasoning performance of substantially larger models such as DeepSeek-R1 and Qwen3-235B.

05

Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data

Large language models (LLMs) are becoming the foundation for autonomous agents that can use tools to solve complex tasks. Reinforcement learning (RL) has emerged as a common approach for injecting such agentic capabilities, but typically under tightly controlled training setups. It often depends on carefully constructed task-solution pairs and substantial human supervision, which creates a fundamental obstacle to open-ended self-evolution toward superintelligent systems. In this paper, we propose Tool-R0 framework for training general purpose tool-calling agents from scratch with self-play RL, under a zero-data assumption. Initialized from the same base LLM, Tool-R0 co-evolves a Generator and a Solver with complementary rewards: one proposes targeted challenging tasks at the other's competence frontier and the other learns to solve them with real-world tool calls. This creates a self-evolving cycle that requires no pre-existing tasks or datasets. Evaluation on different tool-use benchmarks show that Tool-R0 yields 92.5 relative improvement over the base model and surpasses fully supervised tool-calling baselines under the same setting. Our work further provides empirical insights into self-play LLM agents by analyzing co-evolution, curriculum dynamics, and scaling behavior.

06

CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production

This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp, and Messenger. Starting from LLaMA 3.1, we refined models across 15 generations using data from both internal and external real-user traffic. Through continuous deployments from July 2024 to April 2025, we conducted controlled 7-day A/B tests showing consistent engagement improvements: 7 of 8 newly deployed models demonstrated positive lift over the baseline, with the strongest performers achieving up to 8.8% improvement in engagement breadth and 19.4% in engagement depth. We also observed substantial gains in steerability, with instruction following increasing from 59.2% to 84.8% and instruction violations decreasing from 26.6% to 5.8%. We detail the CharacterFlywheel process which integrates data curation, reward modeling to estimate and interpolate the landscape of engagement metrics, supervised fine-tuning (SFT), reinforcement learning (RL), and both offline and online evaluation to ensure reliable progress at each optimization step. We also discuss our methods for overfitting prevention and navigating production dynamics at scale. These contributions advance the scientific rigor and understanding of LLMs in social applications serving millions of users.