NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-04-28中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

OpenAI models coming to Amazon Bedrock: Interview with OpenAI and AWS CEOs

OpenAI is integrating its cutting-edge AI models into Amazon Bedrock, AWS's comprehensive managed service for foundation models. This landmark partnership, highlighted by a joint interview with the CEOs of OpenAI and AWS, marks a strategic expansion of OpenAI's reach, making its powerful technologies more accessible to a vast enterprise customer base already operating within the AWS cloud ecosystem. The collaboration allows AWS clients to seamlessly develop, deploy, and scale sophisticated generative AI applications by leveraging OpenAI's models directly through Bedrock, complementing the existing array of foundation models available on the platform. This move significantly bolsters Bedrock's value proposition, offering developers unparalleled choice and flexibility in selecting the optimal AI models for diverse applications, ranging from advanced natural language processing to creative content generation and complex AI agents. The interview likely explores the strategic motivations behind this alliance, the technical intricacies of the integration, and the anticipated long-term implications for the competitive landscape of cloud-based AI services, underscoring both companies' shared vision for accelerating enterprise-wide adoption of generative AI innovation.

02

Microsoft VibeVoice: Open-Source Frontier Voice AI

Microsoft has announced VibeVoice, an open-source initiative poised to significantly advance the field of voice artificial intelligence. This project aims to push the boundaries of current voice AI capabilities, offering a collaborative platform for researchers and developers to innovate in areas such as realistic speech synthesis, expressive voice generation, and nuanced vocal dynamics. By making VibeVoice open-source, Microsoft is actively fostering community-driven development and accelerating progress in a critical domain of human-computer interaction. The 'frontier' designation implies a strong focus on leveraging cutting-edge techniques and methodologies, likely including advanced deep learning models, to create highly natural and emotionally resonant voice experiences. This strategic move is expected to democratize access to sophisticated voice AI technologies, encouraging wider adoption and exploration of novel applications across various industries, from accessibility tools to interactive entertainment and advanced digital assistants.

03

Who owns the code Claude Code wrote?

The ownership of code generated by artificial intelligence models, specifically 'Claude Code', presents a complex and evolving legal and intellectual property challenge. This query highlights a significant debate within the tech and legal communities regarding traditional copyright frameworks' applicability to AI-created works. Key questions arise concerning whether AI-generated output is copyrightable, and if so, whether ownership rests with the AI developer, the user who provided the prompt, or if it defaults to the public domain. The discussion also touches upon potential copyright infringement risks associated with AI models trained on copyrighted material and the implications for software licensing. Resolving these ownership ambiguities is critical for encouraging responsible innovation, establishing clear legal guidelines for AI-assisted development, and mitigating future disputes in the rapidly advancing field of generative AI.

04

Laguna XS.2 and M.1

Poolside.ai has released or provided a detailed exploration of its "Laguna" project, focusing on the distinct models designated as XS.2 and M.1. Although the direct content is brief, the accompanying blog post titled "Laguna: A Deeper Dive" on Poolside.ai's platform indicates these models represent significant advancements within the domain of AI-driven software engineering and automated code generation. These specific iterations, XS.2 and M.1, likely differentiate themselves by scale, performance, or specialized functionalities within the broader Laguna framework, aiming to significantly enhance developer productivity, automate complex coding tasks, or provide sophisticated AI assistance throughout the software development lifecycle. This initiative underscores Poolside.ai's ongoing efforts to innovate artificial intelligence applications that streamline and revolutionize the software creation process, potentially transforming how development teams tackle intricate projects and leverage AI for code synthesis and optimization. The blog entry is expected to elaborate on their technical architecture, performance benchmarks, and practical applications.

05

GitHub Copilot code review will start consuming GitHub Actions minutes

GitHub has announced a significant change to its billing policy for GitHub Copilot's code review feature. Effective June 1, 2026, the use of GitHub Copilot for code reviews will begin to consume GitHub Actions minutes. This policy update means that organizations and developers leveraging AI-powered code review capabilities will now incur usage costs against their allocated GitHub Actions budget, marking a shift from what may have previously been an unmetered service. This change has implications for development teams, requiring them to account for Copilot's review activity in their operational cost planning and resource management for CI/CD pipelines. It underscores an evolving monetization strategy for advanced AI developer tools integrated within cloud development platforms, tying AI service consumption directly to compute resources.

06

Google and Pentagon reportedly agree on deal for 'any lawful' use of AI

Google has reportedly formalized an agreement with the U.S. Pentagon for the application of artificial intelligence technologies in defense operations. This landmark deal emphasizes that AI deployment will be strictly limited to 'any lawful' uses, signaling a deliberate effort to address the ethical and legal complexities inherent in military AI. The collaboration marks a significant step in the ongoing integration of advanced AI capabilities within national security frameworks. Such partnerships between leading technology companies and governmental defense agencies often ignite considerable public and internal discussions concerning the moral responsibilities associated with developing and deploying AI for warfare. This agreement suggests a strategic move by the Pentagon to harness cutting-edge AI for various operational needs, while Google navigates the intricate landscape of government contracts and the societal implications of its technology being utilized in a defense context. It underscores the broader trend of AI becoming a critical component of modern military strategy, prompting continued scrutiny over its oversight and responsible implementation.

huggingface

6 stories
01

From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company

Individual agent capabilities have advanced rapidly through modular skills and tool integrations, yet multi-agent systems remain constrained by fixed team structures, tightly coupled coordination logic, and session-bound learning. We argue that this reflects a deeper absence: a principled organisational layer that governs how a workforce of agents is assembled, governed, and improved over time, decoupled from what individual agents know. To fill this gap, we introduce OneManCompany (OMC), a framework that elevates multi-agent systems to the organisational level. OMC encapsulates skills, tools, and runtime configurations into portable agent identities called Talents, orchestrated through typed organisational interfaces that abstract over heterogeneous backends. A community-driven Talent Market enables on-demand recruitment, allowing the organisation to close capability gaps and reconfigure itself dynamically during execution. Organisational decision-making is operationalised through an Explore-Execute-Review (E^2R) tree search, which unifies planning, execution, and evaluation in a single hierarchical loop: tasks are decomposed top-down into accountable units and execution outcomes are aggregated bottom-up to drive systematic review and refinement. This loop provides formal guarantees on termination and deadlock freedom while mirroring the feedback mechanisms of human enterprises. Together, these contributions transform multi-agent systems from static, pre-configured pipelines into self-organising and self-improving AI organisations capable of adapting to open-ended tasks across diverse domains. Empirical evaluation on PRDBench shows that OMC achieves an 84.67% success rate, surpassing the state of the art by 15.48 percentage points, with cross-domain case studies further demonstrating its generality.

02

World-R1: Reinforcing 3D Constraints for Text-to-Video Generation

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high computational costs and limit scalability. We propose World-R1, a framework that aligns video generation with 3D constraints through reinforcement learning. To facilitate this alignment, we introduce a specialized pure text dataset tailored for world simulation. Utilizing Flow-GRPO, we optimize the model using feedback from pre-trained 3D foundation models and vision-language models to enforce structural coherence without altering the underlying architecture. We further employ a periodic decoupled training strategy to balance rigid geometric consistency with dynamic scene fluidity. Extensive evaluations reveal that our approach significantly enhances 3D consistency while preserving the original visual quality of the foundation model, effectively bridging the gap between video generation and scalable world simulation.

03

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Unified multimodal models typically rely on pretrained vision encoders and use separate visual representations for understanding and generation, creating misalignment between the two tasks and preventing fully end-to-end optimization from raw pixels. We introduce Tuna-2, a native unified multimodal model that performs visual understanding and generation directly based on pixel embeddings. Tuna-2 drastically simplifies the model architecture by employing simple patch embedding layers to encode visual input, completely discarding the modular vision encoder designs such as the VAE or the representation encoder. Experiments show that Tuna-2 achieves state-of-the-art performance in multimodal benchmarks, demonstrating that unified pixel-space modelling can fully compete with latent-space approaches for high-quality image generation. Moreover, while the encoder-based variant converges faster in early pretraining, Tuna-2's encoder-free design achieves stronger multimodal understanding at scale, particularly on tasks requiring fine-grained visual perception. These results show that pretrained vision encoders are not necessary for multimodal modelling, and end-to-end pixel-space learning offers a scalable path toward stronger visual representations for both generation and perception.

04

Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis

Process Reward Models (PRMs) have achieved remarkable success in augmenting the reasoning capabilities of Large Language Models (LLMs) within static domains such as mathematics. However, their potential in dynamic data analysis tasks remains underexplored. In this work, we first present a empirical study revealing that general-domain PRMs struggle to supervise data analysis agents. Specifically, they fail to detect silent errors, logical flaws that yield incorrect results without triggering interpreter exceptions, and erroneously penalize exploratory actions, mistaking necessary trial-and-error exploration for grounding failures. To bridge this gap, we introduce DataPRM, a novel environment-aware generative process reward model that (1) can serve as an active verifier, autonomously interacting with the environment to probe intermediate execution states and uncover silent errors, and (2) employs a reflection-aware ternary reward strategy that distinguishes between correctable grounding errors and irrecoverable mistakes. We design a scalable pipeline to construct over 8K high-quality training instances for DataPRM via diversity-driven trajectory generation and knowledge-augmented step-level annotation. Experimental results demonstrate that DataPRM improves downstream policy LLMs by 7.21% on ScienceAgentBench and 11.28% on DABStep using Best-of-N inference. Notably, with only 4B parameters, DataPRM outperforms strong baselines, and exhibits robust generalizability across diverse Test-Time Scaling strategies. Furthermore, integrating DataPRM into Reinforcement Learning yields substantial gains over outcome-reward baselines, achieving 78.73% on DABench and 64.84% on TableBench, validating the effectiveness of process reward supervision. Code is available at https://github.com/zjunlp/DataMind.

05

How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models

We measure how much one extra recurrence is worth to a looped (depth-recurrent) language model, in equivalent unique parameters. From an iso-depth sweep of 116 pretraining runs across recurrence counts r in {1, 2, 4, 8} spanning {sim}50times in training compute, we fit a joint scaling law L = E + A,(N_once + r^φ N_rec)^{-α} + B,D^{-β} and recover a new recurrence-equivalence exponent φ= 0.46. Intuitively, φ tells us whether looping a block r times is equivalent in validation loss to r unique blocks of a non-looped model (full equivalence, φ{=}1) or to a single block run repeatedly with no capacity gain (φ{=}0). Our φ= 0.46 sits in between, so each additional recurrence predictably increases validation loss at matched training compute. For example, at r{=}4 a 410M looped model performs on par with a 580M non-looped model, but incurs the training cost of a 1B non-looped one. We demonstrate the utility of φ as a measurement tool on two probes. Truncated backpropagation lowers φ to 0.38, indicating that the loop mechanism is poorly trained under truncation, even though validation loss decreases. Conversely, hyperconnections raise φ to 0.65, a genuine capacity gain. Our method applies to any looped LM and separates true loop improvements from token-budget gains.

06

Discovering Agentic Safety Specifications from 1-Bit Danger Signals

Can large language model agents discover hidden safety objectives through experience alone? We introduce EPO-Safe (Experiential Prompt Optimization for Safe Agents), a framework where an LLM iteratively generates action plans, receives sparse binary danger warnings, and evolves a natural language behavioral specification through reflection. Unlike standard LLM reflection methods that rely on rich textual feedback (e.g., compiler errors or detailed environment responses), EPO-Safe demonstrates that LLMs can perform safety reasoning from a strictly impoverished signal in structured, low-dimensional environments: the agent never observes the hidden performance function R^*, only a single bit per timestep indicating that an action was unsafe. We evaluate on five AI Safety Gridworlds (Leike et al., 2017) and five text-based scenario analogs where visible reward R may diverge from R^*. EPO-Safe discovers safe behavior within 1-2 rounds (5-15 episodes), producing human-readable specifications with correct explanatory hypotheses about hazards (e.g., "X cells are directionally hazardous: entering from the north is dangerous"). Critically, we show that standard reward-driven reflection actively degrades safety: agents reflecting on reward alone use the loop to justify and accelerate reward hacking, proving that reflection must be paired with a dedicated safety channel to discover hidden constraints. We further evaluate robustness to noisy oracles: even when 50% of non-dangerous steps produce spurious warnings, mean safety performance degrades by only 15% on average, though sensitivity is environment-dependent, as cross-episode reflection naturally filters inconsistent signals. Each evolved specification functions as an auditable set of grounded behavioral rules discovered autonomously through interaction, rather than authored by humans as in Constitutional AI (Bai et al., 2022).