NO/FOMO

每天一次,过滤 AI 噪音

值得打开的
AI 日报。

发布日期2026-04-16中文版本
本期阅读
—
累计阅读
—

Hacker News

6 stories
01

Claude Opus 4.7

Anthropic has unveiled its latest large language model, Claude Opus 4.7, signifying a notable advancement in artificial intelligence capabilities. This new iteration is expected to build upon its predecessors' strengths, offering enhanced performance across a range of complex cognitive tasks. While specific technical details regarding Opus 4.7's architecture and benchmarks are highly anticipated, the announcement implies significant improvements in areas such as advanced reasoning, sophisticated problem-solving, and nuanced language understanding. The model is poised to demonstrate heightened proficiency in generating coherent and contextually appropriate text, assisting with intricate coding challenges, and delivering more insightful analytical capabilities. The continuous evolution showcased by models like Claude Opus 4.7 underscores the rapid pace of innovation in AI research, promising to deliver increasingly capable and versatile AI systems for a broad spectrum of real-world applications and industries.

02

Qwen3.6-35B-A3B: Agentic coding power, now open to all

Alibaba Cloud has officially announced the open availability of Qwen3.6-35B-A3B, a substantial advancement in large language models specifically engineered for agentic coding. This formidable 35-billion-parameter model is meticulously designed to deliver robust performance in automating complex programming tasks, encompassing intelligent code generation, sophisticated refinement processes, and the orchestration of multi-step development workflows typically handled by advanced AI agents. Its strategic release as an open-source resource democratizes access to powerful AI-driven coding capabilities, thereby enabling a broader global community of developers and researchers to seamlessly integrate cutting-edge automation into their software development pipelines. This move by Alibaba Cloud not only underscores a strong commitment to fostering innovation within agent-based systems but also holds the potential to significantly revolutionize how software is conceived, written, and maintained. Qwen3.6-35B-A3B is therefore poised to accelerate development cycles, substantially enhance code quality, and open entirely new avenues for highly autonomous software engineering applications across a diverse range of technical sectors.

03

Codex for almost everything

OpenAI's "Codex for almost everything" highlights an advanced artificial intelligence model engineered to translate natural language into executable code, marking a substantial advancement in AI-assisted programming. Built upon the architecture of GPT-3, Codex excels in comprehending and generating code across a diverse range of programming languages, such as Python, JavaScript, Go, and Ruby. Its fundamental innovation resides in its capability to interpret human instructions and subsequently transform them into functional code snippets, complete functions, or even entire programs. This groundbreaking technology is designed to democratize software development by significantly reducing the barrier to entry, thereby empowering both seasoned developers to enhance their productivity and non-programmers to embark on application creation. Codex serves as the foundational AI behind influential developer tools like GitHub Copilot, illustrating its potent practical applications in authentic coding environments and foreshadowing a future where AI assumes an increasingly vital role throughout the software development lifecycle, rendering intricate coding tasks more approachable and efficient for a wider spectrum of users.

04

Cloudflare's AI Platform: an inference layer designed for agents

Cloudflare has officially launched its new AI Platform, an inference layer specifically engineered to support and optimize the deployment and operation of AI agents. This strategic offering aims to provide developers with a robust, scalable, and low-latency infrastructure for running AI models, leveraging Cloudflare's extensive global network. The platform is designed to facilitate efficient execution of inference tasks, bringing AI capabilities closer to the edge and reducing computational overhead for real-time applications. By focusing on an inference layer, Cloudflare addresses the critical need for performant and cost-effective AI model serving, particularly as the complexity and demand for agent-based AI systems continue to grow. This initiative underscores Cloudflare's expansion into core AI infrastructure, offering a foundational service that can accelerate the development and deployment of next-generation AI applications while benefiting from Cloudflare's established security and reliability features.

05

We gave an AI a 3 year retail lease and asked it to make a profit

Andon Labs has initiated a groundbreaking project, granting an Artificial Intelligence agent a three-year retail lease with the explicit mandate to generate profit. This innovative experiment, highlighted by the launch of 'Andon Market,' signifies a substantial advancement in the practical application of AI within commercial domains. The core objective is to assess the AI's capacity for autonomous decision-making and operational management across a comprehensive spectrum of retail activities. This includes responsibilities ranging from inventory optimization and dynamic pricing to sales strategy formulation, customer engagement, and ultimately, ensuring financial viability. The initiative positions AI not merely as an analytical tool but as an active, profit-driven entity capable of navigating the complexities of a real-world business environment. This development reflects a broader industry trend towards deploying sophisticated AI agents for autonomous execution in high-stakes commercial scenarios, thereby challenging conventional paradigms of business management and expanding the perceived utility of artificial intelligence.

06

The future of everything is lies, I guess: Where do we go from here?

This piece initiates a profound inquiry into the prospective state of information environments, provocatively asserting that 'The future of everything is lies.' This statement, coupled with the interrogative 'Where do we go from here?', frames a critical discussion likely centered on the escalating challenges posed by misinformation, fabricated content, and the blurring lines between reality and simulation in the digital age. Within the context of Hacker News, the article is presumed to explore the far-reaching implications of advanced technologies, especially the rapid progress in artificial intelligence and generative models. These technologies, while offering significant advancements, also present an unprecedented capacity to create highly convincing synthetic media, deepfakes, and sophisticated narratives that can undermine public trust and distort perception. The author appears to challenge the technical community and broader society to confront these ethical dilemmas, prompting a deeper consideration of the strategies, frameworks, and technological countermeasures necessary to preserve truth and foster digital literacy in an increasingly complex information ecosystem. The article implicitly calls for urgent discourse on the responsibilities of technology developers and users in shaping a more resilient and trustworthy future.

huggingface

6 stories
01

Seedance 2.0: Advancing Video Generation for World Complexity

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and large-scale architecture for multi-modal audio-video joint generation. This allows it to support four input modalities: text, image, audio, and video, by integrating one of the most comprehensive suites of multi-modal content reference and editing capabilities available in the industry to date. It delivers substantial, well-rounded improvements across all key sub-dimensions of video and audio generation. In both expert evaluations and public user tests, the model has demonstrated performance on par with the leading levels in the field. Seedance 2.0 supports direct generation of audio-video content with durations ranging from 4 to 15 seconds, with native output resolutions of 480p and 720p. For multi-modal inputs as reference, its current open platform supports up to 3 video clips, 9 images, and 3 audio clips. In addition, we provide Seedance 2.0 Fast version, an accelerated variant of Seedance 2.0 designed to boost generation speed for low-latency scenarios. Seedance 2.0 has delivered significant improvements to its foundational generation capabilities and multi-modal generation performance, bringing an enhanced creative experience for end users.

02

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models

AI agents are expected to perform professional work across hundreds of occupational domains (from emergency department triage to nuclear reactor safety monitoring to customs import processing), yet existing benchmarks can only evaluate agents in the few domains where public environments exist. We introduce OccuBench, a benchmark covering 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, enabled by Language World Models (LWMs) that simulate domain-specific environments through LLM-driven tool response generation. Our multi-agent synthesis pipeline automatically produces evaluation instances with guaranteed solvability, calibrated difficulty, and document-grounded diversity. OccuBench evaluates agents along two complementary dimensions: task completion across professional domains and environmental robustness under controlled fault injection (explicit errors, implicit data degradation, and mixed faults). We evaluate 15 frontier models across 8 model families and find that: (1) no single model dominates all industries, as each has a distinct occupational capability profile; (2) implicit faults (truncated data, missing fields) are harder than both explicit errors (timeouts, 500s) and mixed faults, because they lack overt error signals and require the agent to independently detect data degradation; (3) larger models, newer generations, and higher reasoning effort consistently improve performance. GPT-5.2 improves by 27.5 points from minimal to maximum reasoning effort; and (4) strong agents are not necessarily strong environment simulators. Simulator quality is critical for LWM-based evaluation reliability. OccuBench provides the first systematic cross-industry evaluation of AI agents on professional occupational tasks.

03

SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineering

The rise of OpenClaw in early 2026 marks the moment when millions of users began deploying personal AI agents into their daily lives, delegating tasks ranging from travel planning to multi-step research. This scale of adoption signals that two parallel arcs of development have reached an inflection point. First is a paradigm shift in AI engineering, evolving from prompt and context engineering to harness engineering-designing the complete infrastructure necessary to transform unconstrained agents into controllable, auditable, and production-reliable systems. As model capabilities converge, this harness layer is becoming the primary site of architectural differentiation. Second is the evolution of human-agent interaction from discrete tasks toward a persistent, contextually aware collaborative relationship, which demands open, trustworthy and extensible harness infrastructure. We present SemaClaw, an open-source multi-agent application framework that addresses these shifts by taking a step towards general-purpose personal AI agents through harness engineering. Our primary contributions include a DAG-based two-phase hybrid agent team orchestration method, a PermissionBridge behavioral safety system, a three-tier context management architecture, and an agentic wiki skill for automated personal knowledge base construction.

04

From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space

While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base model's existing output distribution. Optimizing the marginal distribution P(y) in the Pre-train Space addresses this bottleneck by encoding reasoning ability and preserving broad exploration capacity. Yet, conventional pre-training relies on static corpora for passive learning, leading to a distribution shift that hinders targeted reasoning enhancement. In this paper, we introduce PreRL (Pre-train Space RL), which applies reward-driven online updates directly to P(y). We theoretically and empirically validate the strong gradient alignment between log P(y) and log P(y|x), establishing PreRL as a viable surrogate for standard RL. Furthermore, we uncover a critical mechanism: Negative Sample Reinforcement (NSR) within PreRL serves as an exceptionally effective driver for reasoning. NSR-PreRL rapidly prunes incorrect reasoning spaces while stimulating endogenous reflective behaviors, increasing transition and reflection thoughts by 14.89x and 6.54x, respectively. Leveraging these insights, we propose Dual Space RL (DSRL), a Policy Reincarnation strategy that initializes models with NSR-PreRL to expand the reasoning horizon before transitioning to standard RL for fine-grained optimization. Extensive experiments demonstrate that DSRL consistently outperforms strong baselines, proving that pre-train space pruning effectively steers the policy toward a refined correct reasoning subspace.

05

LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling

Continuous diffusion has been the foundation of high-fidelity, controllable, and few-step generation of many data modalities such as images. However, in language modeling, prior continuous diffusion language models (DLMs) lag behind discrete counterparts due to the sparse data space and the underexplored design space. In this work, we close this gap with LangFlow, the first continuous DLM to rival discrete diffusion, by connecting embedding-space DLMs to Flow Matching via Bregman divergence, alongside three key innovations: (1) we derive a novel ODE-based NLL bound for principled evaluation of continuous flow-based language models; (2) we propose an information-uniform principle for setting the noise schedule, which motivates a learnable noise scheduler based on a Gumbel distribution; and (3) we revise prior training protocols by incorporating self-conditioning, as we find it improves both likelihood and sample quality of embedding-space DLMs with effects substantially different from discrete diffusion. Putting everything together, LangFlow rivals top discrete DLMs on both the perplexity (PPL) and the generative perplexity (Gen. PPL), reaching a PPL of 30.0 on LM1B and 24.6 on OpenWebText. It even exceeds autoregressive baselines in zero-shot transfer on 4 out of 7 benchmarks. LangFlow provides the first clear evidence that continuous diffusion is a promising paradigm for language modeling. Homepage: https://github.com/nealchen2003/LangFlow

06

HDR Video Generation via Latent Alignment with Logarithmic Encoding

High dynamic range (HDR) imagery offers a rich and faithful representation of scene radiance, but remains challenging for generative models due to its mismatch with the bounded, perceptually compressed data on which these models are trained. A natural solution is to learn new representations for HDR, which introduces additional complexity and data requirements. In this work, we show that HDR generation can be achieved in a much simpler way by leveraging the strong visual priors already captured by pretrained generative models. We observe that a logarithmic encoding widely used in cinematic pipelines maps HDR imagery into a distribution that is naturally aligned with the latent space of these models, enabling direct adaptation via lightweight fine-tuning without retraining an encoder. To recover details that are not directly observable in the input, we further introduce a training strategy based on camera-mimicking degradations that encourages the model to infer missing high dynamic range content from its learned priors. Combining these insights, we demonstrate high-quality HDR video generation using a pretrained video model with minimal adaptation, achieving strong results across diverse scenes and challenging lighting conditions. Our results indicate that HDR, despite representing a fundamentally different image formation regime, can be handled effectively without redesigning generative models, provided that the representation is chosen to align with their learned priors.