NO/FOMO

Independent AI signal, once a day

The AI briefing worth opening.

ISSUE DATE2026-04-22DEFAULT EDITION
This issue
—
All time
—

Hacker News

6 stories
01

Our eighth generation TPUs: two chips for the agentic era

Google has officially unveiled its eighth generation of Tensor Processing Units (TPUs), introducing two distinct chips tailored specifically for the emerging "agentic era" of artificial intelligence. This strategic development underscores Google's commitment to providing cutting-edge infrastructure for the next wave of AI innovation, particularly focusing on systems capable of autonomous reasoning and complex task execution. These new TPUs are engineered to deliver substantial improvements in computational performance and energy efficiency, addressing the escalating demands of training and deploying sophisticated AI models, including large language models and other deep learning applications that power AI agents. By offering two specialized chips, Google likely aims to optimize for different phases of AI development perhaps one for intensive model training and another for high-performance, low-latency inference, thereby enhancing the end-to-end lifecycle of AI applications. This advancement is poised to empower researchers and developers leveraging Google Cloud to accelerate progress in creating more intelligent, adaptive, and agentic AI systems, pushing the boundaries of what is achievable in artificial intelligence.

02

OpenAI: Workspace Agents for Business

OpenAI is introducing a significant initiative centered on "Workspace Agents for Business," designed to integrate advanced artificial intelligence capabilities directly into corporate operational environments. These specialized AI agents are engineered to automate repetitive tasks, streamline complex workflows, and provide intelligent assistance to employees across various business functions. The primary objective is to substantially enhance enterprise productivity and operational efficiency by enabling organizations to leverage OpenAI's cutting-edge AI models to manage data, facilitate communication, and execute processes with greater autonomy and precision. This strategic development signifies OpenAI's expansion beyond foundational models and developer tools, directly addressing the growing demand for sophisticated AI-driven solutions within the business sector. The implementation of workspace agents is poised to transform how companies operate, empowering employees to allocate their focus towards more strategic and creative endeavors, while routine and time-consuming tasks are expertly managed by intelligent automation. This initiative firmly underscores OpenAI's ongoing commitment to making powerful, practical AI accessible and transformative for real-world business challenges and corporate innovation.

03

Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model

Alibaba Cloud has officially unveiled Qwen3.6-27B, a novel large language model specifically engineered to achieve flagship-level coding proficiency despite its comparatively compact 27-billion parameter dense architecture. This introduction represents a notable stride in developing highly capable yet resource-efficient AI models, directly addressing the challenge of delivering top-tier performance without relying solely on massively scaled parameter counts. Qwen3.6-27B is meticulously designed to excel across a spectrum of programming tasks, including complex code generation, intelligent debugging, and comprehensive code comprehension, positioning it as an invaluable asset for both software developers and AI researchers. Its emphasis on attaining leading performance in various coding benchmarks, while maintaining a more manageable model footprint, renders it particularly attractive for deployment in environments where computational efficiency and robust functionality are paramount. This advancement highlights a persistent trend within the AI community to refine model architectures for optimal size and resource utilization, continually expanding the practical applications of dense models in specialized fields such as software engineering and automated development.

04

Kernel code removals driven by LLM-created security reports

Large Language Models (LLMs) are demonstrating a significant impact on software security, particularly in the realm of kernel development. Reports generated by these AI models are now directly leading to code removals within the Linux kernel, indicating their efficacy in identifying vulnerabilities. This development underscores the expanding capabilities of LLMs beyond conventional text generation, showcasing their utility in complex analytical tasks such as security auditing and static code analysis. The integration of AI-driven tools for vulnerability detection marks a pivotal shift towards automated security practices, promising faster identification and remediation of critical flaws. While the specific mechanisms by which LLMs uncover these subtle vulnerabilities are subject to ongoing research, their tangible contribution to enhancing the robustness and security posture of foundational software like the kernel is becoming increasingly evident, highlighting AI's evolving role in maintaining software integrity.

05

Website streamed live directly from a model

The Hacker News entry introduces an intriguing development: a website capable of streaming content directly from an underlying AI model in real-time. This novel application signifies a significant leap in how dynamic web experiences can be created and delivered. Rather than relying on static or pre-rendered assets, the platform leverages an artificial intelligence model to generate or process content on-the-fly, which is then streamed live to users. While the specific type of AI model, its operational intricacies, or the exact nature of the streamed output (e.g., text, visual elements, interactive simulations) remain unspecified, the core concept points to advancements in real-time AI inference and generative capabilities. This approach opens new avenues for personalized user interfaces, highly adaptive content generation, and immersive interactive experiences, blurring the lines between traditional web development and cutting-edge AI integration. Such a system could herald a new era of responsive and intelligently driven web applications.

06

Over-editing refers to a model modifying code beyond what is necessary

The concept of "over-editing" in the context of AI models, particularly those involved in code modification, is introduced and defined as the phenomenon where a model makes changes to code that extend beyond the necessary scope of a given task. This behavior frequently results in extraneous or unintended alterations, which can potentially introduce new bugs, increase technical debt, or complicate the code review process. Over-editing stands in contrast to the ideal of "minimal editing," where AI systems are expected to deliver precise, targeted changes that directly address the user's intent without introducing unnecessary complexity or unwanted side effects. This issue holds significant relevance for the development of AI-powered code assistants, automated refactoring tools, and other software engineering applications where the efficiency and reliability of AI-driven code generation and modification are paramount. Addressing over-editing is crucial for enhancing the practical utility and trustworthiness of AI models within software development workflows, ultimately aiming for solutions that prioritize precision, user control, and targeted intervention over expansive, unrequested modifications.

huggingface

6 stories
01

AgentSPEX: An Agent SPecification and EXecution Language

Language-model agent systems commonly rely on reactive prompting, in which a single instruction guides the model through an open-ended sequence of reasoning and tool-use steps, leaving control flow and intermediate state implicit and making agent behavior potentially difficult to control. Orchestration frameworks such as LangGraph, DSPy, and CrewAI impose greater structure through explicit workflow definitions, but tightly couple workflow logic with Python, making agents difficult to maintain and modify. In this paper, we introduce AgentSPEX, an Agent SPecification and EXecution Language for specifying LLM-agent workflows with explicit control flow and modular structure, along with a customizable agent harness. AgentSPEX supports typed steps, branching and loops, parallel execution, reusable submodules, and explicit state management, and these workflows execute within an agent harness that provides tool access, a sandboxed virtual environment, and support for checkpointing, verification, and logging. Furthermore, we provide a visual editor with synchronized graph and workflow views for authoring and inspection. We include ready-to-use agents for deep research and scientific research, and we evaluate AgentSPEX on 7 benchmarks. Finally, we show through a user study that AgentSPEX provides a more interpretable and accessible workflow-authoring paradigm than a popular existing agent framework.

02

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently fail on (i) the structural stability of sensitive regions such as hands and faces and (ii) physically plausible contact (e.g., avoiding hand--object interpenetration). We present CoInteract, an end-to-end framework for HOI video synthesis conditioned on a person reference image, a product reference image, text prompts, and speech audio. CoInteract introduces two complementary designs embedded into a Diffusion Transformer (DiT) backbone. First, we propose a Human-Aware Mixture-of-Experts (MoE) that routes tokens to lightweight, region-specialized experts via spatially supervised routing, improving fine-grained structural fidelity with minimal parameter overhead. Second, we propose Spatially-Structured Co-Generation, a dual-stream training paradigm that jointly models an RGB appearance stream and an auxiliary HOI structure stream to inject interaction geometry priors. During training, the HOI stream attends to RGB tokens and its supervision regularizes shared backbone weights; at inference, the HOI branch is removed for zero-overhead RGB generation. Experimental results demonstrate that CoInteract significantly outperforms existing methods in structural stability, logical consistency, and interaction realism.

03

ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning

Parameter-efficient fine-tuning (PEFT) reduces the training cost of full-parameter fine-tuning for large language models (LLMs) by training only a small set of task-specific parameters while freezing the pretrained backbone. However, existing approaches, such as Low-Rank Adaptation (LoRA), achieve adaptation by inserting independent low-rank perturbations directly to individual weights, resulting in a local parameterization of adaptation. We propose ShadowPEFT, a centralized PEFT framework that instead performs layer-level refinement through a depth-shared shadow module. At each transformer layer, ShadowPEFT maintains a parallel shadow state and evolves it repeatedly for progressively richer hidden states. This design shifts adaptation from distributed weight-space perturbations to a shared layer-space refinement process. Since the shadow module is decoupled from the backbone, it can be reused across depth, independently pretrained, and optionally deployed in a detached mode, benefiting edge computing scenarios. Experiments on generation and understanding benchmarks show that ShadowPEFT matches or outperforms LoRA and DoRA under comparable trainable-parameter budgets. Additional analyses on shadow pretraining, cross-dataset transfer, parameter scaling, inference latency, and system-level evaluation suggest that centralized layer-space adaptation is a competitive and flexible alternative to conventional low-rank PEFT.

04

ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperation

Current AI agent frameworks have made remarkable progress in automating individual tasks, yet all existing systems serve a single user. Human productivity rests on the social and organizational relationships through which people coordinate, negotiate, and delegate. When agents move beyond performing tasks for one person to representing that person in collaboration with others, the infrastructure for cross-user agent collaboration is entirely absent, let alone the governance mechanisms needed to secure it. We argue that the next frontier for AI agents lies not in stronger individual capability, but in the digitization of human collaborative relationships. To this end, we propose a human-symbiotic agent paradigm. Each user owns a permanently bound agent system that collaborates on the owner's behalf, forming a network whose nodes are humans rather than agents. This paradigm rests on three governance primitives. A layered identity architecture separates a Manager Agent from multiple context-specific Identity Agents; the Manager Agent holds global knowledge but is architecturally isolated from external communication. Scoped authorization enforces per-identity access control and escalates boundary violations to the owner. Action-level accountability logs every operation against its owner's identity and authorization, ensuring full auditability. We instantiate this paradigm in ClawNet, an identity-governed agent collaboration framework that enforces identity binding and authorization verification through a central orchestrator, enabling multiple users to collaborate securely through their respective agents.

05

HP-Edit: A Human-Preference Post-Training Framework for Image Editing

Common image editing tasks typically adopt powerful generative diffusion models as the leading paradigm for real-world content editing. Meanwhile, although reinforcement learning (RL) methods such as Diffusion-DPO and Flow-GRPO have further improved generation quality, efficiently applying Reinforcement Learning from Human Feedback (RLHF) to diffusion-based editing remains largely unexplored, due to a lack of scalable human-preference datasets and frameworks tailored to diverse editing needs. To fill this gap, we propose HP-Edit, a post-training framework for Human Preference-aligned Editing, and introduce RealPref-50K, a real-world dataset across eight common tasks and balancing common object editing. Specifically, HP-Edit leverages a small amount of human-preference scoring data and a pretrained visual large language model (VLM) to develop HP-Scorer--an automatic, human preference-aligned evaluator. We then use HP-Scorer both to efficiently build a scalable preference dataset and to serve as the reward function for post-training the editing model. We also introduce RealPref-Bench, a benchmark for evaluating real-world editing performance. Extensive experiments demonstrate that our approach significantly enhances models such as Qwen-Image-Edit-2509, aligning their outputs more closely with human preference.

06

Micro Language Models Enable Instant Responses

Edge devices such as smartwatches and smart glasses cannot continuously run even the smallest 100M-1B parameter language models due to power and compute constraints, yet cloud inference introduces multi-second latencies that break the illusion of a responsive assistant. We introduce micro language models (μLMs): ultra-compact models (8M-30M parameters) that instantly generate the first 4-8 words of a contextually grounded response on-device, while a cloud model completes it; thus, masking the cloud latency. We show that useful language generation survives at this extreme scale with our models matching several 70M-256M-class existing models. We design a collaborative generation framework that reframes the cloud model as a continuator rather than a respondent, achieving seamless mid-sentence handoffs and structured graceful recovery via three error correction methods when the local opener goes wrong. Empirical results show that μLMs can initiate responses that larger models complete seamlessly, demonstrating that orders-of-magnitude asymmetric collaboration is achievable and unlocking responsive AI for extremely resource-constrained devices. The model checkpoint and demo are available at https://github.com/Sensente/micro_language_model_swen_project.