FairyFuse: Multiplication-Free LLM Inference on CPUs via Fused Ternary Kernels
FairyFuse introduces a novel approach to significantly optimize Large Language Model (LLM) inference specifically for CPU architectures. The technique aims to achieve "multiplication-free" LLM computations by leveraging fused ternary kernels. This innovation targets a crucial bottleneck in deploying sophisticated LLMs on commodity hardware, which typically relies on CPUs and lacks the specialized tensor processing units found in GPUs. By eliminating complex multiplication operations, FairyFuse seeks to drastically reduce computational overhead and power consumption, making LLM inference more efficient and accessible for a wider range of applications and devices. The use of "fused ternary kernels" implies a strategic combination of simplified, low-precision arithmetic operations, possibly quantizing model weights or activations to ternary values (-1, 0, 1) and integrating these operations into highly optimized, single-pass kernels. This not only speeds up computation but also enhances cache utilization and minimizes data movement, which are critical factors for performance on CPU-based systems. The proposed method represents a significant step towards enabling broader adoption and real-time execution of LLMs without requiring specialized hardware accelerators, thereby democratizing access to powerful AI capabilities.