In an LLM inference pipeline, serving a request involves two fundamentally distinct computational phases with opposing hardware bottlenecks: the Prefill Phase (processing the prompt) and the Decode Phase (generating response tokens one by one). For years, both phases were executed on the same physical GPUs—until the physical limits of hardware demanded their complete separation.
The Hardware Incompatibility of Prefill and Decode
The operational profiles of the two phases could not be more contradictory:
- Prefill Phase: Processes thousands of prompt tokens simultaneously. Highly parallel, compute-dense ($O(N^2)$ FLOPS), and operates at near 100% GPU compute utilization.
- Decode Phase: Generates one token at a time sequentially. Highly memory-bandwidth bound, memory-latency sensitive, and leaves GPU tensor cores mostly idle.
[Monolithic Serving: Prefill & Decode Fight for GPU Resources]
GPU 1: [Massive 32k Prefill Arrives] ──► Decode streams freeze! (Spike in Time-To-First-Token)
[Disaggregated Serving: Physical Hardware Specialization]
Incoming Request
│
▼ (Step 1: High Compute)
[Prefill-Optimized GPU Cluster (e.g. H100 with massive tensor FLOPs)]
│
▼ (Transfers Computed KV Cache via Fast RDMA Network)
[Decode-Optimized GPU Cluster (e.g. Large HBM Memory Bandwidth)]
│
▼ (Step 2: Low-Latency Streaming)
Streaming Response to User (Zero Jitter, Predictable Latency!)
The Disaggregated Architecture (DistServe / Splitwise)
Disaggregated serving physically partitions inference clusters into specialized sub-clusters:
- Prefill Workers: Optimized for maximum FLOPS throughput, crunching large prompt contexts instantly.
- Fast KV Cache Transfer: Once the prefill completes, the computed Key-Value cache tensors are transmitted across high-speed RDMA networks (PCIe 5.0 / RoCE) to the decode cluster.
- Decode Workers: Dedicated exclusively to single-token streaming generation, operating with zero interference or latency jitter from sudden incoming prompt bursts.
By decoupling prefill from decode, systems achieve up to 10x reduction in tail latency while increasing overall cluster hardware utilization.