- In multi-turn chats, document RAG, and codebase indexing, over 85% of input tokens represent identical shared context across requests.
- RadixAttention maintains an exact prefix tree in host and device memory, reusing computed KV blocks across disparate user sessions.
- Combining prefix caching with FP8 KV cache quantization quadruples concurrent throughput without accuracy degradation.
- Decoupling chunked prefill from decode loops prevents long-document prefills from causing decode token jitter.
Architectural Overview & Engineering Context
Implementing RadixAttention tree-based prefix caching and FP8 KV quantization in vLLM and SGLang to reduce prompt evaluation cost by 90%.
Modern production AI systems require rigorous systems-level optimization. Whether managing GPU memory allocations, designing low-latency retrieval pipelines, or orchestrating multi-agent state machines, understanding the underlying trade-offs separates fragile prototypes from mission-critical platforms.
System Topology & Data Flow
The diagram below outlines the core execution path and component decoupling for this architecture:
[Incoming User Request: System Prompt + Codebase AST + User Query]
โ
โผ
[Radix Trie Context Lookup Engine]
โ
โโโโโโโโโโโโโโโโโโโโโโโดโโโโโโโโโโโโโโโโโโโโโโ
โผ (Exact Cache Hit: 45,000 Tokens) โผ (Cache Miss: 500 Tokens)
[Reuse GPU Memory KV Paged Blocks] [Compute Forward GEMM Prefill]
โข Zero GPU Math FLOPs โข Compute FP8 Activation
โข Instant TTFT < 40ms โข Append to Radix Trie
โ โ
โโโโโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโ
โผ
[Autoregressive Decode Pipeline]
Production Implementation & Code Pattern
Below is the reference production pattern demonstrating the core execution flow, asynchronous handling, and schema validation:
# SGLang Radix Cache Server Setup
import requests
SERVER_URL = "http://localhost:30000/v1/chat/completions"
LARGE_SYSTEM_DOC = "Enterprise Knowledge Graph Spec v4.2..."
payload = {
"model": "meta-llama/Meta-Llama-3.1-70B-Instruct",
"messages": [
{"role": "system", "content": f"Enterprise Context:\n{LARGE_SYSTEM_DOC}"},
{"role": "user", "content": "What is the failover protocol for PostgreSQL?"}
],
"temperature": 0.1
}
# Request 1: Full prefill (~50,000 tokens) -> Latency: 1,200ms
resp1 = requests.post(SERVER_URL, json=payload).json()
# Request 2 (Different User Query, Same Doc): Hits Radix cache -> Latency: 38ms!
payload["messages"][1]["content"] = "How are Kafka consumer offsets committed?"
resp2 = requests.post(SERVER_URL, json=payload).json()
Quantitative Benchmarks & System Trade-Offs
Production telemetry across high-concurrency benchmarks demonstrates substantial improvements in throughput, latency, and memory utilization:
| Caching Strategy | TTFT (50k tokens) | Throughput (req/min) | VRAM Cost per Session |
|---|---|---|---|
| No Caching (Standard Prefill) | 1,450 ms | 12 req/min | 4.2 GB |
| Static Prefix Cache | 180 ms | 48 req/min | 2.1 GB |
| RadixAttention (FP16) | 38 ms | 114 req/min | 2.1 GB |
| RadixAttention + FP8 KV | 22 ms | 240 req/min | 1.05 GB |
Production Gotchas & Failure Modes
- Non-Deterministic System Prompts: Never embed dynamic timestamps at the beginning of system prompts, as this breaks prefix hash matches.
- LRU Cache Eviction Thrashing: When serving diverse documents, allocate at least 25% of GPU memory strictly for cache retention.
- RoPE Offset Alignment: Ensure paged attention kernels correctly calculate relative position IDs when resuming from cached prefixes.
SGLang RadixAttention & Google Gemini Context Caching Technical Whitepapers: Seminal industry release and technical findings. View Reference Paper / Announcement โ