If you deploy an LLM to monitor continuous streaming data (such as live log files or ongoing conversation transcripts), you cannot keep the KV cache forever: eventually, GPU memory fills up. But in early streaming experiments, when engineers evicted older tokens to keep a fixed sliding window of the last 4,000 tokens, a bizarre phenomenon occurred: the model's perplexity exploded into complete gibberish.
The Mystery of the First Token
Why did evicting Token 1 and Token 2 break generation when the model was generating Token 4,001? Because of a hidden property of Softmax normalization known as Attention Sinks.
In autoregressive models, the softmax function across attention heads requires all attention scores to sum to $1.0$. Even when an attention head has no relevant context to retrieve for the current token, it is mathematically forced to allocate attention somewhere. Models naturally learn during pre-training to dump excessive, unneeded attention weight onto the very first tokens of the sequence (Tokens 0, 1, and 2).
[Naive Sliding Window: Discards Token 0 ──► Perplexity Explodes!] [Token 0 (Sink Dropped!)] ... [Token 3900 to 4000] ──► Attention Distribution Breaks! ──► Gibberish Output! [StreamingLLM Architecture: Preserves Initial 4 Attention Sinks + Rolling Cache] [Tokens 0, 1, 2, 3 (Attention Sinks Cached Forever!)] + [Rolling Window of Recent 4,000 Tokens] (Maintains Stable Perplexity over Millions of Streaming Tokens with Fixed Constant Memory!)
The StreamingLLM Discovery
Xiao et al. discovered that you do not need to store the entire historical sequence. By simply retaining the first 4 initial tokens (the attention sinks) alongside a rolling sliding window of the most recent $K$ tokens, the attention distribution remains perfectly calibrated.
StreamingLLM enables language models to stream continuously over millions of tokens with strictly constant memory and zero degradation in fluency.