← Back to all stories

The Mystery of the Attention Sink: How Initial Tokens Unlocked Infinite Streaming Context

If you deploy an LLM to monitor continuous streaming data (such as live log files or ongoing conversation transcripts), you cannot keep the KV cache forever: eventually, GPU memory fills up. But in early streaming experiments, when engineers evicted older tokens to keep a fixed sliding window of the last 4,000 tokens, a bizarre phenomenon occurred: the model's perplexity exploded into complete gibberish.

The Mystery of the First Token

Why did evicting Token 1 and Token 2 break generation when the model was generating Token 4,001? Because of a hidden property of Softmax normalization known as Attention Sinks.

In autoregressive models, the softmax function across attention heads requires all attention scores to sum to $1.0$. Even when an attention head has no relevant context to retrieve for the current token, it is mathematically forced to allocate attention somewhere. Models naturally learn during pre-training to dump excessive, unneeded attention weight onto the very first tokens of the sequence (Tokens 0, 1, and 2).

[Naive Sliding Window: Discards Token 0 ──► Perplexity Explodes!]
[Token 0 (Sink Dropped!)] ... [Token 3900 to 4000] ──► Attention Distribution Breaks! ──► Gibberish Output!

[StreamingLLM Architecture: Preserves Initial 4 Attention Sinks + Rolling Cache]
[Tokens 0, 1, 2, 3 (Attention Sinks Cached Forever!)] + [Rolling Window of Recent 4,000 Tokens]
(Maintains Stable Perplexity over Millions of Streaming Tokens with Fixed Constant Memory!)

The StreamingLLM Discovery

Xiao et al. discovered that you do not need to store the entire historical sequence. By simply retaining the first 4 initial tokens (the attention sinks) alongside a rolling sliding window of the most recent $K$ tokens, the attention distribution remains perfectly calibrated.

StreamingLLM enables language models to stream continuously over millions of tokens with strictly constant memory and zero degradation in fluency.

Reference Paper / Context: Efficient Streaming Language Models with Attention Sinks (Xiao et al., MIT) — Read source ↗
About the Author

Vikram Samal is an AI systems architect focusing on test-time reasoning, high-throughput inference runtimes, and distributed agent infrastructure. Writing weekly architectural stories on Sundays.

Previous
← Supervisors vs. Swarms: The Architectural Topology of Multi-Agent Systems
Next
The Ghost in the Document: The Reality of Indirect Prompt Injection Defenses →