For years, quantization was treated as a painful compromise: you compressed 16-bit floating-point weights into 4-bit integers only when you could not afford enough GPUs, accepting noticeable degradation in perplexity and reasoning ability. But modern quantization algorithms have transformed 4-bit compression into an almost zero-loss standard.
The Discovery of Salient Outlier Weights
Early naive quantization algorithms compressed all weight parameters uniformly. However, deep neural networks contain a tiny fraction (less than 1%) of critical salient outlier weights that govern feature activation magnitudes. Compressing these critical outliers distorted the model's entire internal representations.
[Naive Rounding: Destroys Salient Outliers]
Weights ──► Uniform 4-bit Rounding ──► Outlier Weights Distorted ──► Model Quality Degrades!
[Activation-Aware Weight Quantization (AWQ)]
Weights + Sample Activations ──► Identify Salient 1% Weights
│
▼
[Protect Salient Weights in Higher Precision / Scaled Channels]
│
▼
4-bit Quantized Weights: 99.8% Quality Preserved!
AWQ and GPTQ: Activation-Aware Protection
Modern quantization techniques like AWQ (Activation-aware Weight Quantization) inspect the activation distributions rather than inspecting weights in isolation. By identifying which weight channels correspond to large activation magnitudes and protecting them with custom per-group scaling factors, the remaining 99% of weights can be aggressively quantized to 4-bit integers with zero perceptible loss in reasoning quality.
Marlin Kernels: Hardware-Accelerated Mixed Precision
Quantizing weights solves the memory capacity problem, but decompression during inference can introduce compute overhead. The invention of the Marlin GPU kernel solved this by performing fused, highly optimized 4-bit weight decompression directly inside the GPU registers during matrix multiplication, achieving near-theoretical maximum memory bandwidth speeds.