Early multimodal vision models suffered from an embarrassing blindness: when shown a high-resolution architectural blueprint, a circuit diagram, or a dense document scan, they failed to read fine text or spot small spatial details. The root cause was a primitive image preprocessing pipeline: fixed-resolution downsampling.
The Destruction of Fixed 224x224 Resizing
Early Vision Transformers (ViTs) required all input images to be resized into fixed square dimensions (e.g. 224x224 or 336x336 pixels). When a 4K panoramic blueprint was forced into a tiny square, aspect ratios were distorted, and fine text and small circuit components were permanently blurred into unrecognizable pixels.
[Naive Fixed Resizing: Fine Details Destroyed]
4K Blueprint (4000x2000) ──► Squashed to 336x336 Square ──► Illegible Mud (OCR & Details Lost!)
[Dynamic Patch Slicing (LLaVA-NeXT / AnyRes)]
4K Blueprint (4000x2000)
│
├──► Overview Thumbnail (Low-res global context: 336x336)
│
└──► Grid Slicing (Native aspect ratio):
┌──────────┬──────────┬──────────┐
│ Patch 1 │ Patch 2 │ Patch 3 │ (Each patch processed at full 336x336 resolution!)
├──────────┼──────────┼──────────┤
│ Patch 4 │ Patch 5 │ Patch 6 │
└──────────┴──────────┴──────────┘
│
▼
[Concatenated Vision Tokens + Spatial Position Embeddings]
The Dynamic Slicing Architecture
Modern Vision-Language Models (such as LLaVA-NeXT and Qwen2-VL) implement dynamic high-resolution patching:
- Global Context Thumbnail: The full image is downscaled to provide a low-resolution global overview.
- Native Aspect Ratio Grid Partitioning: The original high-resolution image is sliced into an optimal grid of local tiles (e.g. $2 \times 3$ or $3 \times 3$), each preserving its native aspect ratio and full optical resolution.
- Spatial Position Encoding: 2D spatial coordinate tokens are appended to each tile, allowing the transformer to reason seamlessly across local high-resolution details and global composition.
Dynamic patching unlocked human-level visual acuity for complex engineering diagrams, medical imaging, and dense document analysis.