In the early era of generative image modeling (Stable Diffusion 1.5 and 2.0), convolutional U-Net architectures were the unquestioned gold standard for noise estimation. But when researchers attempted to scale generative diffusion to high-resolution, temporally coherent video spanning hundreds of frames, convolutional U-Nets hit an architectural wall.
The Inductive Bias Trap of Convolutions
Convolutions operate on local receptive fields: they process small $3 \times 3$ pixel neighborhoods. While local inductive bias works well for low-resolution images, video requires modeling complex long-range spatio-temporal dependencies: a ball thrown in Frame 1 must obey physical trajectory laws when caught in Frame 120.
[Convolutional U-Net: Local Receptive Fields, Poor Long-Range Video Scaling]
Video Frames ──► [Local 3D Convolutions] ──► Fails to capture global physics across long sequences!
[Diffusion Transformer (DiT / Sora Architecture): Full Spatio-Temporal Attention]
Video Sequence ──► 3D Spatio-Temporal Patch Slicing (Space-Time Latent Tokens)
│
▼
[Transformer Blocks: All-to-All Self-Attention over Space & Time]
(Scales predictably with compute according to Power Scaling Laws!)
│
▼
High-Fidelity, Physics-Consistent Generative Video!
The Diffusion Transformer (DiT) Breakthrough
Peebles and Xie replaced the convolutional backbone entirely with standard Vision Transformer blocks. By tokenizing video into 3D spatio-temporal latent patches (space-time cubes), the model processes video generation identically to language modeling: predicting noise tokens across a spatial and temporal sequence.
The Power of Predictable Scaling
Unlike U-Nets, which plateaued as compute increased, Diffusion Transformers adhere strictly to the empirical power-law scaling curves of Transformers: increase compute and model parameters, and video fidelity, physical consistency, and prompt alignment scale predictably.