THE GENERATIVE VIDEO LAB

Motion. From noise.

A frame is a field of possibilities. Explore how noise becomes structure, how patches become tokens, and why a video needs more than beautiful individual frames.

CONCEPTUAL SIMULATION · NO MODEL INFERENCE

01 / SIGNAL EMERGESSIGNAL 65%
“A coral sun over an electric-blue mountain lake”64 patches

xₜ = √ᾱₜ x₀ + √(1 − ᾱₜ) ε

This is the DDPM forward marginal, with ε ∼ N(0, I). Our slider chooses ᾱ directly; the same noise is reused while scrubbing. Revealing a known image illustrates the mixture, not a learned reverse sampling algorithm. Display values are clipped to the screen’s color range.

02 / THE FOURTH DIMENSION

Good frames.
Do they agree?

A changing sun and drifting mountains make a plausible frame sequence feel unstable. Increase consistency to keep identity and geometry steady while the camera moves.

A toy illustration of temporal coherence. Real video models learn spatial and temporal relationships; this control does not measure a model.

FRAME 00024 FPS · 4 SECOND LOOP
03 / DIFFUSION TRANSFORMERS

Pixels → patches → tokens

DiT applies a transformer to patches of a latent representation. Attention mixes information between tokens; the model predicts denoising information conditioned on the noise level.

Our grid overlays image space for visibility. It is an analogy for latent patchification, not the internal activations of a DiT.

04 / A DIFFERENT PATH

Flow matching

One simple conditional path is x(t) = (1 − t)ε + tx₁, from noise at t = 0 to data at t = 1. Its conditional velocity is x₁ − ε. Flow matching trains a velocity field; generation integrates the learned field.

This differs from simply fading a known image into view. Neither demonstration here is a trained generative model.