A frame is a field of possibilities. Explore how noise becomes structure, how patches become tokens, and why a video needs more than beautiful individual frames.
CONCEPTUAL SIMULATION · NO MODEL INFERENCE
01 / SIGNAL EMERGESSIGNAL 65%
“A coral sun over an electric-blue mountain lake”64 patches
xₜ = √ᾱₜ x₀ + √(1 − ᾱₜ) ε
This is the DDPM forward marginal, with ε ∼ N(0, I). Our slider chooses ᾱ directly; the same noise is reused while scrubbing. Revealing a known image illustrates the mixture, not a learned reverse sampling algorithm. Display values are clipped to the screen’s color range.
02 / THE FOURTH DIMENSION
Good frames. Do they agree?
A changing sun and drifting mountains make a plausible frame sequence feel unstable. Increase consistency to keep identity and geometry steady while the camera moves.
A toy illustration of temporal coherence. Real video models learn spatial and temporal relationships; this control does not measure a model.
FRAME 00024 FPS · 4 SECOND LOOP
03 / DIFFUSION TRANSFORMERS
Pixels → patches → tokens
DiT applies a transformer to patches of a latent representation. Attention mixes information between tokens; the model predicts denoising information conditioned on the noise level.
Our grid overlays image space for visibility. It is an analogy for latent patchification, not the internal activations of a DiT.
04 / A DIFFERENT PATH
Flow matching
One simple conditional path is x(t) = (1 − t)ε + tx₁, from noise at t = 0 to data at t = 1. Its conditional velocity is x₁ − ε. Flow matching trains a velocity field; generation integrates the learned field.
This differs from simply fading a known image into view. Neither demonstration here is a trained generative model.