Text to video
Text-to-video adds time to image generation, and the hard part is making frame two agree with frame one about what the world contains.
- 14 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Text-to-video generates many frames at once and forces them to agree with each other.
That agreement is what makes it look like motion, not a slideshow.
The analogy you have already lived
Think about a flipbook you drew as a child. Draw a stick figure on twenty pages, then flip. If each drawing was made without looking at the previous page, the figure jumps and twitches.
Now think about how animators actually work. They draw the first pose and the last pose, then fill in between. Every frame is drawn in relation to its neighbours.
Text-to-video models had to learn that second habit. The first attempts drew each page separately, and the results twitched.
Why it exists
Video is what most people actually watch. Explaining, teaching, advertising and entertaining all happen in video, and producing it is slow and expensive.
The technical reason is more interesting. Generating video is not "generating images, several times". Try it and you get something unwatchable. The background changes colour. A shirt changes pattern. A face becomes a different face.
A picture has one thing to get right. A video has two: what is in it, and that it stays the same from frame to frame. The second is called temporal consistency, meaning things do not change unless they are supposed to.
How it works
The trick is to stop treating a video as a stack of independent pictures. Instead, the model works on the whole clip at once.
noise for frame 1 ┐
noise for frame 2 │
noise for frame 3 ├──> [ model that can look across frames ] ──> clip
noise for frame 4 │ (attention over time)
noise for frame 5 ┘The model can look sideways across time while it removes noise. When it decides the shirt is blue in frame 1, that decision is visible while it works on frame 4.
Most models were built by taking an image generator and adding time-aware layers to it. The image knowledge comes for free, and only the motion has to be learned.
Why this is so expensive
Five seconds at 24 frames a second is 120 pictures. That alone is 120 times the work of one image.
Then the model must compare frames with each other, and comparison cost grows faster than the frame count. So a short clip costs vastly more than one picture. Long clips are harder than short ones by more than their length suggests.
This is why video tools are slower, shorter and more expensive than image tools. They will stay that way for a while.
Where you have already seen it
- Short generated clips in social feeds and advertisements.
- Video editing tools that extend a shot or fill in a removed object.
- Slow-motion features on phones that invent frames between real ones.
- Old film restoration that raises frame rate and sharpness.
What is honestly hard here
Physics. Objects pass through each other. Water flows uphill. A poured drink does not fill the glass. The model learned what videos look like, not how the world works.
Length. Beyond a few seconds, things drift. A character's clothing changes. The camera loses track of the room.
Faces and hands over time. A face that looks right in one frame can become a slightly different face two seconds later.
Trust. Convincing video of things that never happened is now cheap to make. That is a societal problem this technology created and has not solved.
Remember this
- Video generation must get content right and keep it stable over time.
- Models work on the whole clip at once, looking across frames as they generate.
- The cost grows much faster than the number of frames, which limits clip length.
What to learn next
- Audio-visual learning — adding sound to moving pictures.
- Text to image — the single-frame case, where the machinery is easier to see.
- Latency and throughput — why generation cost decides what ships.
Developer — Code and libraries.
You cannot run a real video diffusion model on a laptop, and pretending otherwise wastes your time. What you can do is measure the exact problem those models exist to solve, in a way that takes two seconds and no downloads.
Setup
pip install numpyMeasuring flicker
This builds the same eight-frame clip twice. Once with every frame generated independently, and once with one shared noise draw plus a motion model. Then it measures how much the pixels change between neighbouring frames.
import numpy as np
H = W = 24
FRAMES = 8
def draw(canvas, cx, cy, size=6):
y0, x0 = int(cy) - size // 2, int(cx) - size // 2
canvas[max(y0, 0):y0 + size, max(x0, 0):x0 + size] = 1.0
return canvas
def per_frame_video(seed=0):
"""Generate every frame on its own. No frame knows about the others."""
rng = np.random.default_rng(seed)
frames = []
for t in range(FRAMES):
canvas = 0.15 * rng.normal(size=(H, W)) # fresh background noise each frame
cx = 4 + 2.0 * t + rng.normal() * 1.5 # the prompt says "moving ball",
cy = 12 + rng.normal() * 1.5 # but nothing pins down where
frames.append(draw(canvas, cx, cy))
return np.array(frames)
def shared_latent_video(seed=0):
"""One noise draw for the whole clip, plus a motion model."""
rng = np.random.default_rng(seed)
base = 0.15 * rng.normal(size=(H, W)) # the same background every frame
frames = []
for t in range(FRAMES):
frames.append(draw(base.copy(), 4 + 2.0 * t, 12))
return np.array(frames)
def flicker(video):
"""Mean pixel change between neighbouring frames. Low = smooth, high = jumpy."""
return float(np.abs(np.diff(video, axis=0)).mean())
a, b = per_frame_video(), shared_latent_video()
print(f"independent frames : flicker {flicker(a):.4f}")
print(f"shared latent : flicker {flicker(b):.4f}")
print(f"ratio: {flicker(a) / flicker(b):.1f}x more change per frame\n")
for name, video in (("independent", a), ("shared", b)):
centres = [(np.argwhere(f > 0.9).mean(axis=0)[1] if (f > 0.9).any() else -1) for f in video]
print(f"{name:>11} square centre x per frame:", [round(float(c), 1) for c in centres])
pixels_img = 512 * 512
pixels_vid = 512 * 512 * 24 * 5
print(f"\none 512x512 image : {pixels_img:>12,} pixels")
print(f"five seconds at 24 fps : {pixels_vid:>12,} pixels ({pixels_vid // pixels_img}x more)")independent frames : flicker 0.2176
shared latent : flicker 0.0420
ratio: 5.2x more change per frame
independent square centre x per frame: [2.5, 5.5, 7.5, 9.5, 8.5, 10.5, 13.5, 18.5]
shared square centre x per frame: [3.5, 5.5, 7.5, 9.5, 11.5, 13.5, 15.5, 17.5]
one 512x512 image : 262,144 pixels
five seconds at 24 fps : 31,457,280 pixels (120x more)Read the position lists, not the flicker number
The flicker ratio of 5.2 is the headline. The position lists are the lesson.
The independent clip goes 9.5 then 8.5. The object moved backwards for one frame. Every frame was individually plausible, and the sequence is not.
This is precisely what early text-to-video looked like, and it is why "generate frames independently and stitch" is not a strategy. There is no post-processing step that repairs it, because the information needed to repair it was never generated.
The shared-latent version steps evenly by 2.0 every frame, and its background noise is identical throughout. Both of those come from one design decision: share state across frames.
Line by line, the parts that are not obvious
np.diff(video, axis=0) subtracts each frame from the next along the time axis. Taking the mean absolute value gives a crude but genuinely used flicker measure — the same idea underlies warping-error metrics, which additionally compensate for motion using optical flow.
per_frame_video draws fresh noise inside the loop. shared_latent_video draws once outside it. That one line is the difference between a slideshow and a clip, and it is the toy version of what temporal layers do in a real model.
The > 0.9 threshold finds the square, since background noise stays near zero. np.argwhere(...).mean(axis=0)[1] takes the mean column index of those pixels, which is the square's horizontal centre.
The pixel table at the end is the honest cost statement: 120 times the raw data for five seconds of ordinary video, before any attention across frames is counted.
What real models do about it
Inflated image models. Take a pretrained image UNet and add temporal attention and temporal convolution layers between the existing spatial ones. Train only the new layers on video. This is how Video LDM and AnimateDiff work, and it reuses all the image knowledge.
Full spatiotemporal transformers. Treat the clip as a 3-D grid of latent patches and attend over all of them. More capable, far more expensive, and the direction the largest current models take.
Cascades. Generate a short, low-resolution, low-frame-rate clip, then run separate models to upsample in space and interpolate in time. Each stage is affordable, and errors compound across stages.
What running one actually costs
Open video models are large. Weights in the several-gigabyte range are typical, and generation needs a GPU with a lot of memory — commonly 12 GB or more for a few seconds at modest resolution. On a CPU, a short clip takes hours rather than minutes.
If you have a suitable GPU, diffusers exposes video pipelines with the same API shape as the image ones:
# Requires a GPU with substantial VRAM and a multi-gigabyte download.
# Check the model card for exact size and licence before running this.
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"<a text-to-video checkpoint>", torch_dtype=torch.float16
).to("cuda")
frames = pipe("a paper boat floating down a rain puddle", num_frames=16).frames[0]The checkpoint name is left as a placeholder on purpose. These repositories move and are renamed often, and a stale name in a lesson is worse than none. Pick a current one from the diffusers documentation, and read its VRAM requirement first.
This lesson's illustrative code is illustrative because the real thing genuinely does not fit on the machines most readers have. Saying so is more useful than a snippet that fails after a 10 GB download.
Common mistakes
Generating frames independently and hoping. The output list above shows why. Temporal information must be in the generation process.
Judging consistency by eye on one clip. Measure it. Warping error using optical flow, or CLIP similarity between consecutive frames, both give a number you can track.
Assuming more frames means a longer clip. Most models are trained at a fixed frame count. Asking for more usually degrades quality rather than extending the story.
Ignoring the first-frame anchor. Image-to-video conditioning, where you supply frame one, is dramatically more controllable than pure text-to-video. Use it when you can.
Try it yourself
Change shared_latent_video so the square's position is 4 + 2.0 * t + rng.normal() * 1.5 while the background stays shared. Rerun and watch the flicker land between the two original numbers. You have now separated the two independent causes of flicker: unstable content and unstable motion.
What to learn next
- Audio-visual learning — adding sound to moving pictures.
- Text to image — the single-frame case, where the machinery is easier to see.
- Latency and throughput — why generation cost decides what ships.
Researcher — Mathematics and papers.
The problem statement
A video is a tensor in R^{T x H x W x 3}. Diffusion over it is formally identical to the image case in text to image — the same forward noising and the same noise-prediction objective — with the denoiser now a function of a spatiotemporal input.
The entire difficulty is architectural and computational. Full 3-D self-attention over T * H * W latent positions is quadratic in that product, which is intractable at any useful resolution.
Factorised attention
Video Diffusion Models (Ho et al., 2022) introduced the 3D U-Net with factorised space-time attention: spatial attention within each frame, then temporal attention across frames at each spatial position. Cost drops from O((THW)^2) to O(T (HW)^2 + HW T^2).
The same paper introduced reconstruction guidance for extending clips, conditioning generation of new frames on already-generated ones through a guidance term rather than through retraining.
Make-A-Video (Singer et al., 2022) trained the spatial layers on image-text pairs and the temporal layers on unlabelled video, arguing that text-video pairs are scarce and unnecessary — motion can be learned without captions. Imagen Video (Ho et al., 2022) used a cascade of seven models: a base generator plus alternating spatial and temporal super-resolution stages.
Latent video diffusion
Align Your Latents (Blattmann et al., 2023) inflates a pretrained image LDM by inserting temporal layers and fine-tuning only those, keeping spatial weights frozen. This preserves image quality and requires far less video data.
Stable Video Diffusion (Blattmann et al., 2023) is the most instructive public account of the data pipeline. Its central finding is that data curation dominates: cut detection to remove multi-shot clips, optical-flow filtering to remove static or jittery clips, OCR filtering to remove text-heavy frames, and synthetic captioning. Training on uncurated video produced substantially worse motion regardless of model scale.
A separate subtlety is the VAE. An image VAE applied per frame produces temporally inconsistent latents. Video models fine-tune the decoder with temporal layers to remove the resulting flicker, which is a distinct fix from anything in the diffusion model itself.
Transformer-based video generation
The Sora technical report (OpenAI, 2024) describes patchifying video into spacetime latent patches and training a diffusion transformer over variable numbers of them, allowing variable duration, resolution and aspect ratio in one model. It reports emergent 3-D consistency and object permanence as scale increases, alongside persistent failures on physical interactions such as glass breaking.
The architectural direction — diffusion transformers over spacetime patches, following Peebles and Xie (2022) — is now standard in large video models, replacing the U-Net inflation approach.
Evaluation
- FVD (Unterthiner et al., 2018): Fréchet distance in the feature space of an I3D network trained on video. Inherits every FID caveat and adds sensitivity to clip length and frame rate.
- Warping error: warp frame
ttot+1using estimated optical flow and measure residual. The most direct temporal consistency measure, and it is bounded by flow estimator quality. - CLIPSIM: mean CLIP similarity between the prompt and each frame. Measures alignment, not motion, and a static image satisfies it completely.
- VBench (Huang et al., 2023): decomposes quality into sixteen dimensions including subject consistency, motion smoothness, and dynamic degree. The dynamic-degree axis exists because static output scores well on most other metrics — a model that outputs a still image can win a naive consistency benchmark.
That last point deserves emphasis. Consistency and motion are in direct tension, and any single-number video benchmark can be gamed by generating less motion.
Physics
Kang et al. (2024), How Far is Video Generation from World Model: A Physical Law Perspective, tested generalisation on synthetic scenes with known dynamics. In-distribution physics was learned well; out-of-distribution combinations failed, and the models fell back on case-based retrieval of similar training clips rather than any learned law. Scaling helped combinatorial generalisation and did not produce genuine extrapolation.
Treat claims that video generators are "world models" against that result.
Cost
Generating 16 frames at 576x1024 with an SVD-class model requires roughly 12 to 20 GB of VRAM depending on precision and attention implementation, and takes on the order of a minute on a modern data-centre GPU. Training runs are measured in thousands of GPU-days. This is the single largest practical gap between video and image generation, and it shapes which research questions get asked.
Papers
- Ho et al., Video Diffusion Models, 2022 — arxiv.org/abs/2204.03458
- Singer et al., Make-A-Video, 2022 — arxiv.org/abs/2209.14792
- Ho et al., Imagen Video, 2022 — arxiv.org/abs/2210.02303
- Peebles and Xie, Scalable Diffusion Models with Transformers, 2022 — arxiv.org/abs/2212.09748
- Blattmann et al., Align Your Latents, 2023 — arxiv.org/abs/2304.08818
- Blattmann et al., Stable Video Diffusion, 2023 — arxiv.org/abs/2311.15127
- Huang et al., VBench, 2023 — arxiv.org/abs/2311.17982
- Kang et al., How Far is Video Generation from World Model, 2024 — arxiv.org/abs/2411.02385
What to learn next
- Audio-visual learning — adding sound to moving pictures.
- Text to image — the single-frame case, where the machinery is easier to see.
- Latency and throughput — why generation cost decides what ships.