Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Text-to-Video AI Works Under the Hood

1From Prompt to Conditioning Signal2Latent Diffusion: Generating Video in Compressed Space3Spatio-Temporal Attention: Keeping Frames Consistent4Where the Pipeline Breaks: Failure Modes and Their Causes
Spatio-Temporal Attention: Keeping Frames Consistent

Linking the Same Place Across Frames

2 / 5
Now stack the frames and run attention along the other axis. Pick one spatial location, and look at the patch sitting there in frame one, frame two, frame three, all the way down.
0:00 / 0:00

Temporal attention runs along the time axis. For a given spatial location, the patch at that location in frame one, frame two, frame three, and onward are treated as a sequence, and each one attends to the others. A patch representing an eye in frame one can therefore pull in the appearance of that same eye in later frames, so the representation stays anchored to one identity instead of being re-decided from scratch each frame. This is what keeps a face recognizable as it turns and a shirt the same color as the person walks.

References

  1. [1]Video Diffusion Modelsarxiv.org
Previous2 / 5Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion