Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Text-to-Video AI Works Under the Hood

1From Prompt to Conditioning Signal2Latent Diffusion: Generating Video in Compressed Space3Spatio-Temporal Attention: Keeping Frames Consistent4Where the Pipeline Breaks: Failure Modes and Their Causes
Latent Diffusion: Generating Video in Compressed Space

Destroying the Latent on Purpose

2 / 5
Now that the video lives as a latent cube, the model has to learn how to build one from scratch. Training starts by doing the opposite: taking a clean latent and destroying it.
0:00 / 0:00

Forward noising is the process that turns a clean latent into pure noise. At each timestep a small amount of Gaussian noise is added according to a noise schedule, and after enough steps the original structure is completely gone. The schedule controls how much noise is added at each step — early steps add little, later steps add a lot. This process is fixed, not learned: it defines the training target, because the model's job is to learn how to reverse it. The final state is a latent that carries no information about the original video, only noise.

References

  1. [1]Denoising Diffusion Probabilistic Modelsarxiv.org
Previous2 / 5Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion