Forward noising is the process that turns a clean latent into pure noise. At each timestep a small amount of Gaussian noise is added according to a noise schedule, and after enough steps the original structure is completely gone. The schedule controls how much noise is added at each step — early steps add little, later steps add a lot. This process is fixed, not learned: it defines the training target, because the model's job is to learn how to reverse it. The final state is a latent that carries no information about the original video, only noise.
How Text-to-Video AI Works Under the Hood
Latent Diffusion: Generating Video in Compressed Space
Destroying the Latent on Purpose
2 / 5
Now that the video lives as a latent cube, the model has to learn how to build one from scratch. Training starts by doing the opposite: taking a clean latent and destroying it.
0:00 / 0:00