Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Text-to-Video AI Works Under the Hood

1From Prompt to Conditioning Signal2Latent Diffusion: Generating Video in Compressed Space3Spatio-Temporal Attention: Keeping Frames Consistent4Where the Pipeline Breaks: Failure Modes and Their Causes
Latent Diffusion: Generating Video in Compressed Space

Pushing Toward the Prompt

4 / 5
There is a catch. If the denoiser only ever sees the prompt, it can drift. So at every step, it runs twice: once with the prompt conditioning, and once with no conditioning at all.
0:00 / 0:00

Classifier-free guidance strengthens prompt adherence by running the denoiser twice at each step: once with the prompt conditioning and once without it. The two noise predictions are combined by extrapolating away from the unconditional one, using a guidance scale. A scale of one means no guidance — the output is the conditional prediction alone. Higher scales push the result further toward the prompt, which improves alignment but reduces diversity and can over-saturate colors or freeze motion. Lower scales allow more variation but drift from the prompt. The guidance scale is therefore a direct trade-off knob between prompt adherence and naturalness.

References

  1. [1]Classifier-Free Diffusion Guidancearxiv.org
Previous4 / 5Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion