A generated clip can look technically clean yet drift away from what the prompt described. This is prompt misalignment, and it traces back to conditioning: the per-token embeddings that carry the prompt's content into the generator were too weak, too diluted, or misaligned with the spatial regions they should have steered. Because conditioning enters through cross-attention, a token that never wins attention in a region leaves that region free to be filled by whatever the model's prior prefers. Raising guidance strengthens the conditional signal, but it is not a fix for a prompt the encoder represented poorly.
How Text-to-Video AI Works Under the Hood
Where the Pipeline Breaks: Failure Modes and Their Causes
When the Video Ignores the Prompt
1 / 5
Here is a clip that looks perfectly clean. Sharp, well lit, no flicker. And yet it has ignored half of what the prompt asked for.
0:00 / 0:00