A video model does not paint one continuous scene. It produces a sequence of frames, and each frame is generated with only partial reference to its neighbors. When that reference is weak, fine detail re-rolls slightly from frame to frame even though the overall scene looks stable. This is temporal flicker: texture, edges, and small features that shimmer or breathe rather than staying fixed.
A related failure is morphing, where a shape does not merely shimmer but gradually changes into something else. A hand can gain a finger over a second, a pattern on a shirt can rearrange itself, or the grain of a surface can flow like liquid. The mechanism is the same in both cases: the model is reconstructing detail locally in each frame instead of tracking one persistent object through time.
To check for this, pause the clip and step through consecutive frames rather than watching at normal speed. Look at a small, high-detail region — hair, fabric weave, a patterned background, the edge of a face — and compare it across three or four frames. In real footage, that region stays essentially identical between adjacent frames. In generated footage, it often shifts, blurs, or re-forms. Flicker is most visible in areas the model has to invent rather than areas it can copy from a strong reference, so textured backgrounds and fine edges are the best places to look.