Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Generative AI Creates Text, Images, and Videos from a Prompt

1The Shared Idea Behind All Generative AI2How a Prompt Becomes an Instruction3Generating Text: One Token at a Time4Generating Images: From Noise to Picture5Generating Video: Adding Time6Comparing the Three Modalities and Judging Outputs
Generating Images: From Noise to Picture

Why Start From Noise

2 / 4
Watch the starting grid. It is pure static, random color in every pixel, and there is no picture hidden in it. Now watch what happens as the process runs: structure appears gradually, and the static resolves into a recognizable scene. The important point is the direction of travel. The model is not cleaning up a blurry photo or uncovering something that was already there. It is adding structure where none existed. Notice also that the animation runs twice from two different random grids. Same prompt, different starting static, and you get two different images that both fit the prompt. That is the same idea you saw with text: the model has preferences, but the starting randomness decides which of the many acceptable results you actually get.
0:00 / 0:00

The starting point is a grid of random values, usually called noise. Every pixel gets an arbitrary color, so the initial image looks like television static.

Starting from randomness looks wasteful, but it removes a problem. If generation began from a blank white canvas, the model would have to decide where every object goes, in what order, and how to keep later additions from overwriting earlier ones. If it began from an existing photo, it would have to detect and undo whatever structure was already there. Noise has neither problem: it contains no objects to preserve and no structure to remove. The model is free to impose structure from scratch.

Randomness also serves a second purpose. Because the starting noise differs on every run, the same prompt does not have to produce the same picture. The prompt constrains what the final image must look like, while the noise supplies the variation in pose, layout, and detail. This is the image-side version of the sampling variation you already saw in text generation: the model's preferences are fixed, but the path taken through them is not.

One clarification: the noise is not the picture in disguise, waiting to be revealed. Nothing in the random values corresponds to the final content. The structure is introduced by the model during refinement, not extracted from the starting grid.

Previous2 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion