Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Generative AI Creates Text, Images, and Videos from a Prompt

1The Shared Idea Behind All Generative AI2How a Prompt Becomes an Instruction3Generating Text: One Token at a Time4Generating Images: From Noise to Picture5Generating Video: Adding Time6Comparing the Three Modalities and Judging Outputs
The Shared Idea Behind All Generative AI

Where the Patterns Come From

2 / 4
Picture the training material as a wide stack of examples on the left, and a small block of internal settings on the right. The arrows between them are not copying. They are extracting what repeats: which words tend to follow which words, how edges and textures tend to sit next to each other, how motion usually unfolds. The stack stays behind. Only the distilled pattern survives into the settings. That is why we can say the model learned the regularities without memorizing the library, and it is also why the model can respond to a request it has never literally seen before.
0:00 / 0:00

A model does not arrive knowing anything. Before it can generate, it is trained on a very large collection of examples: books, articles, captioned photographs, video clips, and so on. During training the model is shown these examples again and again and its internal settings are adjusted until it becomes good at anticipating what tends to appear together.

What gets stored is not the examples themselves. It is the regularities inside them. In text, that means which words tend to follow which others, and how sentences and paragraphs are usually shaped. In images, it means how edges, textures, colors, and object shapes tend to be arranged. In video, it adds how things usually move and how a scene tends to change from one moment to the next.

This is why training is often described as compression. A library's worth of material is squeezed into a compact set of internal settings that capture the recurring structure and discard the incidental detail. The model keeps the pattern and lets go of the specific instance. When it later generates, it draws on those captured regularities rather than looking anything up.

Previous2 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion