Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Generative AI Creates Text, Images, and Videos from a Prompt

1The Shared Idea Behind All Generative AI2How a Prompt Becomes an Instruction3Generating Text: One Token at a Time4Generating Images: From Noise to Picture5Generating Video: Adding Time6Comparing the Three Modalities and Judging Outputs
How a Prompt Becomes an Instruction

Meaning Becomes a Position

2 / 4
Picture a map where every word has a location. Words used in similar ways land near each other, so cat and dog sit close, while cat and bicycle sit far apart. Real models use hundreds or thousands of directions instead of the two on this map, so no single direction means something as simple as alive or moving. But the geometry still carries meaning: distance reflects how similarly two words are used, and direction can capture relationships, which is why the step from king to queen looks like the step from man to woman. This is what turns your prompt from a string of symbols into a shape the model can work with. Words pointing the same way reinforce each other, and words pointing in opposite directions create a tension the model has to resolve.
0:00 / 0:00

Once the prompt is a sequence of tokens, each token is converted into a list of numbers that places it at a point in a space with many dimensions. You can think of this space as a map of meaning: two tokens that are used in similar contexts end up near each other, and tokens used in unrelated ways end up far apart.

A useful way to picture it is a two-dimensional map where one axis roughly separates living things from objects and another separates things that move from things that stay still. On such a map, "cat" and "dog" would sit close together, "cat" and "bicycle" would sit farther apart, and "skateboard" would land near other rideable objects. Real models use hundreds or thousands of dimensions rather than two, so no single axis means anything as clean as "alive" or "moving". The directions are learned from data and are not individually interpretable. What matters is the geometry: distance reflects similarity of use, and direction can encode relationships, so the step from "king" to "queen" resembles the step from "man" to "woman".

This is the bridge from text to conditioning. Because meaning is a position, a prompt is not a string of symbols to the model but a shape traced through meaning space. Words that pull in the same direction reinforce each other; words that pull in opposite directions create tension the model has to resolve.

Previous2 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion