Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How Generative AI Creates Text, Images, and Videos from a Prompt

1The Shared Idea Behind All Generative AI2How a Prompt Becomes an Instruction3Generating Text: One Token at a Time4Generating Images: From Noise to Picture5Generating Video: Adding Time6Comparing the Three Modalities and Judging Outputs
Generating Text: One Token at a Time

Why Small Steps Add Up to a Paragraph

3 / 4
Look at the traced sentence and notice what changes between rounds. After "The capital of France is", the field of candidates is narrow and "Paris" dominates. Once "Paris" is in the context, the question the model is answering has quietly changed: it is no longer what the capital is, but what usually follows that completed statement. So the options shift to punctuation or a descriptive clause. That shift is the mechanism behind coherence. The model is not holding a plan for the sentence; it is responding to a context that its own earlier choices keep rewriting. The same feedback explains why an early mistake tends to survive, since every later step is scored against the mistake rather than against the facts.
0:00 / 0:00

Every token the model produces is added to its own input for the next round. That feedback is what turns independent local guesses into a continuous piece of writing: each choice shrinks the set of continuations that still fit, so the text acquires a direction it was never given in advance.

Tracing a short sentence

Take the opening "The capital of France is". The next-token distribution is dominated by "Paris", so that token is chosen. The context is now "The capital of France is Paris". The next distribution is no longer about capitals; it is about what typically follows that completed statement, so options like a period, a comma, or a clause such as ", which sits on the Seine" become plausible. Each step is scored against a slightly longer history, and the sentence's shape is decided along the way rather than at the start.

Because later steps are conditioned on earlier ones, an early wrong turn tends to be carried forward rather than repaired. The model is not re-reading the whole answer and checking it against the world; it is continuing from whatever is already there.

Previous3 / 4Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion