Text is the single most reliable clue in a still image, and the reason is structural. Letters are an exact, learned, high-constraint pattern. A model that is reconstructing plausible-looking marks rather than retrieving a stored character will produce marks that look like letters at a glance but do not resolve into real words when you read them.
What to check:
Read the text, do not skim it. Zoom in and read every word, including small print on signs, labels, T-shirts, and book spines. Common failures are letters that merge, extra strokes, missing strokes, characters that shift alphabet mid-word, and words that are simply not words.
Check repeated text. If the same word appears twice in the image, compare the two instances. A model often renders them slightly differently, because it generated each occurrence independently.
Check logos. A real logo is a fixed shape. AI logos are often near-misses: the right colors and general form, but a distorted letterform, a wrong number of elements, or a shape that changes between two appearances of the same logo.
Check the background. Backgrounds receive less attention than the subject, so they fail more often. Look for objects that make no sense in the scene, furniture that merges into a wall, people in the distance with no faces or with faces that dissolve on close inspection, and architectural lines — window frames, floor tiles, brick courses — that bend or stop for no reason.
Why text fails so badly: a model trained on images learns the visual statistics of letter-like shapes, not a spelling system. It can produce something that satisfies the local statistics of "text" without ever committing to a specific string of characters. That is why the failure is not random noise but confident-looking pseudo-text.
This page closes the still-image pass. You now have four families of clues — anatomy, light, texture, and structured detail — and the next chapter applies the same logic to motion, where the model must stay consistent across time as well as across space.