A defensible judgment comes from a fixed order of inspection, not from the first striking clue you notice. The routine is: check context first, scan the whole frame, zoom on high-risk regions, check motion if the media moves, and read file-level signals last.
Order matters because clues differ in how much they should move your confidence. Context frames what the other signals mean — the same odd hand is more suspicious in a photo claiming to be a news capture than in a stylized illustration. A whole-frame scan catches implausible content before you get lost in detail. Zooming on high-risk regions concentrates attention where generation errors cluster: hands, teeth, ears, eyes, hair edges, text, logos, and reflections. Motion checks apply only to video, where flicker, identity drift, and broken physics appear across frames rather than within one. File-level signals — metadata, content credentials, and watermarks — come last because they are asymmetric: a verified credential naming a capture device is strong positive evidence, while a missing metadata block proves nothing, since most platforms strip it on upload.
At the end you state a confidence level, not a verdict. Something like: "I lean toward AI-generated, moderate confidence, because the reflections do not match the light source and the fingers merge, but the metadata is absent and that tells me nothing either way." Naming the evidence and the level separately keeps the judgment honest and reviewable.