A video adds one thing to an image: repetition over time. It is stored as an ordered sequence of images called frames, played fast enough that the eye reads them as continuous motion. A clip at 30 frames per second shows 30 images for every second of footage, so a 10-second clip contains 300 frames. The order is part of the data, not a convenience. If you shuffled those frames, every individual image would still look perfectly normal, yet the motion would be destroyed and the clip would become meaningless. This is the first place where video differs from a still image in data terms: an image is one grid, while a video is a list of grids plus a time index that says which grid comes first, second, third, and so on. The frame rate fixes the real time gap between consecutive frames.
How AI Sees: A Conceptual Guide to Image and Video Understanding
What a Machine Actually Sees
Video Is Frames in Order
3 / 4
Watch the frames advance one after another. Each one is a complete image, exactly like the grids we just looked at. What makes it a video is the order and the speed. At thirty frames per second, ten seconds of footage is three hundred separate images. Now imagine shuffling them. Every single frame still looks fine on its own, but the motion is gone and the clip stops making sense. That is the crucial difference in data terms: a still image is one grid, a video is a list of grids with a time index attached.
0:00 / 0:00