Skip to content
Learn Motion
ExploreHow it worksMembership
Log in
Learn Motion

How AI Sees: A Conceptual Guide to Image and Video Understanding

1What a Machine Actually Sees2From Pixels to Patterns: Features3How Deep Networks Learn to See4Beyond Labels: Locating and Describing What Is Seen5Adding Time: Understanding Video6How These Systems Learn and How We Judge Them
Adding Time: Understanding Video

Two streams: appearance and motion

3 / 5
The diagram shows two pathways running side by side. The upper one takes raw frames and learns what is present; the lower one takes flow fields and learns what is moving. Each produces its own summary, and the summaries merge into one decision. The reason to keep them apart is that they break in different ways. Appearance is reliable when the camera is still but struggles when the same object appears in many poses. Motion is reliable when the movement is distinctive but struggles when the camera itself moves, because camera motion contaminates the flow. Seeing both lets the model lean on whichever is more trustworthy for this clip.
0:00 / 0:00

A temporal model needs to combine two kinds of evidence. The first is appearance: what the scene looks like in each frame, which is exactly what the image-side machinery from earlier chapters already extracts. The second is motion: how the scene is changing, which optical flow supplies.

The standard conceptual arrangement is two parallel pathways. One pathway processes a small stack of raw frames and learns what objects and scenes are present. The other processes a stack of flow fields and learns what movements are happening. Each pathway produces its own summary of the clip, and the two summaries are merged into a single decision.

The reason for keeping them separate is that they fail in different situations. Appearance is reliable when the camera is still and the object is clear, but it struggles when the same object appears in many poses. Motion is reliable when the movement is distinctive, but it struggles when the camera itself is moving, because camera motion contaminates the flow. A model that sees both can lean on whichever is more trustworthy for the clip in front of it.

Previous3 / 5Next

Learn Motion

Generate a course. Learn it properly.

Operated by Wuhan Daoyin Technology Co., Ltd.

Contact: [email protected]
Privacy PolicyTerms of Service

© 2026 Learn Motion