A still image records appearance at one instant. An action is a pattern of change over an interval, so the same appearance can belong to several different actions. Recognizing the action requires comparing moments, not inspecting one.
Two readings of the same frame
Take a frame showing a person leaning forward with both feet off the ground. Read alone, it is ambiguous: a jump, a stumble, or a dive all produce a similar pose. Add the next half-second and the ambiguity collapses — the body either rises and lands, or drops and catches itself. The frame did not change; the surrounding moments supplied the meaning.
Classifying each frame independently and averaging the results is a common first attempt, and it is not the same as understanding video. It can tell you that people and a ball are present, but not that someone is throwing. Order and change are lost in the averaging.