Pixel values encode brightness and color only. They never encode what the picture is of. Any statement like "this is a dog" is an inference drawn from patterns in the numbers, not a fact read out of them.
Meaning lives in the arrangement, not the values
Because meaning is not stored, it has to be recovered from structure. Sharp changes in brightness along a line suggest an edge; repeating local patterns suggest texture; characteristic combinations of edges and textures tend to accompany particular objects. These are statistical regularities learned from many examples, so a recognition result is a probability judgment, not a certainty.
The same object under different lighting, viewpoint, or distance produces a very different grid of numbers. The object is unchanged; the data is not. This is a central reason recognition is difficult.