A comparison of human perception, macaque IT activity used as a neural benchmark, and artificial neural networks (ANNs) identified a sharp weakness in conventional vision systems: they decoded motion direction well from ordinary-looking, naturalistic clips, but performance fell when an object's recognizable appearance was removed. Humans and macaque IT responses retained motion-direction information across that change, while conventional image and video models transferred it poorly.
The study drew on online perception tests with 107 MTurk participants, a macaque sample of five adult male rhesus monkeys, and representations from 15 video-model and 14 image-model architectures. The ANNs were publicly available pretrained systems used without additional training or fine-tuning.
People performed strongly on all three tasks: mean accuracy was 0.94 for identity, 0.91 for motion direction and 0.82 for speed, with standard deviations (SDs) of 0.02, 0.07 and 0.06, respectively.
Video-model representations also yielded higher mean decoding than image-model representations: 0.73 (SD 0.07) versus 0.66 (SD 0.07) for identity, 0.65 (SD 0.06) versus 0.58 (SD 0.02) for motion direction, and 0.67 (SD 0.04) versus 0.63 (SD 0.02) for speed. The reported differences were statistically significant. In a matched condition with no temporal variation, mean accuracy was lower by 0.11 for identity, 0.10 for direction and 0.06 for speed.
Appearance was the harder test
To test whether motion information transferred across appearance, motion-direction decoders were trained on naturalistic representations and tested without refitting on held-out naturalistic and paired appearance-free representations. That made the comparison about whether the same motion signal survived a change in appearance.
People retained high motion-direction accuracy without appearance cues: 0.90 (SD 0.06), close to 0.92 (SD 0.04) with appearance-based stimuli. Conventional models did not show the same transfer. Video models fell from 0.78 (SD 0.09) on naturalistic videos to 0.53 (SD 0.05) on appearance-free versions; image models fell from 0.71 (SD 0.05) to 0.47 (SD 0.02).
Some model classes held up better. Optic-flow models, built around patterns of movement, scored 0.91 (SD 0.01) on naturalistic videos and 0.79 (SD 0.01) on appearance-free videos. Predictive world models scored 0.92 (SD 0.02) and 0.85 (SD 0.04), respectively. SAM2 dropped from 0.88 (SD 0.03) to 0.28 (SD 0.03).
The brain-model gap appeared in timing
Neural recordings showed a similar pattern. Across 122 IT sites from two macaques, motion-direction decoding was above chance at 0.61 (standard error 0.02) for held-out naturalistic responses and 0.58 (standard error 0.02) for appearance-free responses. The distributions did not differ significantly.
To compare how information was organized, the researchers used noise-normalized Centered Kernel Alignment, or CKA, a score of similarity between two representation patterns. The analysis used 200 moving-object videos, 404 IT neurons and 15 video ANNs.
Early CKA was similar for models using spatial information alone and models integrating time: 0.35 (SD 0.06) versus 0.34 (SD 0.06), a difference that was not statistically significant. Later, CKA was higher for the time-integrating models: 0.30 (SD 0.05) versus 0.28 (SD 0.04), a statistically significant difference.
Among the tested approaches, predictive world-model architectures had the highest reported IT correspondence: CKA was 0.30 (SD 0.02) for spatial versions and 0.36 for spatiotemporal versions. But this comparison did not isolate predictive learning, because these models also differed in scale, training data and implementation.
Another mismatch concerned timing. IT responses changed more at about 120 milliseconds, where the median temporal-change score was 0.27 (standard error 0.006), than at about 330 ms, where it was 0.11 (standard error 0.007); the early-to-late median difference was 0.16 and statistically significant. The neural representation became more temporally stable later, while current video models did not match this transformation.
At later responses, the difference between appearance-related and motion-related information was 0.34 in IT, compared with 0.70 in video-model features. Spatiotemporal models narrowed that gap relative to spatial models, a reported statistically significant difference, but remained unlike IT.
A narrower lesson for AI
The authors interpret the pattern as a candidate principle of robust dynamic vision: motion should become integrated into object representations in a way that survives changes in appearance. They identify predictive world modeling as a promising route, but the result does not isolate predictive learning because the model groups differed in scale, training data and implementation.
The practical lesson is narrower: high scores on naturalistic videos are not enough. Systems should be tested on appearance-disrupted versions and compared with neural dynamics. Because the ANNs were used off the shelf in controlled dynamic-object tasks, the findings do not establish robustness in unconstrained real-world settings.
Paper data and sources
Original title: Primate vision reveals a missing principle for robust dynamic AI
Authors: Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text