A preprint asks what kinds of generality can arise from visual learning and sketches several possible routes to general intelligence. It does not report an experiment showing that any route works. Instead, it uses a multi-perspective conceptual approach to clarify open questions rather than propose a single model or benchmark, bringing together contributors from diverse standpoints and affiliations. The contributors do not converge on a single definition of visual general intelligence or on one model, learning objective or representation.
Three objectives at the center
At the center of the paper's hypothesis are three complementary objectives: sequential prediction, open-ended generation and reconstruction. In ordinary language, that means learning from what unfolds over time, producing open-ended possibilities and rebuilding useful structure from visual experience. The paper treats their convergence as a possible path toward visual intelligence and eventually general intelligence, not as a demonstrated result.
That proposal also depends on what a system can retain and revise. The Spatial AI perspective calls for representations that are persistent but adaptable: they should change with new information and changes in the world while supporting prediction and planning. A lifelong-learning perspective says a visually intelligent system should learn throughout its visual lifetime, detect genuinely new events, decide what should change and retain useful knowledge. These are design requirements proposed in the paper, not capabilities measured by a common experiment.
The case for memory, senses and action
Several contributors would also broaden the problem beyond vision alone. One perspective holds that visual intelligence should be multimodal, generative and computationally efficient, combining vision with touch, audio and proprioception, the body's sense of position and movement. The embodied-intelligence perspective treats action as both an output of visual intelligence and a way to obtain visual evidence. Another position says visual intelligence should recover compositional physical structure so its representations support editing, simulation, verification and action.
Video models are another possible route in the paper. One contributor argues that developing them as visual foundation models could make visual general intelligence achievable. But that is one contributor's perspective, not a conclusion shared by the paper; the wider synthesis leaves the definition, model and learning objective unsettled. The paper also does not show that video models generalize reliably to arbitrary unseen tasks or unfamiliar real-world environments.
The test would need to be broader
If visual intelligence is to be judged, the paper says fixed-task performance will not be enough. A broader test would look for transfer, persistent and revisable world knowledge, creative and counterfactual reasoning, active information seeking, continual adaptation, spatial and physical consistency, and reliable action in unfamiliar situations. The paper does not report a validated benchmark covering those dimensions; it sets them out as targets for future evaluation.
Creativity is one proposed test. The evaluation calls for visual outputs that are coherent, structurally diverse across generations and original relative to the system's training experience. That is a proposed standard rather than an experimental result in the preprint.
A research agenda, not a result
Taken together, the perspectives amount to a research agenda rather than a claim that visual general intelligence exists. Existing systems demonstrate important components of visual intelligence, but not their integration into a general, continually adaptable visual system. The paper does not show that visual experience alone is sufficient for general intelligence or that any proposed learning objective produces it.
The document identifies itself as arXiv:2608.25924v1 [cs.CV], dated 26 Aug 2026. The work was conducted by AIST researchers under FRONTia and supported by JST through ASPIRE, Grant Number JPMJAP2518.
Paper data and sources
Original title: Visual General Intelligence: A White Paper
Authors: Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text