A preprint describes a framework that ranked six AI systems against 18 workplace activities by comparing capability profiles. Under the report’s neutral settings, Gemini 3.1 Pro ranked first for every activity, with Gemini 3 Flash second and four other systems in a lower, overlapping group. The score is a comparative suitability estimate, not a calibrated probability that an AI system will complete a task successfully.
The framework uses a shared set of cognitive capabilities to describe both AI systems and workplace activities. It began with 18 core capabilities, each scored on a six-point demand scale from level 0 to level 5. Reliability screening removed Attention, Inhibitory Control and Prospective Memory, leaving 16 capabilities.
Putting AI and work on one map
After invalid items were removed, the final benchmark battery contained 19,535 items. A clustering step based on the items’ demand profiles produced eight composite dimensions. Capability estimates came from Measurement Layouts, a Bayesian item-response model that combines each item’s performance with its demand annotation.
To check the setup, the report simulated 20 agents whose capabilities were known. It selected soft-min pooling with a setting of 1 and a shared intercept, meaning a common baseline for overall levels. The exercise recovered profile shape with a correlation of 0.92. Recovery of overall levels improved from 0.12 with free intercepts to 0.98 with a shared intercept, while suitability recovery rose from 0.14 to 0.91.
That was an internal recovery test: the data were generated under the same modeling assumptions used for inference. It supports the selected operating setup, but it does not establish that the pooling behavior will hold for real workplace tasks or deployed agents.
What the profiles showed
The catalogue contained six AI systems from two developer families. Gemini 3.1 Pro had the highest aggregate estimate, at 3.70, followed by Gemini 3 Flash at 3.38; the other four systems were approximately 2.6 to 2.9. All six systems had similar profile shapes, with variation across dimensions exceeding variation between systems. Semantic Memory, Language and Social Cognition were strongest, while Action Planning & Simulation, Instrumental Reasoning and Object Permanence were weakest.
The workplace side came from a questionnaire that collected 539 responses across six job domains; 410 remained after quality control. The highest frequency-adjusted importance scores were for Problem solving at 2.04, Decision making at 1.78, Checking at 1.44, Researching at 1.35 and Computer use at 0.96.
Those scores reflect what respondents considered important, not a direct measurement of cognitive demand. Across occupational domains, the task-capability maps shared a core of Planning, Semantic Memory, Working Memory, Language and Procedural Memory. Domain profiles correlated from 0.53 to 0.77, with a mean of 0.63, and their main differences lay in secondary capabilities.
A comparison of company and online respondents also found substantial agreement in the reported profiles: cosine similarity was 0.846 for capabilities and 0.905 for work activities. The report flags an inconsistency in the Pearson figures between the main text and an appendix table, so that validation result warrants caution.
A ranking with boundaries
Using neutral mapping settings of p = 0 and s = 1, the framework put Gemini 3.1 Pro first across all 18 activities. Its log-suitability scores were approximately 4.1 to 4.6, compared with about 3.3 to 4.1 for Gemini 3 Flash. The other four systems formed a lower, overlapping group at roughly 1.7 to 3.4.
The broad ordering was generally stable when the analysis changed how much strengths could compensate for weaknesses and how sharply weights were applied. Nearly 12 of the 15 pairwise orderings remained unchanged, and each task showed between three and seven distinct rankings across the policy sweep. Gemini 3.1 Pro remained at the top for 17 of 18 activities, with Admin the exception at p = 2.
Company X supplied 35 responses after quality control. Its local map identified Communicating and Researching as the clearest deployment opportunities, while Problem solving and Computer use remained high-priority activities where the catalogue fell short. The report presents this as an illustrative example from a small, unevenly distributed respondent group.
The score is best used to compare priorities, not as a guarantee of performance. The questionnaire weights task importance rather than task demand, so suitability is comparative rather than calibrated. The benchmark is primarily text-based, and the analysis profiles foundation models in isolation while underrepresenting multimodal and agentic capabilities.
The document is an arXiv preprint dated 26 August 2026. Its acknowledgements state that Accenture supported the research.
Paper data and sources
Original title: Using profiles of cognitive capability to assess AI suitability for workplace tasks
Authors: Jonathan Prunty, Marko Tešić, Patrick Quinn et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text