An 18-question multiple-choice test of generative AI knowledge showed acceptable statistical fit when researchers checked it on Estonian high school students held out from the test-building process. The evidence was strongest for measuring a shared underlying level of conceptual knowledge; the test was less precise among students at the lower end of that scale.
The researchers describe GenAIT as an objective measure intended primarily for group-level research. Its final 18 multiple-choice items cover technical, practical and human domains and target conceptual GenAI knowledge, making the instrument narrower than a test of every way a student might use or assess an AI system.
Building the questions
To build it, five members of the University of Tartu’s Artificial and Natural Intelligence Lab drafted, reviewed and refined the items through repeated rounds. Each item had four answer options and one correct response.
At the expert-review stage, eight people were contacted and five took part: three AI or data-science experts, one educational technologist and one psychologist. Relevance and clarity were rated on a four-point forced-choice scale.
That review led to the exclusion of Item 19. The pool moved from an initial 21 items to the final 18, while the average content-validity index rose from .96 to .99 and average clarity from .89 to .90. The analysis says expert review generally supported relevance and clarity, but did not establish exhaustive coverage of GenAI literacy.
A large held-out check
For the main validation, all Estonian high schools were invited and 100 agreed to take part. Of 8,698 respondents who completed the questionnaire, 1,266 were removed, leaving 7,432 students for analysis. Demographic data were unavailable.
The remaining respondents were randomly divided into a development group of 3,737 and a held-out validation group of 3,695. Researchers used the first group for item selection and IRT-model comparison, then kept the second separate for independent evaluation.
The main analysis combined confirmatory factor analysis, classical test theory and item response theory, examining the test’s internal structure, reliability and the way individual items functioned. In the item-response analysis, the best-performing model was a three-parameter model, which allows questions to differ in difficulty, their ability to distinguish among students and the chance of a correct answer by guessing.
Where the test held up
In development data, the 3PL model fit significantly better than simpler alternatives, with lower AIC and BIC scores. For the 18-item version, approximate fit figures were RMSEA .016, TLI .982, CFI .986 and SRMSR .020. But the exact-fit test was still significant (M2 = 226.383, 117 degrees of freedom, p < .001), so the model should be read as a useful approximation rather than a perfect description.
The held-out validation produced a similar picture. A one-factor analysis supported approximate unidimensionality — that is, the items were broadly measuring one underlying knowledge trait — and the 3PL fit was acceptable, with RMSEA .013, TLI .987, CFI .990 and SRMSR .021. No problematic local dependence was detected, although the exact-fit test again remained significant (M2 = 193.267, 117 degrees of freedom, p < .001).
Reliability, a measure of how consistently a test distinguishes among students, was .72 by the study’s marginal estimate and .69 by KR-20, with a 95% confidence interval of .68 to .71 for the latter. Precision changed across the score range: conditional reliability was .55 at the 5th percentile and .89 at the 95th percentile, and it exceeded .70 only between latent-trait values of -0.27 and 2.62, a span containing about 61% of students.
In practical terms, the test gave a more dependable read on students around the middle and toward the higher end of the measured trait than on students at the lower end. That uneven precision is why the authors present GenAIT as a promising tool for exploratory group-level research, not as a stand-alone basis for high-stakes individual classification.
At item level, validation estimates showed positive discrimination for every question, ranging from .59 to 4.23. Item difficulty ranged from -1.71 to 2.05, while pseudo-guessing estimates ran from 0 to .33. The spread shows that the questions varied in challenge and in how sharply they distinguished students, but it does not show that students can apply the knowledge in authentic tasks.
Links to AI use
A small pilot offered an early clue about how scores might relate to what the study calls magical perception of AI. Among 71 retained respondents, higher GenAIT scores were associated with lower magical perception: Spearman’s rho was -.35 for a 20-item version and -.30 for the final 18 items. The pilot’s 20-item reliability was KR-20 = .77, but the authors caution that the correlation may be unstable because of the small sample.
In the main survey, raw scores and model-based latent scores ranked students almost identically, with a Spearman correlation of .97; the researchers therefore used raw scores. GenAIT scores were negatively associated with general LLM-use frequency (rho = -.16) and school-related use (rho = -.19). After adjustment, however, the test was not significantly associated with perceived usefulness or ease of use.
Those findings describe relationships, not causes. Because the analysis was cross-sectional, it cannot show whether more frequent LLM use is linked to lower conceptual knowledge, whether lower knowledge is linked to more use, or whether an unmeasured third factor helps explain both.
The measure’s boundaries
Several limits remain. GenAIT measures conceptual knowledge rather than authentic prompting, verification or output-evaluation performance, and the main study did not include another objective AI-literacy measure or a performance-based criterion. With demographic data unavailable, the researchers could not assess whether the test works equivalently across subgroups.
The study therefore presents an initial validation of this instrument, not a universal yardstick. Future work will need to test it against authentic tasks, compare it with other objective measures, examine different languages and demographic groups, and update its content as GenAI evolves. On the evidence available here, its clearest role is measuring differences between groups of students, not deciding what any one student can or cannot do.
Paper data and sources
Original title: GenAIT: Development and Validation of an Objective Generative AI Literacy Test for High School Students
Authors: Brett Puppart, Kristjan-Julius Laak, Jaan Aru
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text