Preprint

Multimodal AI Holds a Narrow Edge Over Text in Interview Test

This preprint describes a 646-participant dataset in which personality ratings from targeted questions were more reliable, while the best text model stayed close to the strongest multimodal system.

The best multimodal model, meaning a system that uses more than one type of interview signal, had only a modest edge over the strongest text-based model in a benchmark built from asynchronous video interviews. HFUT-VisionXL recorded an average mean squared error of 0.123, compared with 0.125 for PersonalityLLM, the best text-based model. Mean squared error was the benchmark's measure of prediction error, with lower scores representing smaller errors.

The result comes from AVI-Personality, a dataset introduced for personality and job-related competency assessment from asynchronous video interviews. It contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. The study pairs those interviews with self-reported personality scores, observer personality ratings and recruiter competency ratings.

Participants were recruited through Prolific. Of 793 recruits, 646 remained after exclusions involving consent, attention, response quality, audio and rater compliance. The final sample included 309 men, 309 women and 28 non-binary participants. Mean age was 36.69 years and mean work experience was 15.96 years.

Six questions set the comparison

Each participant answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. The targeted questions were intended to make particular traits more diagnostic in an answer or behavior.

Observer personality ratings came from 12 raters after nine hours of training, with at least three personality psychologists rating each participant. Five professional recruiters rated four job competencies and one overall interview-performance score after viewing all six questions, and each participant received at least two recruiter ratings on a five-point BARS scale.

Reliability varied across the ratings

To assess rating reliability, the analysis separated variation due to participants, raters and residual noise. It then reported two intraclass correlation coefficients, or ICCs: one for absolute agreement between raters and one for consistency.

Personality ratings were more reliable than competency ratings, and ratings from targeted questions were more reliable than ratings from generic questions. For Extraversion, the mixed absolute-agreement and mixed-consistency scores were 0.817 and 0.881 for targeted questions, versus 0.763 and 0.852 for generic questions. Integrity scores were 0.324 and 0.525.

Targeted questions generally showed stronger alignment between participant self-reports and observer ratings. The correlations were 0.42 for Extraversion, 0.41 for Conscientiousness, 0.25 for Agreeableness and 0.22 for Honesty-Humility. For generic questions, Extraversion was 0.38 and Emotionality was 0.37. These figures describe alignment between two rating sources; they do not show that the question format caused the differences or establish definitive construct validity.

What the recruiter scores showed

All four recruiter-rated competencies were positively associated with overall interview performance. Social versatility was the strongest individual predictor, with a beta coefficient of 0.405, and the regression model explained 87.1% of the variation in the interview-performance score. The predictors and outcome came from the same recruiter-rating process, so this is an internal association within the dataset.

Observer-rated personality models explained 17.7% to 26.6% of the variance in outcomes for personality-targeted questions and 18.1% to 27.1% for generic questions. Models based on self-reported personality explained 2.7% to 11.0%. These were within-dataset regressions without an independent criterion, so the figures are not a test of later job performance.

The benchmark's boundaries

The fairness analysis used separate demographic regression models for age, education, gender and ethnicity. It reported false-discovery-rate-adjusted p-values and partial eta-squared effect sizes. Generic-question observer ratings showed more frequent demographic associations than personality-question observer ratings; gender and education were the most consistent predictors of competency ratings, while ethnicity had limited effects. For generic-question observer ratings of Emotionality, the gender effect had a partial eta-squared of 0.177. These are subgroup associations, not deployment-level fairness tests.

For model evaluation, participants rather than individual clips were split into 70% training, 10% validation and 20% testing groups. The groups contained 452, 64 and 130 participants respectively, and joint sampling preserved demographic and work-experience distributions. Mean squared error was the evaluation metric. At the trait level, EMMR had the best Conscientiousness result at 0.093, while PersonalityLLM led for Honesty-Humility at 0.080 and Agreeableness at 0.059.

The question comparison has important limits. Question order was fixed, and the generic and targeted conditions were not randomized, so the results do not establish that targeted questions caused better measurement. Only four HEXACO traits had dedicated personality questions, competency ratings had lower absolute agreement, and the benchmark did not cover all possible modeling strategies.

The dataset describes a simulated management traineeship application, and the evidence concerns annotation properties and model comparisons within that setting. It should not be read as a hiring recommendation or proof of future workplace performance. The document is an arXiv version 1 preprint dated 26 August 2026, and its abstract states that the dataset was released at the project repository.

Paper data and sources

Original title: AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews
Authors: Tianyi Zhang, Jinwenxi Shang, Antonis Koutsoumpis et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.