Preprint

AI screening settings linked to different test items for review

Preprint: A computational study found that embeddings, structural methods and eligibility rules were associated with different evidence and wording in AI-generated Big Five candidate forms.

Different choices in an AI-assisted screening pipeline were associated with different sets of AI-generated Big Five self-report items available for expert review, even when the source populations were held fixed, a computational analysis reports. The differences extended from the evidence attached to individual items to the wording of completed candidate forms.

Different representations, different evidence

The primary computational design combined two open-weight generators, two generation packages, five traits and 20 independently seeded repetitions. Each generation task produced 80 selected items, for 400 tasks and 32,000 selected item occurrences. Each source population then passed through five embedding configurations—different representations of the same wording—followed by a shared UVA step and two structural evaluations, TMFG and graphical lasso.

The five embedding configurations attached different construct evidence to identical items. On behavioral indicators, Qwen 8B had the highest average scaled target advantage—1.417 versus 1.408 for Qwen 4B—while target-first rates were 92.16% and 92.52%, respectively. Yet the near tie in averages masked item-level movement: target-first status varied across configurations for 35.2% of AI-GENIE source items and 25.5% of construct-indirect items.

The source package was also associated with a different anchor profile. Under Qwen 8B, construct-indirect source items had a .222 higher behavioral-indicator target advantage than AI-GENIE-source items; the reported interval was .202 to .242.

Broad agreement, local divergence

Source-population choice moderated the sensitivity of the structural comparison. Qwen-generated source populations showed greater overlap in retained items and greater initial community correspondence—how similarly items were grouped—than Gemma-generated populations across ten embedding pairs and two structural methods. In the highlighted Qwen 4B-Qwen 8B TMFG comparison, the differences were .080 in Jaccard, the exact-overlap measure, and .065 in initial AMI, the community-correspondence measure, with reported 95% task-bootstrap intervals of .054 to .106 and .048 to .083.

Broad similarity was not the same as local agreement. Qwen 4B and Qwen 8B had geometry agreement of .874, but their TMFG retained-set Jaccard was .565 and their initial community AMI was .720. In plain terms, the representations broadly ranked items alike while still selecting different items and producing different early groupings.

Structural reduction—the step that narrows the item set—usually raised the share of exact four-community solutions from 32.93% before reduction to 73.02% after it. That pattern was not universal: AMI decreased in 59 of 1,999 TMFG contexts and 217 of 2,000 graphical-lasso contexts. Coverage could also be uneven; BGE-M3/TMFG had 119 empty intended-content cells among 1,596 evaluable cells, compared with 27 of 1,600 for Qwen 8B/TMFG.

Complete forms, different wording

Study 2 compared two eligibility rules for candidate forms. Inclusive eligibility accepted an item retained by either structural method; agreement eligibility required retention by both. The agreement rule excluded 46,121 of 131,073 complete inclusive-eligible item occurrences, a 35.2% contraction of the eligible pool.

Both policies nevertheless filled all 40 primary and 40 alternate positions across all 20 content cells, but matched primary forms shared a median of 29 of 40 items. Each policy had a median of 11 unique primary items, and alternate forms differed by a median of 21 items.

The cross-embedding comparison showed much less exact wording overlap than the within-configuration policy comparison. Across 36 of 40 evaluable inclusive-form embedding pairs, matched primary forms shared a median of six of 40 statements, with a Jaccard score of .081. Within one configuration, inclusive and agreement forms shared 29 of 40 statements, with Jaccard .569.

Across generation-task resampling, 19,359 of 20,000 planned paired draws were evaluable. Every evaluable draw filled all 40 primary and 40 alternate positions, while configuration-specific median wording Jaccard ranged from .379 to .778. These results describe variation across computational task mixtures, not uncertainty from respondents.

Before validation

These are pre-review computational findings, not evidence that the items work as a measure. All evaluator evidence came before respondent-data validation and does not establish item functioning, reliability, respondent dimensionality, measurement invariance or validity for a proposed use.

The work is an arXiv preprint, version arXiv:2608.23766v1, dated 24 August 2026. It received in-kind computational support from Advanced Research Computing at the University of Michigan, Ann Arbor.

Paper data and sources

Original title: What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development
Authors: Christopher Brooks
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.