Preprint

AI method maps possible overlap among test questions

Preprint: An LLM-based score aligned more closely with psychometric patterns than BLEU and cosine measures in one item bank, while simulated adaptive tests showed a trade-off.

A preprint reports that an AI method built around a large language model (LLM) tracked residual correlations, or links between questions left after a graded-response model (GRM) was fitted, more closely than two conventional text-similarity measures. The LLM score correlated 0.50 with those residual correlations, compared with -0.10 for BLEU and 0.06 for cosine similarity.

The study used a single 36-item Experiences in Close Relationships response dataset. The source had 51,491 participants, of whom 35,278 remained after cleaning for age and missing responses. Researchers tested whether similarity-based groups of questions showed coherent measurement patterns and ran simulated computerized adaptive testing (CAT), in which item selection changes during a test.

A two-part score for similar questions

The framework used Claude Sonnet 4 and component-specific prompts to generate similarity scores on a 0-to-10 scale in two parts. Structured decomposition (S1) assessed the parts of an item, while semantic relatedness (S2) assessed shared meaning. It combined the dimensions with weights of 0.3 and 0.7, respectively, and fixed the model's API temperature at 0.

Across the item bank, total LLM similarity ranged from 6 to 26.60. Its mean was 15.41, with a standard deviation of 4.02, and it correlated 0.98 with S2 and 0.76 with S1. Its correlations with BLEU and cosine were 0.30 and 0.43, respectively.

The two component scores were moderately aligned, with a correlation of 0.66, and most pairwise differences fell between -5 and 5. The highest average component scores were for item format, at 7.97, and overlapping themes or contexts, at 6.81.

Clusters lined up with measurement patterns

Hierarchical clustering divided the 36 items into four groups. Similarity within the LLM-derived groups ranged from 19.2 to 22.0, compared with an overall mean of 15.41 and a standard deviation of 4.02. The partition was more balanced and differentiated than partitions based on BLEU or cosine.

The groups also showed a more interpretable pattern in discrimination, a parameter describing how strongly an item separates responses along the measured trait. Apart from outlier Item 35, two LLM clusters were mainly low-discrimination and two were mainly high-discrimination. BLEU- and cosine-based clusters were less clearly differentiated.

A multivariate test of the four GRM thresholds found a significant difference across the LLM clusters (F(12, 93) = 4.79, P < .001). The first and last thresholds differed significantly, while the two intermediate thresholds did not. BLEU and cosine clusters also produced significant results, but their P values were .042 and .017, respectively.

Adaptive testing exposed a trade-off

In the CAT simulation, the test length was fixed at 20 items, and similarity-constrained conditions used cluster bounds of 3 to 10 items. The researchers ran 50 replications with 2,000 simulated examinees per replication, comparing unconstrained maximum Fisher Information (MFI) selection with LLM-, BLEU- and cosine-based constraints.

The LLM-constrained condition had the lowest mean bias, 0.029, compared with 0.033 for unconstrained MFI, 0.033 for BLEU and 0.031 for cosine. But MFI had the lowest mean root mean square error (RMSE), a measure of estimation error, at 0.106. Mean RMSE was 0.115 for LLM, 0.113 for BLEU and 0.112 for cosine, while the LLM condition had the most concentrated RMSE across replications.

What remains untested

That scope matters. The validation used a single 36-item ECR item bank and Claude Sonnet 4, while the adaptive-testing results came from simulations rather than operational administrations. The method also used fixed weights of 0.3 for structure and 0.7 for semantic relatedness, so the study does not establish that those choices work equally well elsewhere.

Reported correlations and CAT condition means were not accompanied by confidence intervals or other inferential uncertainty estimates, leaving the precision of the differences unresolved.

Posted as a version-one arXiv preprint dated 25 Aug 2026, the work evaluates similarity metrics and psychometric proxies in one test bank, not direct examinee experience. It therefore does not establish that similarity constraints improve real examinee outcomes. Whether the pattern holds across other LLM architectures and item banks, or under different prompts and weights, remains open.

Paper data and sources

Original title: A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments
Authors: Jing Huang, Jihong Zhang, Hua-Hua Chang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.