An evaluation of open-ended AI beliefs found that the ranking of confidence methods can flip when outputs that do not match a finite reference set are treated as false. In a controlled comparison, a frozen source-prior rule looked better under reference-derived labels, but native confidence looked better when the same contents were judged for literal truth by blinded human auditors.
To make the comparison, the study held audited contents and associated information fixed while changing only the label source. The local literal-truth calibration sample contained 259 usable items out of 280 annotated items. Three independent external volunteers handled the local and OpenToM audit waves while blinded to method, confidence, source, matching and applicable hypotheses.
A label switch changed the score
The comparison used Brier risk, a paired score for judging the two confidence rules. Under finite-reference labels, the frozen source-prior rule, a fixed confidence baseline, lowered Brier risk relative to native confidence by 0.227. Under adjudicated literal-truth labels, it raised risk by 0.152. The same rule therefore moved from the better-ranked option to the worse-ranked option when the label source changed.
The reversal appeared in every authored condition examined: the reference-label difference was negative and the adjudicated-truth difference was positive in 6 of 6 conditions. That consistency held across the six conditions used for the test, but it remains a result from those authored conditions.
A separate released NQ-open analysis covered 301 questions. It compared an average-confidence rule with native confidence using instance-level calibration error, or ICE, the study's measure of how confidence compared with correctness for individual outputs. The average-confidence-minus-native gap changed from -0.045 under exact-match labels to +0.074 under human-correctness labels, and both reported intervals excluded zero. Average confidence thus looked better on the reference labels but worse on the human-correctness labels.
The missing matches were often true
The analysis also found a sharp change in what counted as a positive label. Reference construction reduced the weighted positive-label rate from 0.783 under adjudicated truth to 0.295 under reference labels. The paper calls this prevalence collapse: the reference process counted far fewer of the tested beliefs as positive.
Among reference-unmatched local beliefs, 73% were judged literally true. That means the unmatched category was not a reliable stand-in for falsehood in this audit, even though the reference-labeling rule treated it that way.
When the team separated the distortion into components, omitted truths dominated it. False-positive matches pushed in the opposite direction only weakly in the local analysis and were zero on OpenToM. The main failure identified in these comparisons was therefore the loss of true beliefs that were absent from the finite references.
The pattern held in a second audit
An additional OpenToM audit reproduced the direction of the result. Among 209 usable beliefs, between 90% and 96% of unmatched output was judged literally true.
The paired Brier gaps shifted from -0.381 and -0.430 under reference labels to +0.271 and +0.227 under adjudication. The reported narrative-bootstrap intervals excluded zero.
A small audit offered a repair
The authors then tested whether a limited human audit could help recover the ranking. In frozen-audit retrospective replay, the correct direction was recovered with probabilities of 0.996 for the local unit, 1.000 for OpenToM Vanilla and 1.000 for OpenToM Conservative, with no false terminations.
At 50 usable truth labels, TriSource-Restore, the proposed repair, reduced human-only RMSE, the reported measure of estimation error, in all three units, narrowed intervals by up to 37%, and achieved interval coverage of 0.959, 0.996 and 0.961.
A crossover criterion was also checked against 21 released-system pairs. It correctly classified 13 reversals and 8 non-reversals within that tested set.
A finding with a defined scope
The work is an arXiv version 1 preprint dated 26 August 2026. Its central conclusion is narrower than a verdict on model quality: in the tested open-ended tracking evaluations, changing how unmatched beliefs were labeled changed which confidence rule appeared better.
The 73% local figure and the 90% to 96% OpenToM figure describe audited samples, not a universal rule that every unmatched belief is true. The evidence shows why finite-reference evaluation can mis-rank confidence rules in the tested settings; it does not settle how the pattern will behave in every dataset or deployment.
Paper data and sources
Original title: Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking
Authors: Zhexi Feng, Wuxi Chen, Bingrui Zhang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text