A computational benchmark of language-model auditors found that one listed trained checkpoint scored above an untrained Haiku 4.5 baseline on the reported composite, audit-quality and false-positive calibration measures. The benchmark also produced a sharp contrast between reward strategies: a concerningness-focused run had a higher production misalignment discovery score alongside a steep fall in false-positive calibration.
The study trained auditor policies on a bank of targets containing 26 planted behaviors and about 50 scenario seeds per behavior, or roughly 1,300 behavior-seed pairs. The evaluation behaviors, seeds and target model were held out. One epoch was one pass through the training data, and Claude Sonnet 4 judged rewards during training.
The leading checkpoint combined two kinds of checks
One listed setup was a three-epoch reference-pairwise model with four of eight false-positive calibration cases. It scored 72.7 plus or minus 2.6 on audit quality, 99.6 plus or minus 0.7 on false-positive calibration and 48.7 plus or minus 1.1 on the composite. The untrained Haiku 4.5 baseline scored 65.6 plus or minus 2.9, 98.9 plus or minus 1.3 and 44.2 plus or minus 1.0 on those measures, respectively.
The calibration score was defined as 100 minus the false-positive rate. In this check, a false positive meant accusing a target with no injected behavior of a problem. The calibration evaluation comprised 270 audits across six behaviors, 15 seeds and three rollouts, using 20-turn investigations against Sonnet 4.6 without behavioral injection.
Audit quality was measured in a separate set of 270 audits covering six evaluation behaviors, 15 tailored seeds and three rollouts. Each investigation lasted 30 turns against DeepSeek v3.1, and Opus 4.6 scored it on six dimensions.
A higher production score came with a steep calibration drop
The concerningness-RL arm trained on eight production targets without planted behaviors, using 181 default Petri seeds and the same 38-metric Petri judge as the production-value evaluation. Its production-value cell was in-distribution, meaning it resembled the training setting, while the audit-quality and false-positive cells were out-of-distribution, or outside that setting.
In that run, the production misalignment discovery score rose from 7.8 to 23.0, while false-positive calibration fell from 98.9 to 3.3. The paper characterizes the apparent discovery gain as largely illusory.
The strategy analysis described pairwise regimes as more stable than binary and concerningness training. The reference-pairwise auditor converged on controlled, systematic experiments, while binary and concerningness-RL auditors displayed aggressive, manipulative or other failure-mode patterns.
The tested calibration levels tracked a steep score change
In a two-epoch pairwise sweep, false-positive calibration was 3% when none of eight calibration cases were included, 28% with one case, 93% with two and near-perfect 97% with four. The four-of-eight condition had the highest composite.
The compute-matched comparison of single-target and multi-target training was similarly close. Across four pairs at two epochs, averaged results were nearly indistinguishable, composite differences stayed under one point and no evaluation dimension consistently favored either approach.
A narrow transfer test also favored the trained checkpoint
The researchers tested transfer on the hardest TD+KTO targets from six of AuditBench’s 14 behaviors. The investigations used Claude Code scaffolding, a 25,000-token generation budget and 16 independent audits per model for each behavior.
On the STC transfer test, the three-epoch calibrated multi-target pairwise checkpoint detected the target behavior in 28.1% of evaluations, with a reported plus or minus 5.4 percentage points. Base Haiku detected it in 11.5%, plus or minus 6.2 points. Each rate averaged six behaviors and 16 audits per behavior, or 96 evaluations.
A faster crux evaluation used all eight crux features in a multivariate regression. It predicted 73% of the variation in realism scores, 67% to 68% of production-metric variation and 53% of audit-quality-identification variation. Because the analysis was in-sample, it was a regression proxy rather than independent validation.
What the benchmark covered
The production discovery evaluation used 543 audits built from 181 default Petri seeds and three rollouts. It used 30-turn investigations against Sonnet 4.5 without planted behaviors. The realism evaluation used 100 comparisons built from 20 topic-matched seeds and five rollouts, with 10-turn investigations against Sonnet 4.5 paired with real WildChat conversations.
The benchmark combined held-out planted-behavior evaluations with clean production-model checks and a selected transfer test. The training evaluation held out behaviors, seeds and the target model, but the AuditBench result covered six of 14 behaviors. The tested single-target and multi-target comparisons showed no consistent advantage for either approach.
The work is an arXiv preprint by Paul Rosu and Rowan Wang. The supplied record lists Rosu as an Anthropic Fellow and Wang as Anthropic, and reports no funding source.
Paper data and sources
Original title: Training Alignment Auditors via Reinforcement Learning
Authors: Paul Rosu, Rowan Wang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text