An arXiv preprint reports a higher level of simulated statistical power for RDQ M2-alpha1 than for RBP(0.9) on POI data. At a simulated sample size of 100 queries, the RDQ configuration had median power@100 of 0.353, compared with 0.287 for RBP(0.9). Here, power means the proportion of 300 random query subsamples for each system pair and simulated query count in which the comparison produced a rejection. The pairwise comparisons used two-sided randomization tests with 1,000 permutations and a 0.05 threshold. In ordinary terms, the result says the RDQ setting more often detected a difference in this simulation. It does not say users would prefer the rankings, because the study measured internal statistical sensitivity and consistency rather than user-perceived quality.
A metric built around order
RDQ evaluates a candidate ranking against an ordered reference list and gives no credit to items outside that list. In the M2 variants, alpha controls the strength of a penalty for moving an item away from its reference position. That makes the ordering in the reference list part of what the score measures. The controlled examples examined this scoring behavior, not user judgments.
The evaluation used controlled examples and two tests
The evaluation combined controlled examples with two empirical tests. The POI experiment applied the metric to outputs from 12 systems using a stratified sample of 5,000 queries. The POI sample included 100 queries with three or fewer reference items and 4,900 with more than three. A separate TREC Deep Learning evaluation used four systems and pooled queries from 2019 through 2022. Those four component datasets contributed 43, 45, 57 and 76 judged queries, for 221 in total.
In the controlled examples, only RDQ M2 variants responded to every tested property, with alpha setting the strength of the displacement penalty. That result describes what the scoring rules did under designed conditions. It is not evidence about user preference.
The POI result depended on the setting
On the POI data, RDQ M2-alpha1 had median power@100 of 0.353, versus 0.287 for RBP(0.9). In this simulation, the higher value means that the RDQ configuration produced a rejection more often across the sampled system comparisons. That is a measure of internal statistical sensitivity, not a finding that the rankings were more useful to people.
The result depended on the parameter choice. Across 25 tested M2 combinations, median power@100 ranged from 0.230 to 0.353, and 16 configurations exceeded RBP(0.9)'s 0.287. The reported 0.353 was specific to M2-alpha1 in this comparison, rather than a result shared by every M2 setting.
A separate measure tracked ranking stability
The study also measured ranking stability. This asks how closely a ranking built from a smaller query sample agrees with the ranking produced from all available queries under the same metric configuration. The analysis summarized that agreement with Kendall's tau and with the normalized area under the tau-versus-sample-size curve.
RDQ M2-alpha1 had a normalized area under the tau curve of 0.904. It reached mean tau of at least 0.8 with 200 queries, while RBP(0.9) reached the same threshold at 250. Because each metric was compared with its own full-query ranking, this result shows consistency as the sample grows, not which ranking was best.
The TREC-DL comparison was more mixed
On pooled TREC-DL data, RDQ M1 and native NDCG had similar median power at 25 queries, 0.108 and 0.102 respectively. At 100 queries, native NDCG was higher, at 0.437 versus 0.367 for RDQ M1. The comparison involved four systems and six system pairs, using the pooled 221 judged queries. The component counts were 43, 45, 57 and 76. This separate result leaves the POI finding tied to the specified RDQ M2-alpha1 and RBP(0.9) comparison.
What the results do not answer
The study's main measures answer a narrower question than a user study. Discriminative power and ranking stability measure internal statistical sensitivity and consistency, not whether metric differences correspond to user-perceived quality. Human-preference validation is left for future work.
The POI reference lists add another qualification. They came from a GenAI-assisted labeling pipeline and were treated as a silver-standard proxy rather than human-annotated ground truth. The POI test therefore does not provide human-annotated ground truth.
The work is identified as an arXiv preprint, version 1, dated 26 August 2026. No funding source is reported. The paper states that POISS was not publicly available at submission and was scheduled for release.
Paper data and sources
Original title: Rank-Deviation Quality: A Distance-Aware Metric for Multi-Answer Retrieval and Ranking Evaluation
Authors: Xiaokun Zhou, Alessandro Moschitti, Danielle Class
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text