Peer-reviewed

Exact methods more reliable for sparse diagnostic-test meta-analyses, study finds

The clearest differences appeared in small or sparse studies, while methods were broadly similar in large, non-sparse datasets.

When diagnostic-test studies are small or contain zero cells, the method used to combine their results can change pooled estimates of sensitivity and specificity. A methods analysis found that exact within-study variance calculations and BGLMM generally performed better than the standard approximate BLMM-Logit approach in sparse datasets, particularly for estimating overall accuracy and the uncertainty around it.

The study examines how evidence from diagnostic-test studies is combined; it is not a clinical comparison of diagnostic tests or an assessment of patient outcomes.

How the methods were tested

The researchers compared five approaches: BLMM-Logit, BGLMM, Exact-Logit, Exact-ASR and Exact-FTDA. The proposed exact methods calculate the uncertainty within each study analytically and can be paired with different transformations of sensitivity and specificity.

The simulation explored 324 scenarios and used 1,000 runs with different seeds. The researchers measured bias, or systematic error; RMSE, which captures overall estimation error; how often confidence intervals contained the target value; and the width of those intervals.

Sparse data changed the results

In sparse scenarios, Exact-Logit had the smallest bias for overall sensitivity, while Exact-Logit and BGLMM had the smallest bias for specificity. BLMM-Logit had the largest negative bias. Exact methods and BGLMM also estimated between-study variance components more accurately, whereas BLMM-Logit underestimated them across the scenarios.

Exact-Logit and BGLMM had the smallest RMSE for sensitivity and specificity in sparse meta-analyses, while BLMM-Logit had the largest. For between-study variance components, Exact-ASR and Exact-FTDA had the least RMSE. Under sparsity, those two methods also had the least bias and smallest RMSE for the between-study covariance.

Confidence-interval coverage was more reliable with the exact methods and BGLMM in sparse settings, especially for sensitivity, while BLMM-Logit coverage was poor. The methods differed in interval width: BLMM-Logit, Exact-Logit and BGLMM were narrower in sparse datasets, while Exact-ASR and Exact-FTDA were widest. In small, non-sparse datasets, the arcsine methods were narrowest.

When per-study samples were large and the primary studies had no zero cell counts, all methods generally performed comparably, with no clear separation between them.

The real examples showed the same divide

In a dataset on ultrasonography for appendicitis in children, 23 studies averaged 77 children with appendicitis and 254 without appendicitis per study. BLMM-Logit produced pooled sensitivity estimates at least 2% lower and pooled specificity estimates at least 1% lower than the other methods.

The summary receiver operating characteristic analysis also differed substantially in this dataset. BGLMM and BLMM-Logit had higher areas under the curve and more precise confidence and prediction regions than Exact-ASR and Exact-FTDA, while Exact-Logit had the lowest area under the curve.

The contrast was much smaller in a second dataset involving MMSE testing. It included eight studies, averaging 47 participants with the condition and 95 without. Because it did not have sparse false-negative or false-positive counts, the methods produced similar average sensitivity and specificity estimates and almost identical summary curves and areas under the curve.

A targeted recommendation, not a universal winner

The authors recommend the exact methods or BGLMM for diagnostic-test meta-analyses with sparse data or small within-study samples. They consider the methods broadly similar when data are non-sparse and per-study samples are large.

The analysis assumes that each primary study uses a single test-positivity cutoff and does not account for differences in cutoffs between studies. Exact variance calculations also retain large-sample assumptions for statistical inference.

The findings therefore do not establish one method as best in every data configuration: interval widths varied by dataset, and the methods were often comparable when studies were large and non-sparse.

Paper data and sources

Original title: Diagnostic test accuracy meta-analysis based on an exact within-study variance calculation method
Authors: Dabi O, Negeri Z
Journal/Repository: Research Synthesis Methods
Status: Peer-reviewed
First online: 2026-08-19
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.