Preprint

New statistical framework adds an inconclusive option to research

A statistical preprint proposes a way to classify estimates into rival possibilities while controlling the risk of assigning the wrong class.

A new statistical framework proposes an additional outcome for research findings that do not cleanly support one interpretation: leave the result inconclusive. The approach is designed for situations in which researchers must choose among rival possibilities, rather than simply ask whether an estimate differs from zero.

The method divides the possible values of an estimate into at least two substantively relevant classes that together cover all possible values. It then assigns the estimate to a class only under an error-controlled rule, or withholds a classification when the evidence is not strong enough. That structure is meant to make room for conclusions such as directional, negligible or substantial, instead of reducing every result to significant or non-significant.

The work is a dated preprint, with front matter showing August 23, 2026. It presents a statistical framework, computer simulations and a reanalysis of 32 pre-registered indices from an existing field experiment.

A test built around wrong classifications

The central safeguard is a limit on misclassification: assigning an estimate to a class that does not contain its true value. The framework aims to keep the probability of any such wrong assignment at or below a chosen level, called alpha. Among rules that meet that limit, it selects the one with the strongest worst-case performance for making the correct classification.

The formal setup assumes that the estimate follows an asymptotically normal sampling model and that its variance can be estimated consistently. In plain terms, the guarantees depend on the estimate behaving as the model assumes when samples are sufficiently large and on its uncertainty being estimated reliably. The paper examines three versions of the test: sign, magnitude, and sign-and-magnitude classification.

A familiar confidence-interval shortcut works exactly in the simple case of two classes separated by one boundary. But when there are more boundaries and the estimate is noisy, that shortcut can miss the intended error rate. The full acceptance-region calculation is therefore needed for more complicated classification problems.

Simulations found the safeguards held in tested cases

In one simulation, the researchers generated 30 estimates at a time and repeated the exercise 20,000 times. They set the target alpha at 0.05 and tested several dependence structures, including an adversarial mixture designed to challenge error control. Across those settings, the sign and sign-and-magnitude procedures kept both the false-discovery rate and the family-wise error rate at or below the target.

A separate comparison tested five confidence-interval methods using 500 observations per replication over 20,000 replications. Four methods kept misclassification at or below 0.05. The 90% confidence interval did not: its peak misclassification rate was near 0.083 in the simulation. These are computational comparisons, not independent tests in new data.

The paper also compares its magnitude classification with the equivalence-testing method TOST. At alpha = 0.05, the TOST acceptance region exists only when the standard error is less than about 0.61 times the chosen threshold. The proposed magnitude-classification region exists for any positive standard error, giving it a broader formal range under the stated normal model.

An application produced more than yes or no

To illustrate the approach, the paper reanalyzes a field experiment involving regular Fox News viewers exposed to CNN. It applies a threshold of 0.15 standard deviations for calling an effect negligible and examines 32 pre-registered indices.

The ordinary two-tailed test split the indices evenly: 16 were significant and 16 were not. The sign test classified 18 estimates by direction, with 17 positive and one negative. The sign-and-magnitude test identified two estimates as positive and substantial, 12 as negligible and the remaining 18 as inconclusive.

Taken together, the two classification tests produced conclusive error-controlled classifications for 29 of the 32 indices, or about 0.91. The two-tailed test produced a conclusive result for 16 of 32, or 0.50. The comparison shows what the framework adds: it can separate a negligible result from an inconclusive one while also allowing directional or substantial classifications when the evidence supports them.

That comparison needs a clear qualification. The application did not apply a multiple-testing correction across the displayed columns; each classification controlled error at alpha = 0.05 for its own estimate. The 29-of-32 summary therefore is not a family-wise corrected claim across all the indices.

The application is also a reanalysis, and a few borderline outcomes differ from the original report because R's glmnet and Stata's elasticregress selected different covariates. The result is best read as an illustration of how the framework changes the language of findings, not as a universal demonstration that the method is superior.

A useful promise, with choices still to make

These guarantees are conditional on the stated sampling model and on reliable variance estimation. The formal setup is asymptotic, so the plug-in approach is not an exact finite-sample guarantee when the variance is unknown. The paper also shows why the simple confidence-interval shortcut can fail in noisy, multi-boundary problems, making full acceptance-region calibration necessary.

The authors present classification testing as a way to adjudicate rival possibilities and produce richer error-controlled conclusions. The supporting evidence is analytical, simulated and based on one application, so it does not establish that the method will be superior in every empirical setting.

Paper data and sources

Original title: Classification testing: A new framework for drawing qualitative conclusions from quantitative estimates
Authors: Andrew C. Eggers, Zikai Li
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.