Exact calibration does not guarantee a unique answer when a statistical test asks whether a parameter differs in either direction. A mathematical analysis compares four exact two-sided constructions, known as equal-tail, density-ordered, UMPU and likelihood-ratio p-values, and shows that they can diverge when the underlying distribution is asymmetric. In other words, two methods can both meet the formal calibration requirement yet assign different levels of evidence to the same observation.
The work concerns continuous one-parameter natural exponential families, a class of probability models in which a single parameter governs the distribution. It is a theory paper, not a study of people or medical outcomes: the examples use inverse-Gaussian and hyperbolic-secant models as illustrations rather than empirical evidence.
Symmetry is the dividing line
At one fixed null value, symmetry about the null mean is the key condition behind agreement. The analysis finds that UMPU and equal-tail p-values coincide exactly when the null distribution is symmetric about its mean. With the paper's additional regularity condition on the two density branches, UMPU and density-ordered p-values have the same symmetry requirement.
Under that same density condition, equal-tail and density-ordered p-values also coincide exactly when the fixed-null distribution is symmetric about its mean. When the coincidence is required across the whole natural exponential family rather than at one null value, the restriction becomes much stronger: the family must be Gaussian, apart from a change of location and scale.
The comparison with likelihood-ratio ordering produces a different picture. The paper reports that global agreement between UMPU and likelihood-ratio p-values occurs exactly for normal, gamma and inverse-Gaussian families, up to affine transformation. For equal-tail and likelihood-ratio agreement, the authors derive necessary mathematical conditions but do not claim a complete classification for every one-observation family.
More data does not erase the issue automatically
For independent, identically distributed observations, the established fixed-sample results carry over through the model's canonical sufficient statistic, a compact summary that retains the information needed for the parameter calculation. But if equal-tail and likelihood-ratio coincidence persists along an unbounded sequence of sample sizes, the analysis forces the family to be Gaussian. Within the specified class of models whose variance follows a power law, the same Gaussian conclusion holds for global equal-tail and likelihood-ratio coincidence.
The density-ordered comparison is similarly restrictive only with an extra assumption. Under a differentiated local Edgeworth condition, which controls a refined approximation to the distribution and its density as sample size grows, density-ordered and likelihood-ratio coincidence along an unbounded sequence also characterizes Gaussianity. The authors stress that this condition is additional to the basic assumptions on the exponential family.
The numerical gap can be large
The inverse-Gaussian example makes the abstract distinction concrete. At an observed value of 0.5, the equal-tail p-value is about 0.9803, while the likelihood-ratio and UMPU values are about 0.6171. At the null mean, represented by an observed value of 1 in the example, the likelihood-ratio and UMPU p-values are 1, while the equal-tail value is about 0.5724.
A separate hyperbolic-secant calculation found sizeable differences on an evaluated grid of null quantiles. The largest gaps were 0.291 between equal-tail and density ordering, 0.186 between equal-tail and UMPU, 0.186 between equal-tail and likelihood-ratio, 0.442 between density ordering and UMPU, 0.431 between density ordering and likelihood-ratio, and 0.034 between UMPU and likelihood-ratio. These are descriptive grid values, not analytic maximum differences.
Those numerical differences can affect decisions, not just reported decimals. In the hyperbolic-secant example, the reported exact null rejection probability was 0.05, yet density ordering and UMPU disagreed on about 4.5% of null observations. At the reported parameter value π/3, equal-tail, density ordering and likelihood-ratio procedures each had strictly higher model-based power than UMPU.
What the results do and do not settle
The analysis does not identify one ordering as universally preferable. Its central point is narrower and more consequential: exactness alone leaves room for multiple two-sided p-values under asymmetry. The equal-tail and likelihood-ratio variance equation and density identity are necessary restrictions, not a full unrestricted one-observation classification, and the complete equal-tail conclusion is established only in the specified power-variance subclass or under sample-size stability.
The numerical examples are illustrations of the theorems and calculations, not empirical evidence. No participant sample is reported, and no funding source is reported in the supplied text. The density-ordered sample-size conclusion is conditional on the additional differentiated local Edgeworth assumption rather than the basic natural exponential-family assumptions alone.
Paper data and sources
Original title: Exact two-sided p-values in natural exponential families: coincidence, non-uniqueness, and sample-size stability
Authors: Shaul K. Bar-Lev, Linard Hoessly
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text