A mathematical preprint suggests that the pattern of a change may matter as much as its size when comparing two estimated correlation matrices. In the paper's modeled examples, a proposed intrinsic off-log distance becomes more responsive to small shifts spread across many correlation coordinates as the number of variables grows. A rank-one, factor-driven shift follows a different pattern: the distance stays near the same size while the sample needed to detect it rises with the dimension.
The work is theoretical and asymptotic. It studies two independent, identically distributed samples of p-dimensional observations, with nonsingular covariance matrices and finite fourth moments. Under the null hypothesis, both populations have the same correlation matrix C, although their fourth-moment structures may differ. The analysis covers equal and unequal sample sizes, as well as Gaussian, elliptical and non-elliptical sampling and local alternatives.
The distance is not zero under the null
The paper's central warning is that sampling noise creates a strictly positive distance even when the underlying population correlation matrix has not changed. With a common covariance structure, the expected distance is bounded by 2 times the square root of tr(V_C), where V_C is the covariance matrix for the paper's transformed correlation coordinates. Near Gaussian independence, the stated rule of thumb is approximately 2 times the square root of d divided by n. For unequal group sizes, n is replaced by the effective sample size.
That effective sample size is n_eff = 2n1n2/(n1+n2). With unequal sample sizes, the limiting weighted chi-squared law uses 4 times the eigenvalues of an eta-weighted combination of the two coordinate-covariance matrices. If those covariance matrices are equal, the limit no longer depends on the limiting ratio of the sample sizes.
A reference distribution for estimated differences
For equal sample sizes, the scaled squared distance converges under the null to a weighted sum of independent central chi-squared variables. The weights are 2 times the eigenvalues of the sum of the two covariance matrices for the GFT coordinates, the transformed correlation coordinates used in the calculation. In plain terms, it is a reference distribution built from the sampling covariance of those coordinates.
At Gaussian independence, the coordinate covariance is the identity matrix and the scaled squared distance has a parameter-free limit equal to 4 times a chi-squared variable with d degrees of freedom, where d is the number of distance coordinates. Near that point, the covariance changes only at second order in the departure from independence, and the associated weights are 4 plus a second-order remainder. The fixed 4-chi-squared reference is therefore second-order accurate in that neighborhood.
From the limit to a threshold
The preprint also develops ways to turn the limiting distribution into an upper-tail threshold. It gives a closed-form cumulant-generating function, a function used in the tail calculation, together with its first two derivatives. It then presents a Lugannani-Rice saddlepoint approximation for upper-tail probabilities and a conservative Chernoff envelope with an explicit choice that avoids root-finding. The paper does not report empirical validation of the accuracy of these tools.
The proposed thresholds can use estimated quantities under consistency conditions. In the Gaussian common-covariance setting, a consistently estimated correlation matrix gives consistent covariance weights and eigenvalues. With non-Gaussian data, calibration additionally requires consistent estimates of fourth-moment quantities or of the coordinate covariance itself.
The intrinsic-distance calculation is not automatically the same as a coordinate-based Wald calculation. The two quadratic forms are proportional for every coordinate difference only when V_C is a scalar multiple of the identity. That condition holds at Gaussian independence and for every Gaussian C in the bivariate case, p=2. The preprint leaves their relative local efficiency away from independence unresolved.
Widespread small shifts look different from factor changes
In one coordinate-dense illustration, the significance level is 0.05, the sample size is 1,000, and the target power is 50 percent. Under those conditions, the smallest detectable per-pair shift in the study's GFT coordinates falls from 0.043 when p=5 to 0.019 at p=20 and 0.008 at p=100. The product of that minimum shift and the fourth root of d is approximately 0.069. These are scenario-specific calculations, not empirical power estimates.
The modeled rank-one case points in the other direction. With a systemic correlation shift of 0.05, the GFT displacement norm stays near 0.14 from p=5 through p=100, while the sample size needed for 50 percent power rises from about 1,100 at p=5 to about 5,200 at p=30 and 16,500 at p=100. The required sample size per dimension stabilizes near 235. The calculation uses an equicorrelation scenario and prevailing correlation levels, so it does not establish uniformly strong performance for low-rank changes.
What the theory leaves open
The results describe an asymptotic reference point, not demonstrated finite-sample performance. The theory assumes nonsingular covariance matrices and finite fourth moments, and the preprint does not evaluate finite-sample convergence. The non-Gaussian extension gives kurtosis-dependent chi-squared limits: for equal sample sizes, the coefficient is 2(2 + kappa1 + kappa2), falling to 4(1 + kappa) when the two kurtosis parameters are equal. Under non-elliptical sampling, the limit retains a general weighted-chi-squared form, while the general covariance structure and first-order flatness remain unresolved.
Paper data and sources
Original title: The Sampling Distribution of the Log-Euclidean Distance Between Sample Correlation Matrices
Authors: Argyn Kuketayev
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text