Compressing a time series by averaging several observations can give a long-run cross-sectional estimate a faster theoretical convergence rate than taking a single point, a new econometrics preprint reports. The reason is a two-part statistical effect: K-averaging magnifies the persistent signal as the window grows while diluting noise; point sampling gets the signal magnification but cannot dilute noise.
That advantage is conditional, not a blanket endorsement of averaging. The paper’s results rely on assumptions about independence, exogeneity, stationarity and moments. With K fixed, its cross-section estimator is consistent and asymptotically normal whether the regressor is I(1), a nonstationary process whose level can wander, or I(0), a stationary one.
The statistical trade-off
For a simple I(1) process, cross-sectional variance grows in proportion to elapsed time when the average long-run variance of its innovations is nonzero. In plain language, the spread between units can widen as persistent shocks accumulate.
Averaging changes that pattern. Stationary data fan in as the sample grows, while full-sample averages of I(1) data still fan out but have one-third the variance of point sampling at the same time horizon.
Long differencing, which compares observations separated by a long interval, sits outside the main theorem class for nonnegative compression weights. Its effective fanning rate is K minus M divided by three, so it can be slower than K when the local-average width M is large.
The paper also examines endogeneity, where the regressor and regression error are related. Under its I(1) setup, average endogeneity covariance remains of order one while cross-sectional variance is of order K. As K increases, endogeneity bias therefore vanishes, although correlation can still inflate variance. If stationary dynamic terms are omitted, their contribution becomes negligible as K grows, but the resulting coefficient is a long-run multiplier rather than a response to a single innovation.
Simulations offered the clearest support
In the reported Gaussian simulation, the data contained 50 periods and 50 units. Point-sampling, K-averaging and long-differencing estimates were close to standard normal after normalization. Their standard errors were 0.014, 0.001 and 0.006, respectively, and K-averaging had higher statistical power, meaning it was better at detecting a signal in the simulated settings.
Applications were less uniform
In a consumption and GDP application covering 73 periods and 44 countries, the K-averaging estimate was 0.935, close to 0.950 from the comparison estimator cupBC. Point sampling produced 0.862 and long differencing 0.787. The apparent advantage came with a cost: the K-averaging standard error was 9.66 times the cupBC standard error. The paper attributes the slower convergence to possible omitted fixed effects and reports sensitivity to the compression choice.
A demographic regression using 35 countries from 1960 to 2022 produced positive inflation coefficients for older and younger demographic shares: 0.289 and 0.406. Neither met the paper’s 10% significance threshold. In the growth regression, both coefficients were negative, at −0.154 for the older share and −0.048 for the younger share. The older-share estimate had a t statistic of −2.338, while the younger-share result was only marginally significant. These figures describe regression associations, not proof that demographic composition caused either outcome.
The temperature-growth application, based on 48 U.S. states, was more fragile. Temperature showed a slight upward trend but little cross-sectional fanning. Linear estimates declined as K increased, while adding latitude made them more stable.
A nonlinear specification produced a negative linear temperature term and a positive quadratic term, but the full-sample t statistics were −1.301 and 1.236, so neither term was statistically distinguishable from zero. Long-difference estimates varied from positive to negative, and the long-run trend coefficient was not well determined. The paper’s interpretation is that short-term temperature variation overwhelmed the weak trend.
The gain has clear boundaries
The faster rate also breaks down when coefficients differ across units. Even under the paper’s stated orthogonality condition, heterogeneous coefficients leave the estimator only √N-consistent because the magnified signal is offset by magnified regression noise.
The practical message is to test whether estimates remain stable as the compression window changes. The country and state applications moved in different ways across estimators and values of K, showing that averaging is a conditional tool for studying long-run associations, not a universal replacement for other methods or a causal finding.
Paper data and sources
Original title: Cross-Section Estimation of Long-Run Relations Using Time-Compressed Data
Authors: Serena Ng, Nikolay Gospodinov
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text