Preprint

Preprint model limits the pull of extreme data points

A proposed Bayesian method resisted extreme observations in simulations and had the lowest prediction error in 19 of 20 comparisons, but its intervals narrowed too far in one calibration test.

A statistical preprint describes a Bayesian quantile-regression model designed to stop sufficiently extreme observations from exerting the same pull as regular data. In selected simulations, the method generally produced smaller estimation errors and shorter credible intervals than GAL and skew-t, two cited alternatives. In the supplied prediction comparisons, it had the smallest error in 19 of 20 cases. But a separate calibration exercise found a warning: under one misspecified process, the intervals became too narrow to maintain their stated coverage.

The paper works with quantile regression, which estimates a chosen conditional quantile of an outcome. AL-LPAL combines a conventional asymmetric-Laplace component with a log-Pareto scale mixture of asymmetric-Laplace distributions. Both components share the target quantile. The theory says the resulting distribution has sharp central concentration, log-regularly varying tails in both directions and an unbounded density at the target quantile.

A model built for the tails

The authors' central aim is to distinguish observations that fit the regular pattern from observations that may be extreme. Their theoretical results say that AL-LPAL preserves the prescribed quantile and gives both tails log-regular variation. Its density is unbounded at the target quantile. In plain language, the model puts a sharp peak at the target while retaining unusually heavy tails on either side.

An accompanying posterior-robustness theorem says that, under the stated prior and design assumptions, the posterior, or updated distribution, for regression coefficients and scale converges to the posterior based only on non-outlying observations as contamination diverges in either direction. A separate proposition gives sufficient prior conditions for finite posterior moments of the coefficients and scale, even though ordinary moments of the super-heavy-tailed error may not exist.

To fit the model, the authors use latent-variable Gibbs sampling with an augmented log-Pareto representation. They also develop mean-field variational Bayes as an approximation to the full posterior.

Strong results in selected contamination tests

The main simulations used datasets of 100 observations with either three or 20 covariates. The models were fitted at the 0.1, 0.5 and 0.9 quantiles, and each setting was repeated 500 times. The reported comparisons focused on RMSE, an error measure that penalizes large misses, and the average length of credible intervals.

The clearest numerical gap appeared in a low-dimensional, one-sided contamination setting at the 0.9 quantile. AL-LPAL had an RMSE of 0.113 and an average interval length of 0.265. GAL recorded an RMSE of 6.908 with intervals averaging 5.683, while skew-t recorded an RMSE of 9.201 and average intervals of 7.255. The result is specific to that contamination setting.

The broader two-sided tests pointed in the same direction without showing uniform dominance in every scenario. Across Cases 4 to 6, the proposed methods generally had much smaller RMSE at the 0.1 and 0.9 quantiles than GAL and skew-t, while retaining substantially shorter credible intervals.

In a fixed-data contamination path, the MCMC estimate became very close to the clean-data estimate as contamination increased, while GAH remained increasingly affected. Variational Bayes followed broadly but deviated at the most extreme values. An influence-function analysis, which tracks how much a small change in an observation can affect an estimate, found that AL-LPAL influence was progressively attenuated as contaminated responses moved farther into the tails, with similar qualitative attenuation across mixture weights.

Prediction gains came with a calibration warning

The method also did well in the supplied real-data prediction tests. Across the reported datasets, quantile levels and loss functions, AL-LPAL achieved the smallest prediction error in 19 of 20 comparisons. The exception was the Boston check loss at the 0.75 quantile, where GAL was lower. The evidence is limited to those supplied datasets and their specified preprocessing.

Calibration asks whether an interval captures the target as often as its stated level suggests. In misspecified Case 3, average absolute bias fell from 0.080 at n = 100 to 0.042 at n = 500. Yet empirical coverage fell from 0.869 to 0.739, and dispersion ratios at n = 500 were about 0.60 to 0.62, indicating posterior under-dispersion. The estimates moved closer to the target on average, but the intervals captured it less often in that experiment.

The speed advantage was large. Across representative settings, variational Bayes was approximately 47 to 191 times faster than Gibbs sampling, especially in the moderately high-dimensional setting. That timing comparison does not establish equivalent posterior accuracy, particularly under severe contamination, and variational Bayes does not inherit the exact posterior-robustness theorem.

What remains to be tested

The theorem is conditional on its stated prior and design assumptions. The simulation evidence comes from the specified sample sizes, dimensions, quantiles and replication counts, so the results should be read as setting-specific rather than universal.

Mean-field variational Bayes can depart from MCMC at the most extreme contamination values, as the fixed-data path showed. Its speed advantage therefore does not establish equivalent posterior accuracy. The calibration experiment likewise showed that shorter intervals can come with under-dispersion and below-target coverage under misspecification.

The supplied metadata identifies the work as arXiv version 1, a preprint dated 25 August 2026. No funding source is reported. The paper is aimed at statisticians developing robust quantile-regression tools for data with extreme residuals.

Paper data and sources

Original title: Log-regularly varying scale mixture of asymmetric Laplaces for robust Bayesian quantile regression
Authors: Dongu Han, Genya Kobayashi, Shonosuke Sugasawa
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.