Preprint

Preprint reports near-zero bias for log-quantile fits in simulations

An arXiv v1 preprint reports exact one-step log-scale linearization, low variability in selected simulations and model-specific results in two observational examples.

An arXiv preprint reports that log-quantile least squares, or log-QLS, had the least variability and nearly no bias in its numerical comparison, while a log-scale formulation reached its weighted-least-squares solution in one step for the log-location-scale model family. The findings are theoretical and simulation-based, and the accuracy comparison was distribution-specific.

The study develops four estimators: ordinary quantile least squares, generalized quantile least squares, and their log versions, log-oQLS and log-gQLS. It presents them as frameworks for estimating and validating log-location-scale distributions. The study combines theoretical derivations, simulations involving lognormal and Weibull contamination, and two observational data examples. Published as an arXiv v1 preprint dated 25 August 2026, it says it is to appear in Variance.

The calculation at the center

For log-location-scale models, the authors transform selected quantiles logarithmically. That makes the weighted least-squares linearization exact, so the iterative update converges in one step. The result is stated for that model family.

In the numerical comparison, Nelder-Mead was the weakest tested approach on bias and root mean squared error, a summary of estimation error, while log-QLS had the least variability and nearly no bias. The paper also reports that log-quantile methods were generally more efficient than the corresponding raw-quantile methods, and that log-gQLS often reached 90% or higher relative efficiency against maximum likelihood in the models studied. The authors interpret these results as combining robustness with high efficiency.

What contamination tests found

Robustness in the paper is tied to which quantiles are retained. QLS and log-QLS are assigned a positive asymptotic breakdown point, with BP = min{a, 1 - b}, where a and b mark the lower and upper retained quantile limits. The paper treats this as an asymptotic property determined by the margins of the retained range.

The main contamination simulations mixed clean and contaminating lognormal or Weibull data at 0%, 3% and 8%. In a repeated study, the researchers generated 104 samples of 500 observations, compared maximum likelihood with log-oQLS and log-gQLS, used 25 quantiles and tested three quantile ranges.

At 8% contamination, Banerjee-Iglewicz rules paired with log-oQLS and log-gQLS detected upper outliers in 0.63 and 0.65 of Weibull samples, compared with 0.40 under maximum likelihood. For lognormal samples, the corresponding rates were 0.38 and 0.37, compared with 0.23 under maximum likelihood. Non-extreme declarations were 0 in the repeated study, while oracle upper-outlier rates were 0.77 for Weibull and 0.58 for lognormal. Those rates were conditional on samples containing outliers in both tails.

Choosing the quantiles

The paper links quantile choice to a trade-off between robustness and efficiency. It presents evenly spaced probabilities across a broad retained range as a simple design: the lower retained probability is no higher than 0.05, the upper probability is at least 0.95, and at least 15 quantiles are used. It reports high asymptotic relative efficiencies for this design and identifies it as optimal for log-Cauchy, while noting that efficiency depends on both the distribution and the quantile design.

For model checking, the residual-based goodness-of-fit test estimates its p-value with a specified bootstrap procedure, meaning the calculation is repeated on model-based resamples. The theoretical distribution of the test statistic has not been derived, which the authors identify as an ongoing challenge. At a 5% significance level, the simulations reported correct calibration; under contamination, rejection rates increased toward 1 as sample size approached 1,000, especially for the 8% mixture.

Two data examples

One illustration used 1,006 Google daily returns from January 2, 2020, through December 29, 2023. Log-Logistic was reported as the best fit. Log-Laplace was borderline, with p-values from 0.06 to 0.11, while log-Cauchy and lognormal were strongly rejected. For the poorly fitting lognormal model, the authors attributed the lack of fit to one upper outlier identified by robust log-oQLS and log-gQLS but missed by maximum likelihood.

The second illustration analyzed 32 normalized U.S. hurricane losses. The selected log-gQLS models used 15 quantiles for estimation and 15 for validation. All were reported as acceptable, with p-values above 0.10 and no model-specific outliers. A separate tail check used fits to losses from 1900 to 1999 to estimate exceedance probabilities in 2000 to 2022 at thresholds of 50 billion, 100 billion, 150 billion and 200 billion. At the higher thresholds, log-Gumbel and log-Laplace were judged closest to the later observed probabilities. Lognormal and log-Logistic tended to underestimate the higher thresholds, while log-Cauchy tended to overestimate the most extreme ones.

The authors described a no-trend interpretation of the hurricane data as reasonable but tentative, with the judgment depending on the historical-period fit, normalization choices and later-period comparison. The simulations covered selected distributions and contamination scenarios, while the real-data evidence consisted of one Google-return dataset and normalized hurricane-loss data. The paper's results remain tied to the distributions, quantile designs and data sets it examined.

How far the findings go

In its concluding remarks, the paper says log-quantile estimators are more efficient while retaining robustness, with log-gQLS often reaching 90% or higher relative efficiency versus maximum likelihood in the studied settings. It also reports that goodness-of-fit power improves when sample size exceeds 100 and that computation for samples with billions of observations can take 2 to 3 minutes. These are conclusions drawn from the mathematical derivations, simulations and selected illustrations.

The document remains an arXiv v1 preprint dated 25 August 2026 and says it is to appear in Variance. No funding source is reported in the supplied text. The theoretical distribution of the goodness-of-fit statistic remains unresolved, and the reported p-values use the specified bootstrap procedure.

Paper data and sources

Original title: Quantile and Log-Quantile Least Squares for Robust-Efficient Fitting and Validation of Log-Location-Scale Loss Models
Authors: Mohammed Adjieteh, Vytaras Brazauskas
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.