An arXiv preprint reports that a depth-averaged Bayesian estimator matched or beat every classical method it faced, within one standard error, on eight of 11 synthetic targets with large alphabets. On a uniform target, however, Good–Turing reached 0.005 bits of loss, compared with 0.174 bits for the best prior in the layered family.
A model that averages its own depths
The layered simplex architecture, or LSA, starts with L independent uniform points from the probability simplex—a set of probabilities that adds up to one. It multiplies the points coordinate by coordinate and then renormalizes the result, creating a family of priors with different depths.
Instead of choosing one depth in advance, the reported predictor averages across depths. It stayed within at most (log₂ Lmax)/N bits per symbol of the best single depth on every target. At d = 10⁶ and Zipf α = 3, the best-depth envelope was approximately 49 times better than L = 1.
The calculations used explicit formulas for regret and predictive probabilities; Monte Carlo sampling of count profiles was the only source of statistical error reported. The implementation was checked against closed forms, independent numerical integration and exact identities.
A theory built around discovering new symbols
For Zipf targets with α > 1 and logarithmic depth, the paper gives a leading regret approximation whose sample-size exponent is 1 − 1/α. The authors interpret that exponent as the rate at which new symbols are discovered, linking the estimator’s loss to how quickly the observed stream reveals previously unseen parts of the alphabet.
The finite-grid fits were close to that picture for α ≥ 2: alphabet-scaling slopes ranged from 0.95 to 1.09. For α = 1.5, the slope fell from 0.87 at N = 10² to 0.58 at N = 10⁴. Predicted and measured data exponents differed by at most 0.002 at α = 2, 0.007 at α = 3 and 0.013 to 0.019 at α = 4.
Competitive, but not universal
In the benchmark, divergence between the true and estimated distributions was averaged over 20 independent trials, with all estimators using common sampled count data. The depth-averaged LSA beat or tied every classical method within one standard error on eight of 11 targets. On Zipf α = 1.5 at n = 2 × 10⁴, fixed L = 5 and Good–Turing both recorded 0.066 bits, while the natural oracle recorded 0.061 bits.
The exceptions were uniform, step and Zipf α = 1. The uniform case shows the trade-off clearly: Good–Turing reached 0.005 bits, while the best LSA-family prior reached 0.174 bits. The paper attributes the gap to the exchangeable family’s less efficient use of count-frequency information.
A narrow edge on one real text
The real-text check used the entire King James Bible: 915,860 tokens, a fixed vocabulary of 100,000, and 13,550 observed word and punctuation types—about 14% of the alphabet. On this corpus, the depth-averaged predictor was best at every prefix, with redundancy of 0.095 bits per token versus 0.097 for Good–Turing.
At the full sample, the model assigned posterior mass of 0.9999991 to L = 15. When order-one memory was added, code length fell from 8.572 to 7.100 bits per token; order-one KT coding reached 9.78 bits per token and was worse than the memoryless layered model as the partitions were refined.
Where the evidence stops
The manuscript is an arXiv preprint, arXiv:2608.19908v1, dated 20 August 2026 in its front matter. Its alphabet results are finite-grid fitted slopes, and its data-scaling comparison is a finite-size prediction; together, they support the proposed pattern over the tested settings rather than proving a law for every alphabet or source.
The benchmark therefore presents LSA as a competitive estimator in the tested settings, not as a universal replacement for Good–Turing: its own exceptions and the uniform-target gap remain material.
Paper data and sources
Original title: A Layered Simplex Architecture for Large Alphabets
Authors: Meir Feder, Yaniv Fogel, Ruediger Urbanke
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text