Preprint

A trading framework says average backtests are not enough

Preprint: A conditional model ties tail risk, search costs and position sizing to its assumptions, but its one-index test found no robust position.

A mathematical preprint argues that average backtest returns are not enough to judge a systematic trading rule. Once persistence is treated as a time-invariant mechanism conditional on a hidden, or latent, state, the paper says a quantitative investment system must account for how regimes recur, how much information the data can support, how candidate rules are selected and how positions are sized. In a worked test on one 30-year equity-index run, the full procedure produced no robust position.

The manuscript is labeled arXiv:2608.23416v2 [cs.LG] and dated 2 Sep 2026. It formalizes persistence through five axioms and treats recurrence, invariance, coherence, signal and regime-contingency quantities as declarations rather than estimates. The architecture it derives is therefore conditional on those assumptions and declarations.

A rule for the bad tail

At the heart of the proposal is a change in what counts as a good backtest. The selection theorem says worst-case future loss should be measured by CVaR, a measure focused on the bad tail of outcomes, across risks associated with different regimes at the declared recurrence level. Ranking rules by average backtest performance is treated as the special case in which Lambda equals 1, and the paper says the rankings can reverse.

That choice creates a theoretical cost for searching. In a constructed family of laws satisfying the stated axioms, choosing the rule with the best tail behavior requires a sample bound carrying the recurrence-dependent factor Lambda minus 1. The mean-ranking comparison is stated without that factor. This is a lower bound for a specific family of laws, not a universal estimate for every trading project.

The price of searching

The broader accounting principle is an information budget. Effective sample supports model capacity, the effort spent searching candidate rules and other description length, while persistence is charged according to regime dispersion and label overlap at full weight. The paper says these formulas are exact only for specified model classes or chains and otherwise are bounds or declared quantities.

Five stages, but not a profit claim

From those assumptions, the paper says a correct quantitative-investment procedure needs five stages: a declared representation; capacity-bounded shrinkage and ensemble deployment; contiguous, purged block-CVaR evaluation; budgeted, deflated search; and robust fractional-Kelly sizing. Its theorem says leaving out a stage can worsen risk or expected log growth under a suitable constructed law. The claim is about necessity under an admissible law, not uniqueness of implementation or a required factorization for every good procedure.

What the tests found

The empirical program examined daily timing rules on one equity index, multiple market series, and a cross-sectional library of long-only portfolios and published long-short signals. A tail-price statistic, used to measure the cost of selecting for tail performance, had a median of about 0.9 times Lambda in the daily timing experiment. The same pattern appeared on 19 of 20 market series and in the long-only portfolio library, while published long-short signals were reported at the Gaussian value.

The paper attributes the excess over the Gaussian reference to heteroskedastic score noise, meaning noise in the scores that varies across conditions. It presents these results as descriptive findings from the reported libraries, not as a universal market estimate.

The declaration tests produced a mixed result. At the authors' corrected or honest testing levels, no axiom was refuted on a live market state. But the conservative kappa equals 1 declaration and the exponential-decay instance at its declared rate were rejected, while an initial invariance-defect signal was withdrawn after the testing level was corrected. The authors interpret the record as challenging particular declarations and rate instances rather than the axiomatic structure itself.

A smaller bet, and an empty result

The sizing result caps exposure below full Kelly. It limits the growth-optimal Kelly fraction to the reciprocal of one plus the fraction of the capacity budget already spent. At full budget use, kappa equals 1 and the position is at most half Kelly; in the paper's quadratic model, full Kelly has non-positive expected log growth.

The one-index result is more sobering than the framework's architecture. In the 30-year run at Lambda equals 4, the selected rule failed the search haircut, and no library candidate received a position after the dispersion penalty. The surviving edges reached only about Lambda 1.1 to 1.3. This was one index and one reported configuration, so it is not a general impossibility result for markets.

What remains unproven

The paper draws a firm boundary around its claims. It says the work concerns procedures rather than markets and does not show that the canonical form produces profit. Its listed limitations include worst-case or one-chain theory, noisy block tails, unresolved axiomatisation and an end-to-end result based on one real series.

The acknowledgement discloses that Claude, used through Claude Code, assisted with drafting, review of the current Section 8 and the open-problem list, while content responsibility remained with the author.

Paper data and sources

Original title: The Axiomatic Trader: Latent Regularity, Information Budgets, and the Canonical Form of a Quantitative Investment System
Authors: Jiayu Li
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.