Positive predictive scores did not translate into positive risk-adjusted trading performance in a one-time historical test, according to a new arXiv preprint. All 23 evaluated runs had positive IC, but each had negative net Sharpe under primary costs. None of 460 fee-slippage cells had positive Sharpe.
The distinction matters because IC is a predictive ranking score, while net Sharpe is a risk-adjusted performance measure. In this study, the two pointed in different directions: candidates ranked well against later outcomes, but the reported cost-adjusted performance remained negative.
The data had to be rebuilt around availability
The paper is arXiv version 1, dated 26 Aug. 2026. It used public Binance BTCUSDT USD-M perpetual-futures archives, including 826 ZIP files verified with official checksums. The analysis centered on whether each candidate could be judged using only information available when its decision was made.
The initial protocol required trade, mark, index and open-interest streams to share one exact time grid without a break. That intersection failed after 304.57 continuous days. The researchers revised the protocol to a five-minute primary availability mask, which excluded decisions when required data was not yet available. The retained sample contained 209,951 eligible decisions, 288 masked decisions and 727 complete UTC days.
Bar data counted as available at close time plus 1 millisecond. Realized funding was assigned event time plus five minutes as a research assumption, with admission checked at zero, five, 10 and 15 minutes. The complete-day count stayed at 727 across those settings.
Open interest was optional and was disabled because its publication time was unverified. Its archive contained nine missing records, three off-grid records, one conflicting duplicate group and 501 archive-order reversals. The defects were retained rather than repaired.
The audit passed its finite test
The deterministic evaluator parsed expressions, traced their input lineage, applied availability masks, audited execution, computed labels and metrics, made a selection decision and wrote the result to a ledger. The auditor rejected five known classes of violations.
The conformance test set 40 illegal executable templates against 40 legal ones, with eight examples in each violation class. Its acceptance rules required at least 38 of 40 illegal templates to be detected, no more than two legal templates to be rejected, and exact reasons for all eight examples in every class.
On that finite benchmark, all 40 illegal templates were detected, all eight examples in each class were caught, and none of the 40 legal templates was rejected. Wilson 95% intervals put the illegal-template detection rate between 0.9124 and 1.0000, and the legal-template false-rejection rate between 0.0000 and 0.0876. These figures describe the tested injected violations; they do not establish how the auditor would handle unknown leakage.
False passes fell in the null tests
The null controls were designed without a genuine signal. They used five moving-block permutations of public 15-minute BTC data and five synthetic paths with heavy tails and clustered volatility. Each path tested 100 candidates: 80 legal templates and 20 known-violation templates.
At 5% missingness with availability masking, the mean false-pass rate was 29.10% with no audit, 24.79% with a basic audit and 6.25% with the full audit. The paired full-versus-no-audit difference was -0.2285, with a path-bootstrap 95% interval from -0.2595 to -0.2012. The reported relative reduction was 78.5%, and the exact two-sided sign-test p-value was 0.001953.
The comparison was descriptive. In the executed masking test, exact-grid deletion and availability masking were indistinguishable, while backward fill was not uniformly worse. The study therefore did not identify an independent effect of masking.
A matched search found no clear leader
For the head-to-head search, the methods used seeds 11, 23, 37, 53 and 71, with 15- and 60-minute horizons. Each method was allowed 100 valid candidates in each seed-horizon run, or 1,000 valid candidates across the design.
Across 10 runs per method, the audited agent and random search each produced 39 qualified and eight selected candidates. Tree GP produced 29 qualified and eight selected candidates. The audited agent tied random search on qualification yield; superiority was not shown.
The predictive score did not survive the cost test
The economic evaluation used a timestamp-only 60/20/20 split: 436 training days, 145 validation days and 146 test days. A 60-minute purge and embargo separated the partitions. The final partition was a one-time historical holdout, not a prospective test.
At primary costs, mean test IC was 0.2352 for the audited agent, 0.1907 for random search and 0.1550 for tree GP. Their 95% intervals were 0.1551 to 0.3552, 0.1164 to 0.2681 and 0.1262 to 0.1765, respectively. Positive-Sharpe counts were 0/8, 0/8 and 0/7.
The audited agent had the highest listed mean test IC, but no method produced a positive-Sharpe run in the reported table. Taken together, the findings show a mismatch between predictive ranking and cost-adjusted performance in this historical holdout.
The result remains narrowly bounded
The evidence covers one Binance BTCUSDT USD-M market, known injected rules, the executed null controls, the stated search budgets and a retrospective historical holdout. It does not establish cross-market validity, prospective performance, protection against unknown leakage or profitability.
The paper says detailed protocols, additional tables, artifact hashes and exact reproduction commands are provided in the supplement and reproducibility record.
Paper data and sources
Original title: Point-in-Time Audit Before Alpha: Public-Archive Availability and a Negative Matched-Budget Study on BTC Perpetual Futures
Authors: Baocheng Zeng, Jinhao Yang, Peilin Han, Kangnan He
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text