Calibration left the daily LLM importance score unable to change next-day market-risk forecasts in the study, according to a preprint. All four fitted LLM weights reached zero, so the 856 later scores could not affect the evaluation.
The augmented forecasts therefore exactly matched the option-implied-volatility baselines for every input. Changing the score from 0 to 100 changed nothing, and every per-date loss contrast and bootstrap resample was zero.
The test was designed to add a news signal
The planned test was whether the daily score could improve next-day SPY risk forecasts beyond an option-implied-volatility baseline. QQQ, paired with VXN, served as a correlated transfer check.
The news and market-data corpus covered 2017 through June 2026, with at most 25 selected headlines per trading day. After one retry, the later evaluation retained 856 strict-valid dates out of 865 eligible dates.
The models were fit on 250 usable dates from 2022 and then locked for 2023–2026. The prespecified news coefficients were constrained to between 0 and 5; binary forecasts used log loss and continuous variance forecasts used Gaussian QLIKE. Uncertainty intervals came from 20,000 moving-block bootstrap draws using 20-day blocks for 95% intervals.
The sequence mattered
Full-history scoring came before the 2022 calibration, so the later LLM scores were acquired before the fit showed that the models could not use them.
In a post-hoc repair, the allowed importance-slope range changed from [0, 5] to [−5, 5]. All four optima moved to interior negative values, but every later-period improvement estimate was nonpositive, and no endpoint produced a familywise-corrected improvement.
A simple count gave a mixed picture
A headline-count feature improved the SPY continuous-variance forecast loss by 0.001720. Its 95% familywise interval ran from 0.000719 to 0.002830, excluding zero.
The count was not generally superior: it harmed SPY binary loss and QQQ variance, while its QQQ binary gain was not familywise-resolved.
A checkpoint would have stopped the run
The proposed viability checkpoint would have classified all four mappings as calibration_nonviable, stopped after the 2022 fit and avoided the $16.17 full-history phase. The authors describe this as a prospective design counterfactual; passing the checkpoint would not predict an out-of-sample improvement.
The result comes with procedural caveats
The manuscript describes its design as internally prespecified and hash-frozen rather than preregistered. A coverage correction made after inspecting LLM outputs changed the recorded decision from STOP to GO, and exactly 16 token-truncated outputs were retried after the market result was known.
No ex-ante power analysis was prespecified, and the zero-slope setup left the prespecified comparison with zero ability to detect a score effect.
The public ancillary archive contains redistributable plans, code, derived rows, machine results, verification outputs and a SHA-256 manifest. It supports statistical recalculation from derived rows, but not full pipeline replay; restricted headline text, provider requests and responses, and raw market downloads are excluded.
The manuscript is identified as arXiv:2608.20304v1 and dated 20 August 2026. The author funded the API charges, reported no external funding and reported no competing interests.
Paper data and sources
Original title: Calibration-Induced Degeneracy in LLM Financial Forecasting: An Audit-Trailed Case Study on Next-Day Market Risk
Authors: Arin Mohanty
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text