Preprint

Preprint: Calibration left an LLM risk score unable to change forecasts

All four fitted LLM weights fell to zero in a next-day market-risk study; a headline-count feature helped one SPY variance test but produced mixed results elsewhere.

Calibration left the daily LLM importance score unable to change next-day market-risk forecasts in the study, according to a preprint. All four fitted LLM weights reached zero, so the 856 later scores could not affect the evaluation.

The augmented forecasts therefore exactly matched the option-implied-volatility baselines for every input. Changing the score from 0 to 100 changed nothing, and every per-date loss contrast and bootstrap resample was zero.

The test was designed to add a news signal

The planned test was whether the daily score could improve next-day SPY risk forecasts beyond an option-implied-volatility baseline. QQQ, paired with VXN, served as a correlated transfer check.

The news and market-data corpus covered 2017 through June 2026, with at most 25 selected headlines per trading day. After one retry, the later evaluation retained 856 strict-valid dates out of 865 eligible dates.

The models were fit on 250 usable dates from 2022 and then locked for 2023–2026. The prespecified news coefficients were constrained to between 0 and 5; binary forecasts used log loss and continuous variance forecasts used Gaussian QLIKE. Uncertainty intervals came from 20,000 moving-block bootstrap draws using 20-day blocks for 95% intervals.

The sequence mattered

Full-history scoring came before the 2022 calibration, so the later LLM scores were acquired before the fit showed that the models could not use them.

In a post-hoc repair, the allowed importance-slope range changed from [0, 5] to [−5, 5]. All four optima moved to interior negative values, but every later-period improvement estimate was nonpositive, and no endpoint produced a familywise-corrected improvement.

A simple count gave a mixed picture

A headline-count feature improved the SPY continuous-variance forecast loss by 0.001720. Its 95% familywise interval ran from 0.000719 to 0.002830, excluding zero.

The count was not generally superior: it harmed SPY binary loss and QQQ variance, while its QQQ binary gain was not familywise-resolved.

A checkpoint would have stopped the run

The proposed viability checkpoint would have classified all four mappings as calibration_nonviable, stopped after the 2022 fit and avoided the $16.17 full-history phase. The authors describe this as a prospective design counterfactual; passing the checkpoint would not predict an out-of-sample improvement.

The result comes with procedural caveats

The manuscript describes its design as internally prespecified and hash-frozen rather than preregistered. A coverage correction made after inspecting LLM outputs changed the recorded decision from STOP to GO, and exactly 16 token-truncated outputs were retried after the market result was known.

No ex-ante power analysis was prespecified, and the zero-slope setup left the prespecified comparison with zero ability to detect a score effect.

The public ancillary archive contains redistributable plans, code, derived rows, machine results, verification outputs and a SHA-256 manifest. It supports statistical recalculation from derived rows, but not full pipeline replay; restricted headline text, provider requests and responses, and raw market downloads are excluded.

The manuscript is identified as arXiv:2608.20304v1 and dated 20 August 2026. The author funded the API charges, reported no external funding and reported no competing interests.

Paper data and sources

Original title: Calibration-Induced Degeneracy in LLM Financial Forecasting: An Audit-Trailed Case Study on Next-Day Market Risk
Authors: Arin Mohanty
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.