A compact sequence model called STRATA beat six parameter-matched alternatives at ranking stocks by next-day returns in a 2024 test year, according to a preprint. The result became much less encouraging when the researchers separated the overnight move from the regular session. STRATA's overnight signal had a long-short spread of 75.6 basis points a day and a Sharpe ratio of 24.13, while the executable session had a spread of minus 1.6 basis points a day and a Sharpe ratio of minus 0.32. Its executable-session spread had t = -0.37, and the authors say none of the models supported a profitable-strategy interpretation.
On the full close-to-close evaluation, STRATA led on all four reported measures after the scores were adjusted for eight price-volume style controls. Rank IC, the paper's daily measure of whether higher-scored stocks tended to have higher next-day returns, averaged 0.0728. IC IR, a summary of how consistently that relationship held across days, was 1.128. The Signal LS Sharpe was 12.85, and the stress-day IC IR was 1.030. Each figure is a mean across three random seeds; the corresponding seed-to-seed standard deviations were 0.0020, 0.113, 0.62 and 0.047.
A test built around raw market data
The model takes five complete trading days of raw five-minute bars and order-book sequences and produces its score only after the final bar of the signal day. The target is the next-day adjusted close-to-close return. The inputs stay as unadjusted exchange fields: the pipeline applies fixed per-field rescaling, a log(1 + x) transform to price and size, per-field standardisation and zero-filling for remaining missing values, with all preprocessing statistics estimated from the training split. The design maps raw data directly to a ranking without hand-crafted features.
STRATA has 244,633 parameters. Its stem uses five branches, including four causal depthwise convolutions with zero-sum initial effective kernels and a cross-field contrast. Four selective state-space blocks then feed a four-path readout. The six comparison models were MLP, LSTM, GRU, TCN, Transformer and Mamba, with each baseline kept within 5% of STRATA's parameter count. All seven arms used the same data, preprocessing, loss, optimiser, training budget, early stopping and evaluator.
Training used a shared objective that combined a point-prediction anchor, a soft Spearman surrogate, a Pearson term and a one-sided excess-kurtosis penalty. That shared setup makes the comparison a test of the encoders under common conditions, but it also means the reported results are not a set of independently tuned systems.
The data came from a point-in-time universe of close to 1,000 names per day. The reported median ranged from 995 to 999 across the splits, with a minimum of 878 names. Validation used 2023, covering 242 days and 240,204 samples. The 2024 test year had 241 evaluable days, while truncation left 955,061 training samples.
The lead was not simply a style bet
STRATA's reported advantage over GRU was clear on the paper's headline numbers. Rank IC was 0.0728 for STRATA versus 0.0634 for GRU, a 14.9% margin. Signal LS Sharpe was 12.85 versus 10.92, a 17.7% margin, and stress IC IR was 1.030 versus 0.909, a 13.4% margin. In the bundled Mamba comparison, Mamba's rank IC was 0.0615 against STRATA's 0.0728, a reported gap of 18.5%. Because that comparison changes several design elements together, it does not identify which individual component accounts for the difference.
To address the possibility that the model was mainly picking up familiar style patterns, the evaluation removed non-finite scores, limited extreme values, standardised scores across each day's stocks and replaced the original scores with ordinary-least-squares residuals. In this context, residualised means the part of the ranking left after the eight controls were fitted and removed. STRATA's residualised rank IC was 0.073, while the share of its score variance explained by those controls was 0.172. That was below the 0.180 to 0.220 range for the five sequence baselines. MLP had an even lower style-explained variance of 0.157, but its residualised rank IC was 0.049. The authors limit this conclusion to the eight price-volume controls.
The significance checks pointed the same way, but their strength differed by unit of analysis. Pairing 241 daily ICs, the Newey-West statistics against every baseline ranged from 4.40 to 6.29, with p < 0.001. Tests based on the three random seeds gave p ≤ 0.014 against every baseline, although the paper describes seed-level inference as coarse.
The trading lesson is in the clock
The caution comes from the timing of the target. Because the score is produced after the last bar of the signal day, the close-to-close label includes an overnight move that occurs before the score exists. STRATA's overnight segment had rank IC 0.1275, IC IR 2.431, a long-short spread of plus 75.6 basis points a day, t = 14.26 and Sharpe 24.13. During the executable session, those figures fell to rank IC 0.0215, IC IR 0.358, a spread of minus 1.6 basis points a day, t = -0.37 and Sharpe -0.32.
STRATA still ranked first when the target was restricted to the executable session: its rank IC was 0.0215 versus GRU's 0.0162, a relative margin of 32.7%. But the ranking lead did not translate into a positive spread. The paper reports STRATA's executable-session spread as minus 1.6 basis points a day, with t = -0.37; MLP and Transformer were significantly negative, with t = -2.99 and -2.75. The authors therefore treat the result as evidence about model ranking, not a profitable trading strategy.
What the result can and cannot support
The remaining questions are substantial. The supplied analysis covers one Chinese equity market, one five-minute frequency, one-day horizon and one held-out test year, so it does not establish that the ordering will transfer elsewhere. The eight controls do not test industry, valuation, profitability, growth or leverage exposures. Reported diagnostics also omit costs, slippage, market impact, price limits and borrowing constraints. The use of three seeds makes seed-level uncertainty coarse, and the stress-day selection is a restricted robustness check based on a deliberately non-causal detector.
The preprint, labelled arXiv:2608.28060v2 and dated 2 September 2026, presents STRATA as a leading controlled benchmark under the stated evaluation. Its authors say the evidence supports a prediction-model comparison and that the Mamba result needs ablation before component-level claims can be made. No funding statement or conflict-of-interest disclosure is reported in the supplied text.
Paper data and sources
Original title: A Compact Selective State-Space Model for Cross-Sectional Stock Return Ranking from Raw Intraday Bars
Authors: Mingju Chen, Enze Zhang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text