A preprint dated 26 August 2026 describes an FPGA system for predicting high-performance computing job runtimes. Its headline comparison reported up to 63.85% lower mean absolute error, or MAE, than users' runtime estimates and up to 71.88% lower MAE than a static model. In a separate end-to-end benchmark using the tested ALCF configuration, the system averaged a 17-fold speedup over an AVX software baseline and processed up to 1,711 jobs per second.
Built around information available at submission
The study asks whether a software ensemble can be adapted to fixed-point hardware without overflow, predict job runtimes from scheduling parameters known at submission, and adapt to temporal changes in runtime activity.
For the final analysis, the predictor used six submission-time fields: NODE_TYPE, QUEUE_NAME, NODES_REQUESTED, CORES_REQUESTED, GPUS_REQUESTED and REQUESTED_CORE_HOURS. It also purged rows with zero or negative runtimes and removed post-submission telemetry to avoid failed-record bias and temporal leakage.
The historical data came from the Argonne Leadership Computing Facility, MIT Supercloud and UIUC Blue Waters. ALCF had 4,003,802 raw jobs and 3,978,553 retained jobs. MIT Supercloud had 395,914 raw and 309,037 retained, while UIUC Blue Waters had 8,750,131 raw and 5,379,998 retained. For additive updates, the recent-job buckets were 45,000 jobs for ALCF, 125,000 for MIT Supercloud and 35,000 for Blue Waters.
The hardware used two AMD Alveo U55C FPGAs. One handled Conifer-FPU inference, the prediction step, and the other ran the STANNIC scheduler. In synthetic drift tests, the additive model, called ADD, started with a 64-tree CORE model and added 64 trees trained on a recent bucket of drifted jobs. The evaluation used five-fold cross-validation with stratified sampling.
The gains varied with the workload
On synthetic shifts built from ALCF data, ADD had lower MAE than CORE in all 10 settings. The mean reduction was 33.8%, and the median reduction was 25.9%. The reported Wilcoxon comparison gave p = 0.002, with a 95% confidence interval from -54.1% to -13.6%.
The pattern was less consistent in cross-dataset synthetic-shift tests. ADD reduced Blue Waters MAE in 8 of 10 settings, with mean and median reductions of 46.3% and 48.9%. On MIT Supercloud, it improved 5 of 10 settings, with a 12.0% mean reduction and a 2.7% median reduction. The reported p values were 0.020 for Blue Waters and 0.275 for MIT Supercloud.
Performance also depended on the predictor. Random Forest had the lowest MAE on ALCF and UIUC Blue Waters, while XGBoost was best on MIT Supercloud at 1.77 Core-Days. The comparison did not identify one predictor as the lowest-error option on every dataset.
In chronological ALCF validation, ADD had an advantage by Window 9. Its MAE was approximately 222 Core-Days, versus roughly 460 for CORE and 539 for the RECENT-ONLY model. The CORE/ADD comparison gave p = 0.05.
That adaptation was not smooth under abrupt workload changes. ADD can temporarily overcorrect, a behavior the paper calls inertia, before recovering toward the CORE trend after an update.
Fast processing came with a memory tradeoff
The latency tests favored streaming at the tested arrival rates. Batch latency fell from 245.0 milliseconds at 102 jobs per second to 2.4 milliseconds at 104 jobs per second. Batch processing accounted for 99.8% of total latency at the lower rate, while streaming stayed below 4 microseconds.
End to end, the system averaged a 17-fold speedup over AVX and reached up to 1,711 jobs per second. The measurements also recorded higher host memory use: peak resident set size was approximately 253 MB, versus approximately 35 MB for AVX, or about seven times higher. Both FPGAs ran at 371.47 MHz, with the inference module using at most 4% of lookup tables and the scheduling module at most 2%.
The result is bounded by the test setup. The end-to-end speed and throughput figures came from the tested ALCF configuration, while the memory and FPGA resource measurements were specific to the AMD Alveo U55C implementation. The chronological tests also showed that the additive policy can temporarily overcorrect after abrupt shifts before recovering toward the CORE trend.
Paper data and sources
Original title: BOOSTEDSOSA: Accelerated Inferencing for Low Variance Stochastic Online Scheduling
Authors: Adam H. Ross, Riccardo Revalor, Aryan Singh et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text