A preprint reports that two lightweight models kept nearly the same AUC—the study’s overall ranking measure—as a much larger conditional-flow model across four simulated new-physics benchmarks. The students’ AUCs were within at most 1% of the teacher’s, and the teacher had the highest AUC in all four benchmarks.
The study asks whether that performance can be retained when a large conditional normalizing flow—a model built around event likelihoods—is distilled into FPGA-compatible models for real-time LHC anomaly detection, including cases where input features may be missing. The two students were a boosted decision tree (BDT) and a dense neural network (DNN).
Four simulated signals set the test
The evaluation used simulated and pre-filtered proton–proton collisions containing an electron or muon. The Standard Model sample contained 13.5 million events, split 40% for training, 10% for validation and 50% for testing. Four additional benchmarks represented LQ → bτ at 80 GeV, A → 4l at 50 GeV, h+ → τ ν at 60 GeV and h0 → τ τ at 60 GeV.
Each event was represented by a 55-component unnormalized feature vector. The engineered inputs comprised 73 normalized particle-level features plus three multiplicities, while the conditional flow modeled the event features given those multiplicities. The students were trained to regress the flow’s joint negative log-likelihood with a weighted squared-error objective and normalized inverse-likelihood importance weighting.
The teacher had 1.3 million trainable parameters. The BDT used post-training quantization (PTQ); the DNN used HGQ quantization-aware training and two hidden layers of 32 nodes each.
Overall scores stayed close
Across the signals in the order LQ → bτ, A → 4l, h+ → τ ν and h0 → τ τ, the teacher’s AUCs were 88%, 90%, 90% and 76%. Both the BDT and DNN reported 88%, 89%, 89% and 76%, respectively.
At the fixed false-positive rate of 10−5, the teacher’s true-positive rates were 0.05%, 1.5%, 0.05% and 0.10% in the same order. The BDT’s were 0.04%, 1.1%, 0.05% and 0.10%, while the DNN’s were 0.04%, 0.58%, 0.03% and 0.07%.
Hardware figures come from synthesis
For a Xilinx Virtex UltraScale+ FPGA at 200 MHz, RTL synthesis estimated 5-nanosecond latency for the BDT and 25-nanosecond latency for the DNN. Each was reported with an initiation interval of one clock, and the quoted resource use for both was below 1% of the FPGA.
The hardware figures are estimates rather than measurements from an operational trigger. The study’s evidence also comes from simulated, pre-filtered data and four benchmark signals, so it does not establish performance on live LHC data or across the full range of possible new-physics signatures.
Questions the study leaves open
No confidence intervals or repeated-training summaries were reported. The paper notes that the extreme-tail true-positive rate is sensitive to finite statistics, feature engineering, random initialization and hyperparameter choice.
The authors describe the approach as a practical path to real-time LHC hardware-trigger use. For now, the reported result is a simulation and synthesis result, leaving its behavior on real data and its measured system-level performance after deployment as open questions.
Paper data and sources
Original title: Distilling Normalizing Flows for Real-Time Anomaly Detection at the LHC
Authors: Tara P. A. Tahseen, Noah Clarke Hall, Nikolaos Konstantinidis, Paula Martínez Suárez
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text