A preprint evaluating HRV Studio found close agreement with NeuroKit2 for many heart-rate variability calculations when the software used matched settings. The strongest matches were in time-domain and nonlinear measures such as RMSSD, SDNN and SD1, while total power and very-low-frequency (VLF) spectral measures showed larger differences. Together, the results describe reproducible calculations under matched conditions, but not a uniform level of agreement across every measure.
The work was a software validation built around computation and quality control, rather than a clinical validation. Its staged framework combined cross-platform agreement with spectral-method checks, synthetic interval perturbations, duration testing and arrhythmia-focused stress tests. The authors treated the synthetic and arrhythmia exercises as tests of computational or engineering behavior, so a finite output and a warning do not establish that a physiological interpretation is valid.
The clearest agreement came in time-domain measures
For the main comparison, HRV Studio was tested against NeuroKit2 on 15,179 five-minute recordings, producing 106,253 rows of calculated metrics. The primary quantitative endpoint was comparator-referenced relative error: how far an output was from the comparison result, expressed as a percentage. Median errors were 0.18% for high-frequency (HF) power, 1.35% for low-frequency (LF) power, 1.41% for the LF/HF ratio, 0.42% for LFnu and 0.81% for HFnu. Total power had a median relative error of 13.57%, while VLF had an error of 37.79%.
Time-domain and nonlinear measures were closer still. RMSSD, SDNN and SD1 had near-zero median errors and Pearson correlations of 1.000; pNN50's median error was 0.135%. SD2 initially showed a 2.729% median error with a correlation of 0.988, but after its calculation was harmonized, the error fell to 0.053% and the correlation rose to 0.99998. That change shows how an implementation detail can become part of the cross-platform result.
A second run used 7,598 ten-minute recordings and produced 53,186 metric rows. It preserved the same metric-by-metric pattern, while total-power error fell to 6.03% and VLF error to 22.72%. Correlations were 0.9998 or higher for most listed measures, 0.999 for total power and 0.998 for VLF. The longer sensitivity run reduced the two largest discrepancies, but did not remove them.
Matching intervals narrowed the Kubios gap
The external Kubios benchmark was smaller and more targeted: it contained 44 manually reviewed recordings, and HRV Studio was rerun on the exact NN sequences selected by Kubios. Aligning those retained intervals made the comparison focus more directly on differences in the calculations themselves.
After NN-sequence harmonization, median errors were 1.46% for LFnu, 3.79% for HFnu, 5.55% for LF/HF, 2.51% for total power, 3.49% for LF and 2.54% for HF. VLF remained the largest residual error at 14.01%, while Pearson correlations across the listed measures ranged from 0.973 to 0.998. The frequency bands in this check followed Kubios-compatible rather than standard HRV Studio conventions.
Time-domain and nonlinear outputs in the same benchmark were even closer. Median errors for RMSSD, SDNN, pNN50 and SD1 were below 0.001%; SD2 had 0.0705% error and a Pearson correlation of 0.9999. The result was strongest when both programs worked from the same NN sequence.
Frequency methods did not behave the same
The study also compared different ways of estimating the spectrum. In the targeted Smoothness Priors analysis, autoregressive (AR) estimates had median errors from 1.16% to 4.38% across the reported measures, while fast Fourier transform (FFT) errors ranged from 4.10% to 18.41%. FFT VLF had a Pearson correlation of 0.586, compared with 0.989 to 0.999 for AR. Some Kubios FFT implementation details, including the window function, were not exposed.
All 100 files in the FFT/AR method checks produced finite outputs. But the diagnostics separated numerical completion from methodological stability: FFT had a median power-spectral-density-to-variance ratio of 96.757 and instability flags in all 100 files. AR had a median ratio of 1.000 and flags in three of 100 files, but its results varied by more than 50% across model orders in 25 files. Welch was retained as the primary reporting method.
Short recordings raised the biggest warnings
Recording length mattered especially for frequency ratios. In a separate duration test, LF/HF median relative error was 60.29% at 30 seconds, 27.66% at 60 seconds, 28.61% at two minutes and 11.98% at five minutes, before edging up to 12.57% at 10 minutes. VLF was unavailable at 30 seconds; its error was 100% at 60 seconds, 33.71% at five minutes and 38.75% at 10 minutes. The numbers did not decline steadily as duration increased.
Across 60 synthetic perturbation cases, outputs were finite before and after correction, and QC indicators were present in all 60 corrected cases. Correction changed RMSSD and several spectral measures. These were computational stress tests using controlled interval changes, not tests that establish physiological validity.
The MIT-BIH arrhythmia stress test used 12 segments from eight recordings. After correction, finite outputs and QC warnings were present in all 12 segments. The authors treated this as engineering stress testing rather than clinical validation, so it does not show that HRV metrics from arrhythmic recordings are physiologically reliable.
A validation with clear boundaries
Taken together, the evidence supports a bounded conclusion. HRV Studio showed strong agreement with NeuroKit2 and Kubios for many time-domain, nonlinear and selected spectral measures when effective NN sequences, preprocessing and spectral conventions were harmonized. VLF and total power remained more sensitive to processing and duration, while FFT and AR estimates remained method-dependent.
For researchers comparing HRV software, the practical message is to treat interval selection, processing settings, estimator choice, recording length and QC warnings as part of the result. A finite numerical output alone does not establish that the underlying physiological interpretation is trustworthy.
The record is a preprint, and the supplied front matter does not state funding. The evidence is about software agreement and QC behavior under the tested conditions, not clinical performance, diagnostic validity or patient benefit.
Paper data and sources
Original title: Validation of HRV Studio: A Transparent and Quality-Control-Aware Platform for Heart Rate Variability Analysis
Authors: Cyrus Mexon Evrard Djindot, Faliang Liu, Sylvain Laborde et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text