A new preprint introduces EvEMTBench, an open dataset built from simulated power-system faults and operating events for machine-learning-based protection research. The paper presents it as a way to address the shortage of open material for this work. Its central contents are synchronized, grid-wide, point-on-wave voltage and current measurements sampled at 9,600 hertz, with comprehensive labels across a wide range of events.
The key caution is that the records are produced by simulation. The authors warn that a gap between simulated signals and real-world conditions may reduce performance in practical applications and may require domain adaptation or fine-tuning for a particular grid. The dataset is therefore presented as material for future evaluation, not as a report of machine-learning performance.
The signal detail is the point
The records were generated automatically in DIgSILENT PowerFactory using detailed electromagnetic-transient models. Those models include both synchronous-machine generation and inverter-based generation. Electromagnetic-transient simulation is designed to capture changing electrical signals as the simulated system passes through a fault or another event.
For each case, voltage and current measurements are exported from each main-voltage cubicle for the final 0.5 seconds of the run. The signals are sampled at 9,600 Hz for both 50-Hz and 60-Hz grids. Because they are point-on-wave measurements, they preserve the changing waveform itself, while the measurements remain synchronized across the grid.
The event library spans fault and operating events. It includes conventional and inverter-based generation loss, standard short circuits, high-impedance ground faults and incipient faults.
The paper offers three dataset types: multigrid, benchmark and adaptgrid. They are set up respectively for testing cross-grid generalization, benchmarking and grid-specific adaptation. These describe the intended uses of the data designs; they are not reported performance results.
The reported generation design specifies 10,000 events across 105 topologies per voltage level for multigrid, and 10,000 events in each selected literature grid for adaptgrid. Those figures describe the planned generation scale. The final exported totals are slightly lower when simulations fail, and the paper does not report the failure counts or exact final record totals.
Designed to stay within bounds
To keep randomly generated multigrid cases within stated operating ranges, the framework checked that the load-flow calculation converged, imposed a maximum overload condition of 110% and constrained voltage deviations. The permitted range was 0.95 to 1.05 per unit for medium-voltage grids and 0.90 to 1.10 per unit for high- and extra-high-voltage grids. Per unit is a normalized way of expressing voltage against a system’s nominal value.
The authors then examined descriptive plots rather than reporting formal statistical tests. In plots based on the first two cycles, a large majority of root-mean-square bus-voltage values sat close to the nominal level of 1.0 per unit, while a smaller share varied more. The TestGrid110kV case was described as a high-load scenario with a lower median voltage.
They also inspected example transient waveforms and report that the signals produced by the sampled events resemble the qualitative waveform patterns expected from the literature. This is an engineering plausibility check, rather than a quantitative validation against field measurements.
Only cases that completed initialization and the electromagnetic-transient simulation were exported. Nonconvergent or numerically unstable runs were discarded, leaving slightly fewer than 10,000 records per dataset under the reported design. Since the paper does not give the number removed, the release description cannot establish a failure rate.
An infrastructure release, not a model result
EvEMTBench combines the simulated signals with comprehensive labels for the events. The release is organized into data, labels and graphs folders, and the paper says the package comprises 11 compressed archives.
The dataset is publicly available through FAUDataCloud. Its open format is intended to give researchers a common basis for machine-learning protection work, including the cross-grid, benchmarking and grid-specific uses defined by the three dataset types.
The paper does not report that a machine-learning model trained on these records detects, classifies, localizes or predicts faults accurately. It reports the creation of the dataset and descriptive checks of simulated behavior, while the proposed benchmarking role remains an opportunity for later evaluation.
The gap between a simulation and a grid
The main uncertainty is the path from synthetic signals to real systems. The authors say performance in real-world applications may deteriorate across the simulation-to-real gap and may require domain adaptation or grid-specific fine-tuning. Their waveform check reports qualitative agreement with expected literature patterns, not independent field validation.
The filtering step adds another qualification. Unsuccessful initialization, nonconvergent cases and numerically unstable simulations are absent from the exported records. The result is slightly less than 10,000 records per dataset, but the supplied analysis does not provide the failure counts or final usable-record totals.
Taken on its own terms, EvEMTBench is an open foundation for controlled studies of power-system protection methods. Its value will depend on how well models trained on the simulated records transfer to real waveforms and whether later evaluations supply the performance evidence that this dataset release does not yet provide.
What has been released
The document is a preprint identified as arXiv:2608.19777v1 and dated 20 August 2026. The paper reports funding from the Deutsche Forschungsgemeinschaft, the German Research Foundation, under identifier 535389056.
Paper data and sources
Original title: A simulation based dataset of faults and events for machine learning in power systems
Authors: Georg Kordowich, Jonathan Loebel, Julian Oelhaf et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text