Preprint

Voltage clustering shrinks a year's grid operating points by 99.66%

Preprint: In a synthetic high-renewables grid model, voltage-response clustering preserved annual voltage-security behavior more accurately than injection-based and heuristic sampling.

An arXiv preprint reports that clustering power-grid conditions by their voltage responses reduced a simulated year's operating-point set from 8,760 hourly cases to a 30-point representative set. The reduced set reproduced the full year's steady-state voltage behavior with 98.3% reconstruction accuracy, while shrinking the annual set by 99.66%.

The study asks a straightforward question: is it better to choose representative conditions from voltage-response space, meaning what the network's voltages actually do, than from power-pattern space, meaning the input injections used to create those voltages? In the paper, voltage security is tracked through the distribution of voltage-risk results in normal operation and after a contingency, so the test concerns whether a small sample retains the behavior of the full year.

Building a year of grid conditions

To build that year, the framework used chronological stochastic simulation, followed by a DC unit-commitment step and sequential AC power-flow calculations. The operating state and control quantities were carried from one hour to the next. The case study used an adapted IEEE Voltage Test System with composite load and DER A modeling, and reported instantaneous renewable penetration in the 80% to 90% range across 8,760 hourly operating points.

Each hourly condition was represented by seven system-wide voltage-risk indices derived from the AC response. They covered voltage extremes, the spatial spread of voltage values, the magnitude and extent of violations, and reactive-power margins. The researchers standardized the measures and used principal component analysis, or PCA, to express them in fewer combined dimensions. Four retained components captured 99.24% of the variance; the first three accounted for 54.13%, 33.04% and 11.26%.

Clustering then used Ward-linkage hierarchical agglomerative clustering, or HAC. The algorithm starts with separate groups and repeatedly joins the pair of groups whose merger adds the least within-cluster variation. From each cluster, the method chose one representative operating point: the point closest to that cluster's centroid in PCA space.

The smaller sample faced two rivals

To test whether the feature space mattered, the researchers compared voltage-response HAC with two alternatives. Injection-space clustering used a normalized 74-dimensional vector of net active-power injections, followed by PCA and K-Medoids++. The heuristic method selected points covering high and low load, solar PV and wind levels across seasons, day and night, and weekday, weekend and holiday patterns.

All methods were assessed with a 30-point representative budget. Each selected point was weighted by the number of original hours in its cluster, and the reduced and full distributions were compared through means, standard deviations, quantiles, distributional distances, extreme-tail measures, and pre- and post-contingency behavior. That let the study test both the center of the annual pattern and its more severe edges.

The lead emerged before and after an outage

Before any contingency was applied, voltage-response HAC preserved 96.1% of the distribution's geometry and had an overall reconstruction error of 1.7%. The corresponding error was 4.3% for injection-space K-Medoids++ and 12.5% for heuristic sampling. These figures compare each reduced set with the full simulated annual distribution.

The voltage-focused method also led in the detailed outage test. The most critical screened single-line outage was identified as line 4032-4044. For that outage, heuristic sampling had an error exceeding 0.5, injection-space clustering had an error of about 0.07, and voltage-space HAC had an error of about 0.02, with 93.4% reconstruction. In this comparison, the smaller the error, the more closely the reduced distribution followed the full-set voltage-security pattern.

A robustness check removed the system's only synchronous condenser. After that change, the heuristic and injection-space methods consistently underestimated VSPI, the study's voltage-security measure, while voltage-space clustering remained accurate with a marginal conservative bias. This finding comes from that added scenario within the same modeled system.

A result with clear boundaries

The headline reduction has a narrow evidence base. The study uses one adapted IEEE Voltage Test System and simulated chronological operating data, with renewable penetration in the 80% to 90% range. Its detailed post-contingency comparison keeps one screened line outage, so performance on other network topologies, real operating data and broader contingency sets remains untested.

The results also depend on the selected feature spaces, clustering algorithms and 30-point comparison budget. The paper therefore supports a case-study finding: under the tested conditions, voltage-response clustering preserved more of the simulated voltage-security distribution than the two comparators, but the study does not establish that 30 points is the right choice for every system.

The document is labeled arXiv:2608.28296v1 and dated 28 August 2026. Funding information is not reported in the supplied paper text or metadata.

Paper data and sources

Original title: Hierarchical Agglomerative Clustering for Efficient Annual Voltage Security Assessment in Very-High RES Penetrated Power Systems
Authors: Rock Agon, Robin Preece, Jovica V. Milanovic
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.