Preprint

In one five-tier failure simulation, Standard Retry had lowest success

Preprint: A review of 200 GitHub repositories was paired with a three-scenario simulation of four retry strategies.

Standard Retry recorded the lowest mean success among the four strategies tested in S3, one of three modeled scenarios in the study. Mean success was 41.5%, compared with 55.4% for No Retry, 55.3% for Circuit Breaker and 54.9% for Adaptive Retry Budgeting, or ARB. The paper described Standard Retry’s result as a 25% relative degradation versus No Retry.

The study also reported a retry amplification factor, or RAF, in the same scenario. Observed RAF was 1.34 for Standard Retry, compared with 1.00 for No Retry, 1.00 for Circuit Breaker and 1.01 for ARB. Standard Retry therefore had the highest observed RAF among the compared strategies in S3.

A code review puts the simulation in context

The simulation was paired with a review of GitHub repositories. The search targeted actively developed projects with more than 50 stars and descriptions or topics referring to microservices or distributed systems. It returned 1,000 candidates, and the first 200 entries analyzed were all Python.

To identify retry behavior, the authors used regex-based static analysis. The process produced 162 raw detections. After 41 non-production detections were excluded and eight repeated detections were collapsed, 113 production configurations remained.

Explicit retry logic was detected in 23 of the 200 repositories, an 11.5% detection rate. A seeded check of 30 initially undetected repositories found retry logic in 10, producing a false-negative rate of 33.3%, with an interval of 19.2% to 51.2%. A separate check of 30 detected items found 29 genuine, giving 96.7% precision. Extrapolating from the false-negative rate placed estimated prevalence near 41%, with a range of 28.5% to 56.8%.

Among the 113 cleaned production configurations, 35 had no backoff, equal to 31.0%. Linear backoff accounted for 47.8%, exponential backoff for 21.2%, and verified jitter appeared in one configuration, or 0.9%.

At the project level, the five reported anti-patterns were measured among the 23 projects with detected retry logic. No backoff appeared in 60.9%, missing jitter under the paper’s broad definition in 95.7%, and aggressive retry in 30.4%. Static configuration and no cross-service coordination were each present in 100% of those projects.

How the model was built

To examine how retry choices interact across tiers during partial failures, the authors modeled a five-tier linear chain. Each service had capacity for 1,000 requests per second against a 500-request-per-second base load, setting initial utilization at 50%. Each tier used three retries with exponential backoff delays of 100, 200 and 400 milliseconds, plus jitter.

The evaluation compared No Retry, Standard Retry, Circuit Breaker and ARB under scenarios S1, S2 and S3, with 100 trials per strategy and configuration. Circuit Breaker was set to trip at a 50% failure threshold and remain open for 30 seconds.

ARB assigned each service tier a retry budget expressed as a fraction of base load. It adjusted that budget using the observed failure rate and used an overload signal so upstream tiers could cut their budget before observing their own failures.

The model and simulator did not line up exactly

The analytical model predicted RAF values of 1.88 for S1, 1.23 to 6.42 for S2 and 10.30 for S3. It described correlated failure as the dangerous case because amplification multiplies across affected tiers. These were analytical predictions, not simulator output.

The simulation produced lower RAF values for Standard Retry, ranging from 1.18 to 1.34. Those figures were below the model projections of 6.42 at the S2 peak and 10.30 for S3. The paper attributed the gap mainly to finite queues that shed load.

The S3 results did not show a higher raw success rate for ARB than for Circuit Breaker: the figures were 54.9% and 55.3%, respectively. ARB’s method centered on adjusting retry budgets through observed failures and overload signals.

What the figures cover

The repository findings came from the first 200 entries in a 1,000-candidate search, all Python. The validation used samples of 30 detected items and 30 initially undetected repositories. The simulation used one five-tier linear chain, three scenarios and 100 trials per strategy and configuration, so its success and RAF figures describe that modeled setup.

The work is an arXiv version-1 preprint dated 26 August 2026. The supplied front matter lists no funding or conflict-of-interest statement.

Paper data and sources

Original title: Retry Amplification in Distributed Systems: A Systematic Analysis of Retry Policies and Their Role in Cascading Failures
Authors: Rishabh Mehan, Jasmit Kaur Saluja
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: 10.2139/ssrn.6313332
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.