Preprint

Preprint suggests dynamic privacy budgets can preserve learning in online tests

A mathematical framework varies privacy risk across visitors, experiments and a firm’s portfolio, while its formal guarantee covers displayed outputs—not later clicks.

An arXiv preprint proposes a way for companies running online experiments to treat privacy risk as a limited budget that can be spent at three levels: on individual visitors, on each experiment and across a firm’s portfolio. In simulations, allowing that spending to change as a test unfolds was associated with a larger performance advantage over a fixed schedule in longer website experiments.

The website-design scenarios used 10,000, 100,000 and 1,000,000 visitors as stand-ins for small, medium and large firms, and tested experiment-level privacy budgets of 0.05, 0.5, 1, 3 and 5. Mean-performance differences between the dynamic and constant strategies were statistically significant at 100,000 and 1,000,000 visitors, although the paper does not report the exact effect sizes or p-values.

How the model spends privacy

At the center is differential privacy, a formal limit on how much a tracker’s beliefs may update after seeing a displayed experimental output. The model links that limit to the explore-versus-exploit balance: stronger privacy requires more random exploration, while weaker privacy permits more exploitation.

The theoretical analysis adds an important qualification: when privacy randomization is matched to the exploration needed for segment-specific learning, the order of the regret bound is unchanged. Here, regret is the modeled loss from not always selecting the best option. A stricter privacy setting can force more exploration and increase regret.

More choices change the calculation

For recommendations, the authors used a source experiment in which ZOZOTOWN recorded 1,374,327 visitors across 80 fashion-item arms. The simulations varied the number of available arms between 2, 4 and 8, and used horizons of 100,000 and 1,000,000 visitors.

With a constant strategy, performance first improved as the privacy budget γ increased, then flattened and could decline. Dynamic performance rose toward an optimal-performance plateau. The number of arms also changed the exploration trade-off: at γ = 3, exploration probability rose from 10% with 2 arms to 30% with 8 arms; with 4 or 8 arms, constant-strategy clicks generally rose with γ.

A firm-wide budget

At the firm level, the authors used an input dataset of 78 ASOS online RCTs. Those experiments averaged 21 million visitors each, ranged from 69,000 to 149,197,471 visitors, and lasted an average of 43.5 days, with a maximum of 131 days.

Under the same firm-wide privacy budget, regret-based allocation often produced more portfolio clicks than splitting the budget evenly. The advantage was greatest at intermediate budgets, including Γ = 5 and Γ = 20; at very large budgets, both approaches moved toward the upper-bound benchmark.

A narrower promise than privacy in general

The reported evidence comes from analytical bounds and simulations rather than a live privacy intervention. Its formal guarantee bounds belief updating from the displayed experimental output, but does not protect the subsequent interaction or click channel.

The authors present the framework as a governance tool for managing impression privacy at customer, experiment and firm levels. The paper is an arXiv version-1 preprint dated 20 Aug 2026.

Paper data and sources

Original title: A Privacy Budgeting Framework for Online Experimentation
Authors: Gilian R. Ponte, Alina Ferecatu
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.