A trade-off that unfolds over time
A preprint examines whether privacy can improve the long-term usefulness of a data-collection system. Its model asks when a finite privacy budget, the setting that controls the amount of privacy noise, can maximize utility in repeated mean estimation, where a system repeatedly estimates an average. The central idea is that leakage may be linked strongly enough to later nonparticipation that losing contributors matters alongside the noise added to each release. The authors use theory and simulations to compare private and non-private estimation under that feedback.
The first model is a Bernoulli warm-up. It estimates the underlying rate of a binary outcome from independent and identically distributed samples whose values are locally privatized with randomized response. The model also includes joining and leaving, allowing leakage in one round to affect the population available for later estimates. Privacy is therefore treated as an accuracy question and a population question at the same time.
Retention can change the best setting
In the Bernoulli simulations, the researchers tested 10 privacy-budget values, clipped the tracked population between 1 and 20,000, used a 15-step moving average for variance and averaged 30 independent trials. Two representative regimes paired joining and leaving probabilities at 0.6 and 0.3, then reversed those values. The reported 30 is a count of simulation trials, not a human-participant sample.
The results split according to the feedback regime. When leaving exceeded joining, the modeled population collapsed. When joining exceeded leaving, it stabilized or reached the maximum size. In regimes where a finite optimum existed, the best privacy budget balanced estimation noise against user retention, and the preferred privacy level became stronger as the leaving rate increased.
The binary analysis also gives a survival boundary. The permitted randomized-response perturbation depends on how much joining exceeds leaving and on the combined rate. If leaving is at least as common as joining, the critical allowance is clipped to zero. This is an expected-growth condition in a stylized model, not a measured threshold for real users.
A second model tracks leakage more closely
The extension uses a Gaussian mechanism and studies membership-inference leakage through a log-likelihood ratio, a score comparing the likelihood of an output under datasets with and without a particular record. For arbitrary data distributions, exact leakage probabilities are generally intractable, so the analysis uses a Rényi-divergence upper bound, a mathematical ceiling on the leakage probability, obtained through Markov's inequality. The bound offers a general leakage-control method, although it can be conservative.
For the Gaussian mechanism, once the dataset is fixed, that log-likelihood ratio follows an exact normal distribution. This supports normal-tail calculations for the chance that the score crosses the membership threshold. The exact result depends on the dataset and mechanism assumptions, so it describes that setup rather than every possible privacy system.
The population update makes the feedback explicit: expected next-round population equals the current population plus recruitment, minus losses weighted by the departure and recruitment rates and the previous round's marginal leakage probability. A sufficient lower noise threshold guarantees a growth factor of at least one, meaning the expected population does not decline in that round. The threshold depends on clipping, behavioral parameters and the data distribution.
The long-run picture has two extremes
The finite-sample error bounds put mechanism variance in the numerator and expected population size in the denominator. In practical terms, the modeled utility calculation weighs the randomness added for privacy against how many contributors remain available for estimation. The bounds rely on the study's clipping, population and approximation assumptions.
At the zero-noise edge, the model's expected-population approximation links vanishing algorithmic noise to expected error that diverges over time whenever the purge rate is positive. With positive recruitment and a per-round growth factor above one, expected error instead converges geometrically toward zero. In the simplified infinite-horizon model, the fastest rate occurs as the noise scale tends to infinity.
For a finite horizon, the reported theoretical noise-scale optima closely aligned with parameters found by grid search. That result suggests the proposed upper-bound objective can help select a setting within the modeled problem. It does not produce one universal privacy budget: the Bernoulli simulations found finite optima only in some feedback regimes, and the authors' broader conclusion is conditional on sufficiently strong leakage-participation feedback.
Evidence for a model, not a verdict on users
The authors interpret the combined theory and simulations as evidence that private mean estimation can outperform non-private estimation when leakage-participation feedback is sufficiently strong. Because the evidence comes from stylized models and simulations, the conclusion is conditional on the specified mean-estimation setting, population update and distributional assumptions.
The general leakage bound may be conservative, while the population-survival threshold is presented as sufficient. Those caveats matter because the paper's central trade-off is produced by the model's rules for leakage, recruitment and departure, not by a reported participant experiment.
The acknowledgments cite support from the National Research Agency under France 2030, a visit to the Simons Institute and the Verified Deep Learning: Formal Methods Perspective project.
Paper data and sources
Original title: Performative Privacy: When Differential Privacy Maximizes Utility
Authors: Uddalak Mukherjee, Edwige Cyffers, Yann Chevaleyre
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-28
DOI: Not available
Original paper · Full text