Preprint

Preprint tests method for separating exposure effects from time trends

Simulations found the design stayed unbiased under one pattern of time-induced confounding, but the authors caution that it is not a general solution.

A statistical method aimed at separating short-lived exposure effects from changes over time stayed unbiased in a difficult computer simulation, while several standard self-controlled approaches did not, according to a new preprint. The method, called symmetric pair matching or SPM, was reported as unbiased in that test, while standard SCCS, CCO and an SCCS model using a five-knot natural spline were biased. Its efficiency was comparable to an SCCS model using 15 knots, one of the methods that also adjusted for the simulated time pattern. The result comes from generated data and speaks to the tool's simulated performance.

The paper asks whether SPM can combine control of time-invariant confounding, or persistent differences that do not change over time, with adjustment for changing exposure and outcome patterns without explicitly modeling those time effects. That question is important because the simulation shows that some self-controlled approaches can become biased when exposure and outcome patterns move out of phase. The authors present SPM as a design-based option, while cautioning that it is not a general solution for inference.

Pairs that work in both directions

SPM constructs reciprocal matched pairs from exposed individuals. A pair is kept only when the two post-exposure risk periods—the windows being compared after exposure—do not overlap and both periods are fully observed. In practical terms, exposed people serve as controls for one another, and the same comparison can be read in both directions.

The point estimate is built from event counts. It averages products of individual-level event counts to estimate the exposure effect. Uncertainty can be measured by bootstrapping at the individual level, meaning that people are resampled to produce standard errors and confidence intervals.

The appendix gives a consistency theorem that allows one person to contribute to more than one valid pair, even when those pairs are dependent. The theorem is asymptotic: it describes the estimator's behavior as information grows, but it does not guarantee unbiased estimates in every finite sample.

A deliberately difficult test

The main Monte Carlo experiment generated 1,000 datasets for each of two scenarios, with 20,000 simulated individuals in every dataset. The synthetic data included a binary exposure, recurrent outcomes and a standard-normal time-fixed confounder. Cox hazard models generated the exposure and outcome processes. One scenario used constant baseline hazards; the other set the exposure and outcome patterns out of phase, creating a direct test of the method's handling of temporal effects.

Each dataset was analyzed with standard SCCS, SCCS using a natural spline with five knots, SCCS using 15 knots, standard CCO, CTC and SPMD using all valid pairs. The researchers compared each estimate with the known effect and examined bias, variance and mean squared error. Bias is the tendency to miss the known effect; variance is how much estimates move around from dataset to dataset.

When the baseline hazards were constant, all methods were unbiased on average. SPMD and the SCCS variants had comparable variance, while CCO and CTC had greater spread across the simulated datasets. In that simpler setting, SPMD landed in the same broad precision range as the SCCS variants.

The harder scenario produced a clearer separation. Standard SCCS, CCO and the five-knot spline SCCS were biased when the exposure and outcome patterns were out of phase. The 15-knot spline SCCS, CTC and SPMD adjusted for the temporal effects, and SPMD's efficiency was comparable to that of the correctly specified spline SCCS. The authors frame this as a result under the simulated pattern, not a guarantee across all possible time trends.

Useful, but tightly bounded

That result depends on a demanding set of assumptions. SPM requires a common transient effect and a known common risk-period length, recurrent outcomes, exposure independent of previous events, a time-invariant exposure effect, a common outcome trend over time and censoring not driven by shared time-dependent factors. These requirements define a narrower use case than a general-purpose method.

The authors specifically caution that SPM is not a general solution for inference and is not appropriate for long-term or sustained treatment effects. It is aimed at self-controlled analyses of transient exposures where temporal effects matter and the stated conditions can be defended.

The evidence here is limited to mathematical derivations and synthetic finite-sample simulations. It does not establish how SPM performs in real observational data, and it does not show that any drug, vaccine or other exposure causes a clinical outcome. Those questions remain outside what this preprint demonstrates.

Appendix F repeated the simulations across different effect values and sample sizes and reported bias, mean squared error, Monte Carlo standard errors and missing or infinite estimates. The supplementary settings were not combined into one pooled performance estimate.

The authors provide the SPMD R package as a freely available GitHub implementation, allowing other analysts to test the design. The front matter identifies the work as arXiv:2608.25979v1, dated 26 Aug 2026. Its financial disclosure is recorded as none reported, and the authors report no potential conflicts of interest.

Paper data and sources

Original title: The Symmetric Pair Matching Design: A Self-Controlled Method with Automatic Adjustment for Time Effects
Authors: Robin Denz, Filippo Saatkamp, Katharina Meiszl, Nina Timmesfeld
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.