Preprint

Covariance method matches GLMM influence checks, weaker N-mixture agreement

A preprint tests PosCoSeA, which estimates leave-one-out changes and bootstrap uncertainty without repeated model refitting.

A posterior-covariance method closely matched exact leave-one-out checks in simulations of generalized linear mixed models, but its agreement was weaker in N-mixture models, according to a new preprint. Leave-one-out analysis tests how much a result changes when one observation is removed; the method estimates that change without repeatedly refitting the model.

The reported computational comparison was striking in the stated setup: about 60 seconds for the posterior-covariance method, compared with about 50 hours for maximum-likelihood refitting and about 250 days for repeated MCMC refitting. Those figures are reported costs for this analysis, not a general measure of how long every model would take.

A different way to test sensitivity

PosCoSeA uses the posterior covariance—the way posterior quantities vary together—to approximate leave-k-out influence diagnostics and bootstrap resampling without repeated model refitting. For the bootstrap, resampled data sets are represented by multinomial observation counts, producing an approximate bootstrap distribution without additional MCMC runs.

The GLMM simulations used sample sizes of 50, 100 and 200 observations, with five random-effect groups. Each replicate drew 10,000 posterior samples and compared covariance-based changes with exact leave-one-out refitting across 100 simulation replicates.

The closest match came in the GLMM tests

The covariance-based influence scores tracked the exact changes closely. Mean Spearman rank correlations were 0.978 for 50 observations, 0.982 for 100 and 0.983 for 200, with reported standard errors of 0.002, 0.001 and 0.001 respectively. Spearman correlation here measures how similarly the two approaches rank observations by influence.

The same simulations compared several ways of measuring uncertainty. At sample sizes of 50, 100 and 200, the empirical standard errors were 0.285, 0.186 and 0.129. Mean posterior standard deviations were lower, at 0.133, 0.091 and 0.063, while bootstrap standard errors were 0.252, 0.178 and 0.126, closer to the empirical figures.

Coverage measures how often a stated 95% interval contains the reference value. In these GLMM simulations, Bayesian coverage was 0.646, 0.663 and 0.680 as the sample size rose from 50 to 200. Bootstrap coverage was 0.913, 0.933 and 0.946, and was reported as closer to the empirical uncertainty in these tests.

Performance was less even in N-mixture models

The N-mixture simulations used 50, 100 or 150 sites, with three surveys per site. They began with 10,000 posterior draws, discarded 5,000 as burn-in, and continued sampling when the convergence diagnostic R-hat exceeded 1.01 until it fell below that threshold.

The leave-one-out approximations still showed a linear relationship with exact refitting, but the agreement was weaker than in the GLMM simulations, particularly for the abundance-intercept and total-abundance targets. The paper identifies possible Monte Carlo error from MCMC sampling as an unresolved contributor to that weaker agreement.

Uncertainty coverage also varied substantially by scenario and target. With 50 sites under log-normal misspecification, posterior versus bootstrap 95% coverage was 0.90 versus 0.98 for the abundance-intercept target and 0.80 versus 0.95 for the βλ target. With 100 sites under beta-binomial misspecification, the corresponding figures were 0.21 versus 0.88 for the abundance-intercept target and 0.20 versus 1.00 for the total-abundance target.

A real-data example widened the uncertainty range

The researchers also reanalysed a Swiss tit field survey covering 263 sites. The analysis used 13 detection covariates and deliberately fitted a standard Poisson abundance model rather than the original zero-inflated Poisson model because the data showed some overdispersion.

The reported national population estimate was 738,000. Its Bayesian 95% interval ran from 563,000 to 792,000, while the cited zero-inflated-Poisson interval ran from 491,000 to 706,000. The approximate bootstrap interval was wider, running from 428,000 to 900,000.

Influence was concentrated in a small part of the survey. About 80% of the sites changed the posterior mean by less than 1% when removed, while five sites changed it by roughly 4% to 6%. The influential sites differed depending on which posterior target was examined.

What the evidence shows

The evidence comes from simulations and one ecological-data application: GLMM and N-mixture tests, followed by a Swiss tit reanalysis. The reported runtimes are specific to the setup, and the model comparisons did not produce one uniform level of agreement.

The paper provides an R package for the method, and the empirical data used in the real-data analysis are available from the R package AHMbook. The research was partially supported by JSPS KAKENHI grants 21K15170 and 24K15120, and the authors declared no conflicts of interest.

The document is an arXiv version 1 preprint dated 26 Aug 2026.

Paper data and sources

Original title: {poscosea} : A Computationally Efficient Sensitivity Analysis for Bayesian Models using the posterior covariance representation
Authors: Yusaku Ohkubo, Yukito Iba
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.