Preprint

Bayesian sampler tests report gains in one model, mixed results in another

Preprint: An arXiv study reports higher RI-CLPM efficiency, mixed OMRF results and an unresolved gap in posterior agreement.

Two continuous-time Bayesian samplers recorded the highest reported efficiency in one set of tests of statistical model selection, while a second set produced a more mixed picture. In RI-CLPM simulations, Bouncy Particle, or BPS, reached a median effective sample size rate of 38.1 per second and Boomerang 35.7, compared with 8.66 for Zig-Zag and 3.02 for the NIMBLE reference implementation. In OMRF simulations, the reported subsampled-to-full efficiency ratios were above one for BPS and Zig-Zag but not for Boomerang, and posterior inclusion probabilities did not fully line up with the reference results.

The work examines Bayesian spike-and-slab variable selection, a way of estimating which effects should be included in a model while allowing others to be left out. It evaluates piecewise deterministic Markov process samplers, or PDMPs, as a continuous-time alternative to conventional Markov chain Monte Carlo, usually shortened to MCMC. The framework extends sticky PDMP methods to dependent priors and to unbiased stochastic gradients for factorized likelihoods, while targeting the same distribution.

The tests covered two psychometric model families. RI-CLPM simulations used four variables at four time points, sample sizes of 100 or 300, and 12 selectable cross-lagged effects, with 10 replications for each condition. OMRF simulations used four response categories, sample sizes of 200 or 500, and 10, 20 or 30 nodes, again with 10 replications per condition. All 480 planned RI-CLPM fits and all 1,520 planned OMRF fits were completed, and none were excluded.

The clearest result came from RI-CLPM

To measure efficiency, the paper used median effective sample size per second of retained sampling time. Effective sample size accounts for the fact that repeated draws from a sampler may carry overlapping information, so a higher rate indicates more statistically useful output per second. BPS and Boomerang led the RI-CLPM results at 38.1 and 35.7, while Zig-Zag recorded 8.66 and NIMBLE 3.02. NIMBLE's rate fell from 7.89 at N = 100 to 1.58 at N = 300, whereas PDMP efficiency changed little.

The selection results also improved with more data in the RI-CLPM simulations. Across samplers and non-null conditions, median ROC AUC, a measure of how well scores separate effects that belong in the model from those that do not, was 0.85 at N = 100 and 1.00 at N = 300. Using a posterior inclusion threshold of 0.50, sensitivity rose from 0.33 to 1.00, while specificity was 1.00 at both sample sizes. Of 1,920 decisions in null conditions, nine were false positives. Median RMSE, the reported measure of estimation error, fell from approximately 0.049 to 0.022.

The choice of prior changed that balance. Under the dependent prior, the median posterior model size in null conditions was 0.38 versus 0.93 at N = 100, and 0.22 versus 0.64 at N = 300. The same prior had lower balanced accuracy in common-sender and mixed-sign conditions at N = 100, although the differences were smaller at N = 300. The reported pattern was therefore a trade-off: sparser null models came with weaker recovery on some non-null conditions.

Despite the efficiency gap, RI-CLPM posterior summaries were close to those from condition-matched NIMBLE runs. The median mean absolute difference in inclusion probabilities was 0.019 for adaptive Boomerang and 0.028 for both BPS and Zig-Zag. The corresponding differences for posterior means were 0.0023, 0.0033 and 0.0030. A mean absolute difference summarizes the average size of the disagreement without treating positive and negative gaps as canceling each other.

Subsampling did not help every sampler

OMRF results depended more sharply on the sampler. The study compared full gradients with subsampled gradients and expressed the result as a subsampled-to-full effective sample size rate ratio. For BPS, the ratios were 2.90, 2.52 and 2.34 at 10, 20 and 30 nodes. For Zig-Zag, they were 4.35, 2.93 and 2.10. Boomerang showed 1.00, 0.95 and 0.54, indicating no gain at 10 nodes and lower reported efficiency at 20 and 30.

A Boomerang example shows why output and runtime cannot be read separately. At 30 nodes and N = 500, the subsampled run produced 158 effective samples compared with 106, but sampling time increased from 42 seconds to 208 seconds. The larger effective sample count came with a much longer run in that comparison.

In structure recovery at 30 nodes under the independent slab, Boomerang had a median balanced accuracy of 0.69, compared with approximately 0.50 for BPS and Zig-Zag. The reported relative model-size errors were approximately 0.3, 3.7 and 3.4. Within each dynamic, recovery was similar for full and subsampled gradients.

That OMRF analysis also exposed the main unresolved issue. Median mean absolute differences in posterior inclusion probabilities between full and subsampled runs were 0.108 for Zig-Zag, 0.167 for BPS and 0.120 for Boomerang. At 10 nodes, the differences from NIMBLE were 0.351, 0.424 and 0.140 in the same order. The paper says recorded Monte Carlo uncertainty did not fully explain the discrepancies, so it remains unclear why the posterior summaries diverged.

A dense fitted network, with a warning

The method was also used for an empirical OMRF analysis of complete responses from 1,804 participants on 14 Mental Health Continuum-Short Form items. The analysis used six ordered response categories, 500 warmup units and 2,000 retained units. All 91 edge-inclusion probabilities exceeded 0.50, the expected model size was 85.3 edges, and the minimum and median effective sample sizes were 796 and 1,878. The reported fit therefore favored a dense network of item relationships.

Still, the speed comparison with NIMBLE needs qualification. The samplers ran in different software environments and used different computational representations, so the runtime differences are not pure algorithm comparisons. The preprint's results support a conditional reading: BPS and Boomerang looked strongest on RI-CLPM efficiency, while OMRF performance varied by dynamic and the posterior disagreement was not resolved.

Paper data and sources

Original title: Accelerating Bayesian Variable Selection using Piecewise Deterministic Markov Processes
Authors: Don van den Bergh, Maarten Marsman
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-27
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.