Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Comparing SMC and Persistent Sampling

This notebook extends the Use Tempered SMC to Improve Exploration of MCMC Methods and Tuning inner kernel parameters of SMC, exploring the theory and implementation of Persistent Sampling (PS) as described in Karamanis et al. (2025), and comparing it with standard tempered Sequential Monte Carlo (SMC) methods. We will compare four different methods:

Introduction

Sequential Monte Carlo (SMC)

SMC samplers propagate N particles through a sequence of probability distributions pt(θ)p_t(\theta) for t=1,…,Tt = 1, \ldots, T, using three main steps:

  1. Reweighting: Adjust particle weights using importance sampling

  2. Resampling: Discard low-weight particles and replicate high-weight ones

  3. Moving: Apply MCMC steps to diversify particles

For Bayesian inference with temperature annealing:

pt(θ)=L(θ)βtπ(θ)Ztp_t(\theta) = \frac{\mathcal{L}(\theta)^{\beta_t} \pi(\theta)}{Z_t}

where 0=β1<⋯<βT=10 = \beta_1 < \cdots < \beta_T = 1 interpolates between prior π(θ)\pi(\theta) and posterior.

Persistent Sampling (PS)

PS extends SMC by retaining and reusing particles from all prior iterations, constructing a growing weighted ensemble. Key differences:

Mixture Distribution: At iteration tt, particles from previous iterations are treated as samples from:

p~t(θ)=1t−1∑s=1t−1ps(θ)\tilde{p}_t(\theta) = \frac{1}{t-1} \sum_{s=1}^{t-1} p_s(\theta)

Persistent Weights: Using multiple importance sampling, weights for particle θt′i\theta^i_{t'} at iteration tt are:

Wtt′i=L(θt′i)βt1t−1∑s=1t−1L(θt′i)βs/Z^s⋅1Z^tW^i_{tt'} = \frac{\mathcal{L}(\theta^i_{t'})^{\beta_t}}{\frac{1}{t-1}\sum_{s=1}^{t-1} \mathcal{L}(\theta^i_{t'})^{\beta_s}/\hat{Z}_s} \cdot \frac{1}{\hat{Z}_t}

Resampling: N particles are resampled from (t−1)×N(t-1) \times N persistent particles

Key Advantages:

Trade-offs:

Imports and Settings

Experimental Setup

We will use the same setup as in Use Tempered SMC to Improve Exploration of MCMC Methods. We have seen that SMC can efficiently sample from a multimodal distribution.

SMC with Fixed Schedule

Now we’ll run standard SMC with a fixed tempering schedule. The tempering schedule and inference loop can be reused for PS.

Final ESS (SMC): 9991.06

CPU times: user 1min 41s, sys: 231 ms, total: 1min 41s
Wall time: 1min 3s

Persistent Sampling with Fixed Schedule

Now we run the same loop for the persistent sampling. Note that there are a few conditions and caveats for Persistent Sampling to function correctly:

Finally, the effective sample size of the final persistent ensemble can (and ideally will) be much larger than the number of particles per iteration.

Final ESS (Persistent Sampling): 237061.25 

CPU times: user 1min 3s, sys: 282 ms, total: 1min 4s
Wall time: 31.5 s

Adaptive SMC

Now we run the adaptive algorithms where the tempering schedule is chosen automatically. For the adaptive algorithm inference loop, we use a while loop that terminates when a tempering paramter of 1 is reached, or a predefined number of iterations is exceeded.

Number of iterations (Adaptive SMC): 9
Final ESS (Adaptive SMC): 9811.12 

CPU times: user 8.63 s, sys: 103 ms, total: 8.73 s
Wall time: 4.12 s

Adaptive Persistent Sampling

The adaptive Persistent Sampling algorithm works similar to the adaptive SMC algorithm. However, there are a few noteworthy percularities:

Number of iterations (Adaptive PS): 11
Final ESS (Adaptive PS): 70700.14 

CPU times: user 39.5 s, sys: 139 ms, total: 39.7 s
Wall time: 20.9 s

9. Posterior Comparison

Compare the posterior samples from all four algorithms.

<Figure size 2000x500 with 4 Axes>