The proposed decoder recasts degenerate maximum-likelihood decoding as a comparison of partition functions. Each logical error class is represented by the partition function of an unconstrained positive Markov random field over generator variables. In simpler terms, the calculation assigns a weighted total to each class and compares those totals. The sampling branch uses annealed importance sampling to estimate them, reusing the same random numbers across classes to sharpen paired comparisons. It then applies paired-bootstrap certificates, statistical checks designed to show when one class is ahead. The paper also says that combining this approach with constant-factor estimators such as WISH can support an exact proof of optimality.
Strong performance on small reference tests
The study reports four computational experiments run on a 64-core commodity server. The first three used independent bit-flip noise for code-capacity tests, while the fourth added phenomenological and circuit-level syndrome noise. For estimator validation, the researchers tested 200 syndromes at p = 0.05 and another 200 at p = 0.10. As annealed-importance-sampling sweeps increased from 4 to 64, mean absolute log-partition error fell from 0.82 nats to 0.068 nats. The 95th-percentile error stayed below 0.18 nats, and reusing common random numbers reduced ratio-error variance by a factor of 1.2 to 2.2.
At a certificate significance level of 0.05 and 32 sampling sweeps, both certificates fired on all 400 validation decisions, with zero empirical errors in that set. The Bethe branch, which uses a region-based approximation to compare classes, produced class log-ratios that differed from exact values by an average of 0.10 nats. Its decisions also matched exact maximum-likelihood decoding on all 400 syndromes.
The closest match came on rotated surface codes. The sampling decoder and exact degenerate maximum-likelihood decoder agreed on every tested distance-3 trial and through p = 0.06 for distance 5. At p values of 0.08, 0.10 and 0.12, they differed on no more than 0.3 percent of trials. Certification was 100 percent at p values up to 0.06 and 98 percent at p = 0.12.
Bicycle-code results were more conditional
In the reported code-capacity comparisons for bivariate-bicycle, or BB, codes, AIS was never worse than the BP+OSD-0 baseline. On the [[72, 12, 6]] code, it repaired eight decisions that the baseline got wrong and broke three decisions that the baseline got right. It also reached the performance of BP+OSD-CS-7 at error rates of p = 0.02 and p = 0.05. Many of these comparisons, however, used a candidate set centered on BP+OSD results rather than searching every logical class.
A direct comparison with exact maximum-likelihood decoding showed why the reference is costly. Exact decoding took between 5 and 26 seconds per syndrome. At p = 0.03, its error rate was 0.0417, compared with 0.0500 for the BP+OSD variants, which were suboptimal on two of 120 syndromes. AIS agreed with exact decoding on 238 of 240 syndromes and recorded no true violations among the 231 decisions that received a bootstrap certificate.
The certificate method was also more reliable than a paired t-test in the BB experiments. Paired-bootstrap certification ranged from 94 percent to 100 percent across the tested codes and rates. The paired t-test certified 58 percent to 92 percent of decisions on [[72, 12, 6]], but its certification rate fell to essentially zero on the gross code, the larger-width instance.
The region hierarchy has a sharp boundary
The region-based branch uses Bethe free energy as a class-comparison score and expands the calculation through the Kikuchi hierarchy and cluster elimination. On the smaller BB instance, mini-bucket elimination with regions of size 16 and 20 reproduced exact decisions on all 40 syndromes. It issued deterministic certificates for 37 and 38 of them, respectively. On the gross code, however, MBE(20) broke 49 of 60 correct decisions, issued no deterministic certificates and took 272 seconds per decode. The analysis attributes that failure to upper bounds that were too loose to rank the classes reliably.
Results under noisy-syndrome simulations were stronger. With surface-code circuit noise, AIS was at least as accurate as both tested baselines and certified 99 percent to 100 percent of decisions at 0.8 seconds per decode. At p = 0.004, its block-error rate was 0.0133 versus 0.0183 for BP+OSD-0; the results were tied at p = 0.008 and statistically indistinguishable at p = 0.012. For phenomenological BB noise, AIS agreed with spacetime BP+OSD-0 on all 720 trials and certified every decision within the single-logical candidate set.
Hardware tests exposed a modelling problem
The surface-code hardware pilot was less decisive. The measured detector rate was 0.229, compared with a calibration prediction of 0.076. The raw logical flip rate was 0.253, the best decoder reached 0.246, and AIS reached 0.260 plus or minus 0.025 on a 300-shot subsample. AIS certified 96 percent of its decisions, but its reported performance was statistically indistinguishable from the alternatives in that subsample.
The BB hardware pilot produced a lower reported block-error rate after decoding, 0.709 versus 0.983 before decoding. Its mean detector rate was 0.355 against a calibration value of 0.238. The raw per-logical flip rate was 0.439, compared with a reported decoded per-logical error of 0.349. AIS certified 100 percent of decisions and agreed with spacetime BP+OSD on all 200 decoded syndromes. The experiment was deep and near saturation, and the analysis notes a mismatch in the calibration model.
A useful framework, not a finished decoder
The evidence is finite and code-specific. The strongest exact comparisons cover selected surface-code distances and BB instances, while several BB tests rely on restricted candidate sets. The gross-code region method failed at the tested width, and both hardware pilots showed mismatches between measured and calibrated detector rates. Those results leave larger instances, fuller logical-class searches, tighter bounded-region estimates and richer hardware noise models unresolved.
The work is identified as arXiv:2608.25545v1, dated 26 August 2026. The authors state that code, scripts and raw results are available upon request. The research was funded by Germany's Federal Ministry of Research, Technology and Space and the state of North Rhine-Westphalia through the Lamarr Institute for Machine Learning and Artificial Intelligence.
Paper data and sources
Original title: Certified decoding of quantum LDPC codes
Authors: Ragavi Krishnamoorthy, Florian Gerhardt, Johannes Knaute et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text