A result from selected hard cases
DCGC with LLaDASFT, a masked-diffusion framework, reported an average accuracy of 24.8 across six benchmarks and the best score on five of them. Its reported scores were 44.9 on GSM8K, 22.3 on MATH, 10.7 on MBPP, 13.1 on HumanEval, 35.7 on MMLU-STEM and 22.5 on MMLU-Pro.
The framework uses an imperfect upstream solution as auxiliary context for a global correction of a complex reasoning trace. The study asks whether masked diffusion models can use such drafts for verifier-free correction, meaning correction without a separate check of logical correctness.
The benchmark slice matters
The main evaluation used solver-failure hard sets, selected from problems the initial Llama-3.1-8B-Instruct solver had failed to solve correctly. The sets contained 216 GSM8K samples, 282 MATH, 216 MBPP, 69 HumanEval, 956 MMLU-STEM and 6,675 MMLU-Pro samples. The main comparison was made within those selected failures.
Two inputs, two contexts
DCGC combines mixed-format supervised fine-tuning, or SFT, with Dynamic Dual-CFG. Some training examples pair a problem with a solution, while others use draft-conditioned correction triples. The SFT data targeted approximately 10,000 unique problems per domain and contained 55,440 training samples and 2,913 validation samples, with 5% reserved for validation.
Dynamic Dual-CFG keeps two contexts in view: the problem alone, and the problem together with the draft. It scales the additional draft residual using the relative confidence gap between those contexts. In plain terms, the method adjusts the draft contribution according to the difference in confidence between the two versions. The paper cautions that this confidence gap is not a logical-correctness verifier.
Different comparisons, different averages
With standard sampling, LLaDASFT reported 18.9 average accuracy versus 4.7 without SFT. On GSM8K, the figures were 32.4% and 7.4%; LLaDASFT also reported 30.9% on MMLU-STEM and 19.2% on MMLU-Pro.
Among the listed guidance alternatives, DCGC Relative reported 24.8 average accuracy, compared with 20.2 for Dual-CFG Static, 21.5 for Dual-CFG Independent, 22.4 for Single-CFG Problem and 18.9 for Standard Sampling. DCGC Relative had the highest average among those alternatives.
Results varied by draft condition
The paper reports sensitivity to draft relevance on MATH. Accuracy was 22.3 with the original draft, 10.3 with a shuffled draft and 11.0 with a domain-shifted draft.
When correction was called
In a gold-agnostic protocol, the system used five upstream samples for self-consistency. A query was sent to correction when no answer received at least three votes.
With that protocol, DCGC reported full-set accuracy of 43.20 on MATH, 86.96 on GSM8K and 69.17 on MMLU-STEM. On the cases sent for correction, the corresponding scores were 23.5, 39.8 and 33.8. Compared with Majority Voting over five samples, the full-set accuracy-point differences were +2.01 on MATH, +0.16 on GSM8K and +0.51 on MMLU-STEM.
A post-hoc split
A post-hoc easy-versus-hard analysis reported DCGC hard-subset accuracy of 31.51 on GSM8K, 18.11 on MATH and 31.45 on MMLU-STEM. The easy-subset figures were 53.33, 44.62 and 40.35, respectively.
Supporting evidence from another backbone
An additional experiment with the DREAM backbone reported average accuracy of 4.7 with unsupervised standard sampling, 15.0 with Dynamic Dual-CFG before SFT, 17.4 with SFT standard sampling and 26.4 with DCGC. DCGC's reported scores were 42.6 on GSM8K, 18.1 on MATH, 13.4 on MBPP and 31.4 on MMLU-STEM. The manuscript describes this experiment as supporting evidence.
On MATH, disagreement positions had a mean confidence gap of 0.214, compared with 0.026 at agreement positions. Draft reuse was 0.321 at disagreement positions and 0.192 at agreement positions. The example-level original-draft Spearman correlation was 0.195 and remained 0.247 after controlling for generation length. The paper says the confidence gap is not a logical-correctness verifier.
Paper data and sources
Original title: DCGC: Draft-Conditioned Global Correction for Complex Reasoning with Masked Diffusion Models
Authors: Minhae Oh, Nakyung Lee, Jungwoo Lee
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text