Main Results
RCD consistently beats Sequential Denoising (SeqD) across every model / block-size setting, with the largest gains on competition-level AIME24/25, where accuracy often more than doubles. Confidence threshold 0.85; SeqD/RCD use a 16,384 sequence length, Chat uses 512 tokens (1,024 for AIME).
| Model | Variant | GSM8K† | MATH500 | AIME24 | AIME25 |
|---|---|---|---|---|---|
| SDAR-4B-b32 | Chat‡ | 86.13 | 50.20 | 5.83 | 2.50 |
| SeqD | 81.73 | 61.20 | 6.04 | 11.88 | |
| RCD | 84.91 | 65.40 | 9.17 | 17.08 | |
| SDAR-4B-b64 | Chat‡ | 85.90 | 49.80 | 6.25 | 1.67 |
| SeqD | 78.85 | 56.80 | 4.17 | 7.29 | |
| RCD | 87.04 | 68.40 | 11.04 | 14.79 | |
| SDAR-8B-b32 | Chat‡ | 88.40 | 50.00 | 6.46 | 4.17 |
| SeqD | 86.50 | 65.80 | 11.67 | 14.79 | |
| RCD | 90.45 | 76.20 | 18.96 | 20.00 | |
| SDAR-8B-b64 | Chat‡ | 88.32 | 51.60 | 5.20 | 2.50 |
| SeqD | 82.87 | 64.20 | 7.08 | 9.79 | |
| RCD | 86.05 | 74.40 | 18.75 | 16.04 |
† Potential data contamination was observed in the original Chat models on GSM8K, which may inflate the Chat baseline. ‡ Chat variants are instruction-following models; SeqD/RCD are further adapted for mathematical reasoning.
| Model | Variant | GSM8K | MinervaMath |
|---|---|---|---|
| LLaDA | Base | 70.30 | 31.40 |
| SeqD | 75.74 | 31.10 | |
| RCD | 78.09 | 37.00 |
LLaDA (global-attention dLLM): sequence length 512, single-token-per-step decoding.
- Doubles hard-task accuracy: on SDAR-8B-b64 AIME24, RCD lifts accuracy from 7.08 → 18.75.
- Scales with size & block: gains of 4–11 percentage points when scaling from 4B to 8B, widening as block size grows.
- Generalizes across architectures: on global-attention LLaDA, RCD adds a ~6 point absolute gain on MinervaMath.