The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models
Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 80.0%
Avg response (correct): 327.8 words
Avg response (incorrect): 248 words
Length gap: -79.8 words (incorrect longer = positive)
Points estimate: 5 toy-scale (1 pt each) + 0 verified (2 pts each)
| Claim | Verdict | Evidence |
|---|---|---|
| Arbitrary-order decoding has flatter Pass@ scaling than autoregressive order on reasoning benchmarks, indicating lower reachable reasoning potential under practical sampling (Figure 3) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that Arbitrary-order decoding has flatter Pass@ scaling than autoregressive order on ..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale. |
| Problems solved by arbitrary-order decoding are largely a subset of those solved by autoregressive-order decoding in the Pass@ solution-coverage analysis (Figure 4) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 80.0% accuracy. Correct responses averaged 327.8 words vs 248 for incorrect responses. The claim that Problems solved by arbitrary-order decoding are largely a subset of those solved... is directionally consistent with our results at toy scale. |
| The paper identifies entropy degradation at logical-fork tokens as a mechanism by which arbitrary-order decoding bypasses hard decisions and narrows exploration (Figure 7) | TOY | Correct responses averaged 327.8 words vs 248 words for incorrect responses (gap = 79.8 words). This is negatively consistent with the claim that The paper identifies entropy degradation at logical-fork tokens as a mechanism b.... Tested on 5 MATH-500 problems; toy-scale verdict. |
| JustGRPO reaches 89.1% GSM8K accuracy with standard GRPO on LLaDA-Instruct while avoiding diffusion-specific RL adaptations (Table 1) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that JustGRPO reaches 89.1% GSM8K accuracy with standard GRPO on LLaDA-Instruct while..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale. |
| JustGRPO-trained models remain compatible with parallel decoding, with larger accuracy gains at higher parallel token counts than the original instruct model (Figure 8) | TOY | Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that JustGRPO-trained models remain compatible with parallel decoding, with larger ac..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale. |
Authored by Jude Ighomena, Copyright Janna AI Research Labs