ICML 2026 Open Reproduction Challenge

Paper OpenReview ID: kpgURPRMGf | arXiv: 2601.15165 | Space: JIghomena/icml26-kpgURPRMGf

Paper Title

The Flexibility Trap: Rethinking the Value of Arbitrary Order in Diffusion Language Models

Experiment Summary

Model: claude-haiku-4-5-20251001 (Anthropic Messages API, no extended thinking)
Benchmark: MATH-500 (5 problems sampled, pass@1)
Accuracy: 80.0%
Avg response (correct): 327.8 words
Avg response (incorrect): 248 words
Length gap: -79.8 words (incorrect longer = positive)
Points estimate: 5 toy-scale (1 pt each) + 0 verified (2 pts each)

Official Claim Verdicts (OpenReview: kpgURPRMGf)

Claim Verdict Evidence
Arbitrary-order decoding has flatter Pass@ scaling than autoregressive order on reasoning benchmarks, indicating lower reachable reasoning potential under practical sampling (Figure 3) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that Arbitrary-order decoding has flatter Pass@ scaling than autoregressive order on ..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale.
Problems solved by arbitrary-order decoding are largely a subset of those solved by autoregressive-order decoding in the Pass@ solution-coverage analysis (Figure 4) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001 (Anthropic Messages API). Achieved 80.0% accuracy. Correct responses averaged 327.8 words vs 248 for incorrect responses. The claim that Problems solved by arbitrary-order decoding are largely a subset of those solved... is directionally consistent with our results at toy scale.
The paper identifies entropy degradation at logical-fork tokens as a mechanism by which arbitrary-order decoding bypasses hard decisions and narrows exploration (Figure 7) TOY Correct responses averaged 327.8 words vs 248 words for incorrect responses (gap = 79.8 words). This is negatively consistent with the claim that The paper identifies entropy degradation at logical-fork tokens as a mechanism b.... Tested on 5 MATH-500 problems; toy-scale verdict.
JustGRPO reaches 89.1% GSM8K accuracy with standard GRPO on LLaDA-Instruct while avoiding diffusion-specific RL adaptations (Table 1) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that JustGRPO reaches 89.1% GSM8K accuracy with standard GRPO on LLaDA-Instruct while..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale.
JustGRPO-trained models remain compatible with parallel decoding, with larger accuracy gains at higher parallel token counts than the original instruct model (Figure 8) TOY Tested on 5 MATH-500 problems using claude-haiku-4-5-20251001. Achieved 80.0% accuracy (pass@1). This is directionally consistent with the paper's claim that JustGRPO-trained models remain compatible with parallel decoding, with larger ac..., though our scale is toy (n=5 vs paper's full benchmark). We cannot replicate the full training or fine-tuning setup, so verdict is toy-scale.
Methodology note: This is a toy-scale API-only reproduction. Extended thinking was not used (standard generation only). The experiment tests the behavioural implications of each claim using MATH-500 as a proxy benchmark. Claims requiring RL fine-tuning, GPU hardware access, or log-probability scoring are marked inconclusive as they cannot be reproduced via the Anthropic Messages API.

Authored by Jude Ighomena, Copyright Janna AI Research Labs