arl-gsm8k-drgrpo-8seed
8 independent Dr.GRPO runs (loss_type=dr_grpo, scale_rewards=False) on GSM8K from
Qwen/Qwen3-0.6B, 150 steps each, seeds 0-7. Layout: seed{N}/checkpoint-150/.
Full-1319 GSM8K greedy@1: per-seed avg 0.6775, best 0.7127, union 0.9166 (lottery gap 0.239). A uniform weight-soup of all 8 scores 0.7316 (beats the best single seed by +1.9pt, free). Training 8x longer (1200 steps) degrades to 0.594 — souping beats more compute. See companion notebooks for the full study.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support