arl-gsm8k-drgrpo-8seed

8 independent Dr.GRPO runs (loss_type=dr_grpo, scale_rewards=False) on GSM8K from Qwen/Qwen3-0.6B, 150 steps each, seeds 0-7. Layout: seed{N}/checkpoint-150/.

Full-1319 GSM8K greedy@1: per-seed avg 0.6775, best 0.7127, union 0.9166 (lottery gap 0.239). A uniform weight-soup of all 8 scores 0.7316 (beats the best single seed by +1.9pt, free). Training 8x longer (1200 steps) degrades to 0.594 — souping beats more compute. See companion notebooks for the full study.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ksgk-fy/arl-gsm8k-drgrpo-8seed

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1336)
this model