Pursuit-Graded Reward: verifier-free policy learning that survived the controls
FCE beat a permuted-reward control by 20.8 points across four SmolLM seeds and raised Qwen2.5-1.5B from 54.6% to 66.8% on GSM8K. Majority, harder-domain transfer, and unanimous-wrong-consensus claims remain explicit non-wins.