Pursuit-Graded Reward (PGR/FCE)
A dense GRPO reward built from the structure of sampled reasoning trajectories rather than a gold answer checker. FCE uses within-group trajectory evidence to assign step-level learning signal, then tests whether that signal survives permuted-reward, base-model, majority, gold-reward, and self-certainty controls.
Across SmolLM training seeds 43-46, FCE beat the permuted control by 20.8 percentage points (95% CI 16.4-24.7), with every seed positive. On Qwen2.5-1.5B, FCE reached 66.8% on held-out GSM8K versus 54.6% for the base model and 49.2% for the permuted control. It did not significantly beat ordinary majority or gold GRPO; MATH-Hard and MMLU policy-transfer gates did not pass; support-only rewards cannot identify fully unanimous wrong consensus.