Research · methods and evidence

Structure first. Claims second.

My work starts from sparse recovery and keeps returning to the same practical question: what structure can we exploit without pretending the evidence says more than it does?

Active research

Current threads

Evidence labels describe today’s boundary, not the ambition.

Verifier-free reasoning RLreplicated policy-learning result

Pursuit-Graded Reward (PGR/FCE)

A dense GRPO reward built from the structure of sampled reasoning trajectories rather than a gold answer checker. FCE uses within-group trajectory evidence to assign step-level learning signal, then tests whether that signal survives permuted-reward, base-model, majority, gold-reward, and self-certainty controls.

Evidence boundary

Across SmolLM training seeds 43-46, FCE beat the permuted control by 20.8 percentage points (95% CI 16.4-24.7), with every seed positive. On Qwen2.5-1.5B, FCE reached 66.8% on held-out GSM8K versus 54.6% for the base model and 49.2% for the permuted control. It did not significantly beat ordinary majority or gold GRPO; MATH-Hard and MMLU policy-transfer gates did not pass; support-only rewards cannot identify fully unanimous wrong consensus.

Transformer optimizationactive research

SemanticDose and Halley-Gram-Muon

Architecture-aware Muon variants that change update geometry only where the parameter semantics support it. SemanticDose retires by cumulative displacement; Halley replaces the polar map with a cheaper rational iteration.

Evidence boundary

Public implementation and frozen experiment artifacts. Attention generalization remains gated on the new MHA/GQA/MQA queue.

Training-free decompositionpublic artifact

SVD-OMP

An orthogonal SVD dictionary with input-specific OMP support selection for decomposing Transformer MLP parameters without fitting a learned dictionary.

Evidence boundary

The public evaluation reports reconstruction, faithfulness, coherence, reproducibility, and support-stability comparisons on Goodfire 67M matrices. Causal interpretation remains future work.

AI-control evaluationprovisional evidence

SecretLoyaltyBench

A behavioral audit for principal loyalty and hidden-objective failures, with explicit separation between evaluated organisms and untested transfer claims.

Evidence boundary

The public artifact preserves held-out behavioral results and known limitations. Automatic-judge calibration and independent audit are still pending.

Selected publications

Published work

Operating rule

Negative results stay in the record.

Tuning is not confirmation. A local benchmark is not an external validation. A useful model-internal signal is not automatically a causal explanation. Those distinctions make the successful results more useful, not less.