Verified Research Reasoning 1,000+ trajectory program |
RLVR, process supervision, critics, first-bad-step detection, proof repair |
Review and machine-checkability recorded per unit; not blanket formal verification |
3 public inspection records · Licensed reviewed packs and holdouts |
OlympiadNet-Math 20,000+ tiered records |
Olympiad SFT, proof attempts, preference data, promoted RLVR tasks |
Quality tiers and promotion gates; only a strict subset is positive-weight RLVR |
10 public inspection records · Tiered private corpus |
ArxivNet 57,648 canonical records |
Long-context math SFT, reconstruction, PRM, critics, preferences |
Structurally validated candidate layer; mathematical review required |
Private candidate and review-ready exports |
EdgeReason 5,000 RL tasks |
Tool policy, routing, structured output, DPO, compact reasoning |
Deterministic manifests and adversarial tests; no live tool execution in v0 |
Public · Apache-2.0 |
ErdősBench Research reasoning benchmark |
Research behavior, proof gaps, counterexamples, calibrated progress |
Judge and proof audit; not formal theorem certification |
Public smoke test · Private full benchmark |
SimoBench Olympiad proof benchmark |
Olympiad proof construction, partial credit, false-solve analysis |
0–7 mathematical grading against references |
Public research benchmark · Private refreshes |
UlamGym Ultra-hard research tasks |
Long-horizon mathematical proof search, self-verification, and recovery |
Private task, structured checks, task-specific probes, and expert review |
Private access |
MathCode Stateful coding environment |
Stateful coding-agent work, typed tools, repository edits, and terminal rewards |
Public contract grader on one expired task |
Public research demo · Custom license |
| Mathlib-derived informal SFT |
Theorem-use language, reasoning atoms, pattern recognition |
Machine-checked source material; informal generated targets |
Private access |
| UlamAI Prover traces |
Tactic SFT, proof repair, critic training, RLVR, regression |
Lean-accepted steps and replayable proof-state outcomes |
Open-source prover · Custom/private traces |
| Custom data packs |
Model-specific weaknesses and private regression suites |
Verifier contract designed around the target task |
Private engagement |