Verified Research Reasoning 1,000+ trajectory program |
RLVR, process supervision, critics, first-bad-step detection, proof repair |
Review and machine-checkability recorded per unit; not blanket formal verification |
3 public inspection records · Licensed reviewed packs and holdouts |
OlympiadNet-Math 20,000+ tiered records |
Olympiad SFT, proof attempts, preference data, promoted RLVR tasks |
Quality tiers and promotion gates; only a strict subset is positive-weight RLVR |
10 public inspection records · Tiered private corpus |
ArxivNet 57,648 canonical records |
Long-context math SFT, reconstruction, PRM, critics, preferences |
Structurally validated candidate layer; mathematical review required |
Private candidate and review-ready exports |
Private Mathematics Problems 100,000+ internally created, unpublished problems |
AIME and olympiad training or evaluation, graduate and PhD reasoning, RLVR, and open-research work |
Exact-answer contracts where appropriate; RL tasks are math problems with executable verifiers; open problems use scoped computation and expert review |
Private packs and holdouts · Public samples: AIME++, Math RL Tasks, and SOTA Math |
AIME++ Exact-answer mathematics |
Competition-to-hard exact-answer RLVR, scalable curricula, and sealed private evaluation |
Integer normalization and exact match; boxed-answer parsing is reported separately |
Private packs · View sample |
AIME-Graduate Graduate mathematics with an exact-answer contract |
Graduate-level RLVR with a deterministic terminal reward |
Integer normalization and exact match; boxed-answer parsing is reported separately |
Private packs · View sample |
AIME-Researcher Research-level mathematics with an exact-answer contract |
Research-level exact-answer RLVR and frontier difficulty profiling |
Integer normalization and exact match; boxed-answer parsing is reported separately |
Private packs · View sample |
Math RL Foundational research-oriented RL suites |
Broad mathematical curricula and deterministic verifier integration |
Task-specific exact, structured-output, or certificate verifiers |
Private suites · View sample |
MathH RL Medium-to-hard research-oriented RL suites |
Verifier-backed capability evaluation and advanced certificate curricula |
Task-specific exact and certificate verifiers, including theorem-tactic and structured-witness checks |
Private suites · View sample |
MathR RL Frontier research-oriented RL suites |
Research-math agent training, difficult private evaluation, and high-value certificate generation |
Strict task-specific exact, witness, construction, and certificate verifiers |
Private suites · View sample |
CritPt-derived Physics RL environments |
Benchmark-targeted post-training and clean-room private evaluation |
Deterministic or exact checks, packaged verifier code, executable tests, and schema checks |
Private Python environments · View sample |
GDPVal-derived Synthetic professional-work RL environments |
Long-horizon professional agents, artifact production, quantitative reasoning, and cross-file consistency |
Deterministic dense rewards across correctness, instruction following, traceability, artifact integrity, consistency, quality, and safety |
Private Harbor-compatible environments · View sample |
GDP.pdf-derived Long-horizon multi-document agent RL |
Retrieval, rule reconciliation, quantitative reasoning, evidence attribution, optimization, and structured decisions |
Deterministic dense and sparse grading requires correct values, physical-page or source-record evidence, and passing prerequisite chains |
Private DocumentEnv delivery · View sample |
GPQA Diamond-derived Graduate-science RL environments |
Benchmark-targeted post-training, clean-room evaluation, and repeatable science-reasoning rewards |
Deterministic or exact checks, packaged verifier code, executable tests, and schema checks |
Private Harbor-compatible environments · View sample |
Humanity's Last Exam-derived Frontier-reasoning environments |
Benchmark-targeted post-training, clean-room evaluation, and difficult reasoning curricula |
Deterministic or exact checks, packaged verifier code, executable tests, and schema checks |
Private Harbor-compatible environments · View sample |
Terminal-Bench-derived Hard terminal-agent environments |
Terminal-agent post-training and clean-room private evaluation |
Deterministic or exact checks, packaged verifier code, executable tests, and structured-output checks |
Private Harbor-compatible environments · View sample |
Terminal-Bench-Science-derived Scientific-agent environments |
Scientific-agent RL, hidden-instance generalization, reusable solver evaluation, and mathematical science |
Exact checks, certificate or witness validation, stateful simulator scoring, executable tests, and schema checks |
Private Harbor-compatible environments · View sample |
Terminal-Bench Science Hard Hard scientific-agent RL environments |
Numerical implementation, experimental design, system identification, uncertainty-aware control, transfer, and robustness |
Hidden numerical packets and executable scientific-campaign evaluators with aggregate, robustness, integrity, and protocol gates |
Private scientific environments · View sample |
SciCode-derived Scientific coding RL environments |
Scientific coding-agent RL and weighted numerical evaluation |
Weighted numerical checks, executable tests, and structured-output checks |
Private Python environments · View sample |
SciDiamond-derived Synthetic advanced-science RL |
Advanced-science RLVR curricula and locked-test evaluation |
Deterministic exact four-option rewards, packaged verifier code, and executable tests |
Private Harbor-compatible environments · View sample |
Tau-Banking-derived Stateful banking-agent environments |
Stateful tool-use RL, banking-policy compliance, and database-state evaluation |
Deterministic outcome and compliance checks, packaged verifier code, and executable tests |
Private Python environments · View sample |
CyberSec-derived Synthetic cybersecurity repair environments |
Defensive coding-agent RL and local incident or security-system repair |
Deterministic partial-credit checks, packaged verifier code, and structured-output checks |
Private Python environments · View sample |
AA-LCR-derived Long-context reasoning environments |
Long-context structured reasoning, precise extraction, and private-gold evaluation |
Deterministic JSON verification with shaped field rewards and strict pass/fail |
Private Python environments · View sample |
AA-Omni-derived Factual-reliability and abstention environments |
Factual-reliability RL, calibration, and abstention behavior |
Deterministic verification with correct, incorrect, partial, and not-attempted outcomes |
Private Python environments · View sample |
ErdősBench Research reasoning benchmark |
Research behavior, proof gaps, counterexamples, calibrated progress |
Judge and proof audit; not formal theorem certification |
Public smoke test · Private full benchmark |
SimoBench Olympiad proof benchmark |
Olympiad proof construction, partial credit, false-solve analysis |
0–7 mathematical grading against references |
Public research benchmark · Private refreshes |
| Custom data packs |
Model-specific weaknesses and private regression suites |
Verifier contract designed around the target task |
Private engagement |