Datasets · Benchmarks · Verifiers

Find where reasoning fails.
Train the fix.

Ulam builds reasoning datasets that preserve the evidence between prompt and answer: attempts, proof units, tool calls, negative traces, critiques, repairs, preferences, and verifier outcomes. Inspect our public releases or commission reviewed corpora, private holdouts, and reward-ready exports.

1,000+trajectory research program · 3 public inspection records
20,000+tiered OlympiadNet records across reviewed and candidate layers
57,648ArxivNet canonical proof-process records
1 RL gymUlamGym

Flagship reasoning datasets and UlamGym

From inspectable public schemas to large candidate layers and reviewed private packs—built for training, criticism, verification, and evaluation rather than answer-only supervision. Our autoresearch RL gym extends that work into a stateful, verifier-backed research loop.

Research proof processes

Verified Research Reasoning

1,000+ trajectory program

Proof-process data that records where research reasoning becomes valid, conditional, incomplete, or wrong: dependency-linked proof units, gaps, first-bad-step labels, negative traces, rewards, critiques, repairs, preferences, and adversarial tests.

The public inspection release includes 3 canonical trajectories, 29 PVUs, 13 negative traces, 17 evaluation tasks, schema, transcripts, and validators. Machine-checkability applies only where indicated.

Public inspection release · Licensed packs
Olympiad proofs · SFT · RLVR

OlympiadNet-Math

20,000+ tiered records

Proof and final-answer data with source solutions kept separate from model attempts, plus proof units, negative traces, preferences, adversarial tests, review queues, and promotion gates.

Strict RLVR rows are an explicitly promoted subset—not a label for the whole corpus. The public release provides 10 inspection examples, schema, and quality-gate definitions.

Public inspection release · Tiered private corpus
Research corpus · SFT · PRM

ArxivNet — arXiv Math Blueprints

57,648 canonical records

Turns 5,387 TeX source files into structured proof-process data for long-context SFT, proof reconstruction, theorem dependencies, process supervision, critic training, preferences, and adversarial verifier tests.

April 2026 build: 115,296 SFT rows, 166,817 process-step candidates, 159,329 negative traces, 57,648 preference pairs, and 115,296 adversarial tests. Structural validation does not establish mathematical truth.

Large candidate layer · Review required
Autoresearch · RLVR · Verifiers

Autoresearch RL Gym

1 RL gym

Verifier-backed environments for agents that must run a research loop—not simply submit a final answer. The gym combines a bounded action interface, stateful observations, measurable progress, and acceptance checks that turn failure into a training signal.

  • UlamGym gives agents ultra-hard research tasks for long-horizon mathematical proof search, self-verification, and recovery.
Research environment

Benchmarks and environments

The evaluation layer shows whether a model can use the data: reason under uncertainty, construct valid proofs, and act correctly through tools.

Research benchmark

ErdősBench

226 research candidates testing obstruction finding, counterexamples, finite experiments, theorem use, proof gaps, partial progress, and calibrated claims.

14-problem public smoke test; private full evaluation.

Olympiad benchmark

SimoBench

126 synthetic olympiad-style problems, selected from 1,260 variants, with reference solutions and 0–7 proof scoring for valid, partial, and false solves.

Mathematical grading, not machine-checked formal proof.

Coding-agent environment

MathCode Mini

A runnable compatibility environment with 8 typed tools, real repository edits and tests, state hashes, replayable events, and a terminal reward contract.

One public expired task—not an active leaderboard.

Training signals beyond final answers

Use inspectable public data, prover-backed traces, or a private build targeted to your model's actual failures.

Natural-language mathematics

Mathlib-derived informal SFT

Compact examples for theorem use, reasoning atoms, and mathematical pattern recognition. EdgeReason includes 10,000 rows derived from Mathlib material.

The source mathematics is machine-checked; generated natural-language targets are informal training data, not Lean-certified proofs.

Formal proof traces

UlamAI Prover + LLM

LLMs propose tactics or proof edits; Lean checks them. Retain proof states, premise retrieval, accepted and rejected actions, errors, repairs, backtracking, and terminal outcomes.

Useful for tactic SFT, critics, proof repair, preference data, RLVR, and regression. Formal status applies to Lean-accepted proof steps.

Private and targeted

Custom failure-driven builds

Bring representative attempts, tool logs, proofs, or evaluation results. Ulam localizes recurring failures and returns positives, negatives, critiques, repairs, preference pairs, verifier contracts, and sealed holdouts.

Delivered as JSONL or Parquet with stable IDs, provenance, splits, quality status, and training-view exports.

One improvement loop

Benchmarks reveal the failure. Verifiers make it measurable. Training data turns the repair into a repeatable signal.

problem → attempt/action → observation → verifier verdict → first failure → critique/repair → training export

01 MeasureRun ErdősBench, SimoBench, MathCode, or customer tasks.
02 DiagnoseLocate the first meaningful error and its dependencies.
03 BuildCreate positives, negatives, repairs, preferences, and variants.
04 VerifyAttach deterministic, Lean, simulator, or reviewer outcomes.
05 Re-runMeasure improvement on sealed holdouts and refreshed tasks.

Dataset catalogue and technical details

Counts support the story; verification and intended use determine whether an asset belongs in training.

Compare products, verification boundaries, and access

Asset Best for Verification boundary Access
Verified Research Reasoning
1,000+ trajectory program
RLVR, process supervision, critics, first-bad-step detection, proof repair Review and machine-checkability recorded per unit; not blanket formal verification 3 public inspection records · Licensed reviewed packs and holdouts
OlympiadNet-Math
20,000+ tiered records
Olympiad SFT, proof attempts, preference data, promoted RLVR tasks Quality tiers and promotion gates; only a strict subset is positive-weight RLVR 10 public inspection records · Tiered private corpus
ArxivNet
57,648 canonical records
Long-context math SFT, reconstruction, PRM, critics, preferences Structurally validated candidate layer; mathematical review required Private candidate and review-ready exports
EdgeReason
5,000 RL tasks
Tool policy, routing, structured output, DPO, compact reasoning Deterministic manifests and adversarial tests; no live tool execution in v0 Public · Apache-2.0
ErdősBench
Research reasoning benchmark
Research behavior, proof gaps, counterexamples, calibrated progress Judge and proof audit; not formal theorem certification Public smoke test · Private full benchmark
SimoBench
Olympiad proof benchmark
Olympiad proof construction, partial credit, false-solve analysis 0–7 mathematical grading against references Public research benchmark · Private refreshes
UlamGym
Ultra-hard research tasks
Long-horizon mathematical proof search, self-verification, and recovery Private task, structured checks, task-specific probes, and expert review Private access
MathCode
Stateful coding environment
Stateful coding-agent work, typed tools, repository edits, and terminal rewards Public contract grader on one expired task Public research demo · Custom license
Mathlib-derived informal SFT Theorem-use language, reasoning atoms, pattern recognition Machine-checked source material; informal generated targets Private access
UlamAI Prover traces Tactic SFT, proof repair, critic training, RLVR, regression Lean-accepted steps and replayable proof-state outcomes Open-source prover · Custom/private traces
Custom data packs Model-specific weaknesses and private regression suites Verifier contract designed around the target task Private engagement

Training views: SFT, DPO, critic and repair data, RLVR/GRPO tasks, process-reward steps, private evaluation, and regression tests.

Reward endpoint: UlamGym can serve verifier rewards while private manifests, rubrics, and answer keys remain server-side.

Start with a dataset you can inspect.

Review Verified Reasoning, OlympiadNet, and EdgeReason—or discuss ArxivNet, larger reviewed packs, sealed holdouts, and custom data built around your model's actual failures.