Conventional math evaluation
Answer-levelUseful for local competence. It often hides whether the model noticed its own gaps, tested fragile steps, repaired a failed line of reasoning, or simply stopped when the prose looked complete.
// Long-horizon mathematical RL
UlamGym is built from manually hand-crafted, research-level mathematics tasks. The first task is private and intentionally long-horizon. In internal attempts, multiple frontier models have announced a solution—then retracted it when later checks exposed a gap.
// 01
UlamGym is designed around the part of mathematics that short-form evaluations erase: building a research plan, tracking dependencies, challenging one’s own lemmas, revising after contradictions, and knowing when confidence has outrun evidence.
THE UNIT OF EVALUATION IS
A RESEARCH TRAJECTORY.
Useful for local competence. It often hides whether the model noticed its own gaps, tested fragile steps, repaired a failed line of reasoning, or simply stopped when the prose looked complete.
The environment makes premature closure, circularity, hidden assumptions, self-correction, and uncertainty management visible as first-class research behavior.
// 02
A UlamGym episode maintains a durable research state. The policy decides what to formalize, which branch to pursue, what deserves verification, when to challenge a result, and whether a claimed solution should survive another round of scrutiny.
ORIENT → CONJECTURE → PROVE
→ AUDIT → REVISE.
Stage 01 / research map
The model identifies definitions, equivalent formulations, limiting cases, plausible invariants, known tools, and the parts of the statement that are most likely to conceal a trap.
Stage 02 / hypothesis generation
The policy generates lemmas, reductions, invariants, and counterexamples while attaching explicit preconditions and confidence. Strong conjectures are useful because they create sharp tests, not because they sound inevitable.
Stage 03 / dependency construction
Every local argument is connected to the exact hypotheses it uses. Open obligations, imported theorems, numerical checks, and heuristic steps remain labeled rather than being smoothed into a polished narrative.
Stage 04 / adversarial verification
The agent replays the proof from the bottom up, searches for assumption drift, asks whether a lemma is being used in its own proof, and tests the most brittle transition under adversarial examples.
Stage 05 / recovery
The environment distinguishes productive self-correction from mere restart. The policy must localize the failure, preserve valid substructure, update confidence, and choose whether to repair, branch, weaken the claim, or stop.
// 03
The current UlamGym task is not publicly disclosed. Its value depends on preserving the statement, task-specific probes, expert annotations, and the history of model attempts as a controlled research surface.
NO PUBLIC STATEMENT.
NO PUBLIC LEADERBOARD.
A manually constructed, long-horizon mathematics task designed to force deep dependency management rather than fast pattern completion. The public page does not reveal its statement, field, organizer knowledge state, or the details of prior model attempts.
// 04
When a model retracts a solution, UlamGym does not reduce the episode to zero. It asks what kind of mistake occurred, when it became detectable, whether the model found it itself, and whether the revision improved the research state.
ERRORS ARE TRAJECTORIES.
NOT JUST LABELS.
// 05
UlamGym is intended to reward verified mathematical progress, useful error localization, productive revision, and calibrated stopping—while penalizing unsupported closure and repeated dependence on invalid steps.
CONCEPTUAL REWARD DESIGN.
PRIVATE SPEC EVOLVING.
Research-episode objective
A model can make valuable progress without solving the full problem: proving a reusable lemma, finding a decisive counterexample, isolating a minimal invalid core, or correctly retracting an overclaim can all improve the research state. The private scoring protocol remains task-specific.
progressdiagnosisrevisionbeliefcost// 06
The research policy operates through a persistent task session. Ulam controls the private statement distribution, task-specific checks, expert annotations, evaluation rubric, and the final interpretation of what the trajectory establishes.
A PRIVATE RESEARCH ENVIRONMENT.
NOT AN OPEN LEADERBOARD.
// Persistent research boundary
Lab-operated model// 07
UlamGym focuses on situations where a model can be locally impressive, globally wrong, and uncertain about the difference. The environment is intentionally small, expensive, and diagnostic.
DEPTH OVER VOLUME.
FORENSICS OVER A SINGLE SCORE.
Every task is designed and reviewed as a research object rather than sampled from a template generator.
The environment preserves research memory, abandoned branches, proof debt, and revisions across an extended episode.
Every conclusion can be traced back to the exact lemmas, checks, assumptions, and model actions that support it.
A withdrawn claim is analyzed for timing, cause, completeness, calibration, and the quality of the next research move.
Structured probes and executable checks are combined with expert mathematical review where generic automation ends.
Private access, controlled traces, and no public leaderboard reduce contamination and repeated tuning to the task.
Compare not only who “solves,” but how models form beliefs, defend weak steps, discover contradictions, and recover.
Verified lemmas, invalid cores, counterexample episodes, and repaired trajectories can become targeted post-training assets.
// 08
UlamGym is not currently a public benchmark. Early work should freeze a model and scaffold, run controlled episodes on the private task, and close with a forensic comparison of claims, errors, retractions, and verified progress.
ONE TASK CAN STILL PRODUCE
A DEEP MODEL DIAGNOSIS.
Choose the model, scaffold, tool policy, episode budget, assistance rules, and the behaviors the comparison is meant to distinguish.
Execute independent long-horizon episodes from clean state while preserving every claim, check, revision, and resource event.
Review the proof frontier, minimal invalid cores, retraction timing, surviving lemmas, confidence shifts, and repeated failure motifs.
Convert verified mistakes and productive recoveries into targeted data, critics, preferences, probes, or a refreshed task variant.
// 09
The output is not just “solved” or “failed.” A UlamGym study can show where a model’s proof frontier advanced, why a solution claim collapsed, how uncertainty changed, and which behaviors are promising targets for post-training.
A RETRACTION CAN BE
MORE INFORMATIVE THAN A PASS.
Chronological record of hypotheses, proof steps, checks, tool calls, claims, reversals, and stopping decisions.
Verified lemmas, unresolved obligations, imported facts, heuristic steps, and the exact status of the final claim.
Machine-inspectable support relationships that expose hidden assumptions, circularity, and proof debt.
When confidence peaked, what evidence reversed it, which components survived, and whether the correction generalized.
Confidence versus verification state across lemmas, global claims, contradictions, and final stopping behavior.
Minimal invalid cores, repaired proof segments, counterexample searches, critic traces, and preference-ready comparisons.
Customer-safe output
The report centers the pinned model setup, verified mathematical frontier, claim history, error taxonomy, calibration, resource use, and limitations. The private task statement and sensitive traces remain excluded unless explicitly agreed.
// 10
The current program is intentionally explicit about what exists today, what remains private, and what a result would—and would not—show.
ONE PRIVATE TASK TODAY.
MORE ONLY WHEN THEY ARE READY.
No. The statement, mathematical domain, organizer notes, task-specific checks, and prior model traces are private. This prevents contamination and preserves the value of the task as a controlled research surface.
UlamGym currently contains one private task, while we hand-craft additional research-level tasks. These are not interchangeable benchmark rows: a single carefully authored problem can support many long-horizon episodes, model comparisons, failure analyses, verifier refinements, and training loops. The design goal is depth and diagnostic value, not catalog size.
A retracted solution is not a solved task, but the retraction itself can be a positive research behavior. UlamGym distinguishes early self-correction, late correction, external correction, incomplete correction, and doubling down after contradictory evidence.
A research-grade resolution whose dependencies survive the private evaluation process. Depending on the task and organizer state, a rigorous partial result or a precise explanation of the remaining obstruction may also be valuable—but unsupported confidence is never equivalent to a proof.
Not exclusively. Structured and executable checks are used wherever possible, but research mathematics often exceeds the coverage of a single proof assistant or generic verifier. The environment combines explicit proof-state structure, task-specific probes, and expert review.
Yes, under an agreed data-use scope. Verified lemmas, counterexample searches, minimal invalid cores, repaired arguments, confidence corrections, and critic comparisons can become SFT, preference, verifier, or RL signals.
No single private task can establish broad human-level research ability. It can provide unusually deep evidence about one model’s proof search, self-verification, epistemic calibration, recovery, and behavior on an ultra-hard task near the edge of current capability.
// UlamGym private research
Bring a frontier model, a frozen scaffold, and a precise question about mathematical reasoning. Ulam will operate the private task, preserve the full research trace, audit the proof frontier, and turn the failures into the next training opportunity.