Ulam / Scientific reasoning evaluation

STEMBench.

Research progress.
Across six STEM domains.

A correctness-first evaluation of AI research across biology, chemistry, physics, astronomy, Earth science, and engineering. Compare supported scientific progress, reproducibility, and how honestly a model states what remains unresolved.

120tasks per full run
6domains · 20 tasks each
5runs · 600 scored responses
0full-target successes across all runs
01 / Domain comparison

A field-by-field view.

Experimental fractional progress scores, from 0 to 4. Higher is better.

STEMBench120 · v1.0.0Audit snapshot ·

5 of 5 runs · 20 tasks per domain

2.172.99+
Below scale (< 2.17) Bold = highest in row among visible runs

Each domain contains 20 tasks. Values follow the published audit aggregates. Scores are on a 0–4 scale; color focuses on the 2.17–2.99 range to show differences clearly. Lower scores use navy; scores above the display range use yellow. This color range stays fixed when filtering. The average includes every task, including zero-score attempts. Fractional scores describe progress, not a percentage solved or a probability of correctness.

02 / Full-run results

The score, with its context.

Ranked by overall fractional average within the selected run setting.

0 · No usable progress 1 · Framing / protocol 2 · Meaningful partial 3 · Substantial progress 4 · Full target

Integer progress scores remain authoritative. All five runs have zero full-target successes. Small differences in the experimental fractional average should be read alongside the progress distribution, run setting, and review uncertainty.

03 / Scoring & audit

Credit for work that holds up.

Progress, scientific support, and remaining obligations are assessed separately.

Integer progress comes first

The 0–4 judgment records how far the response advances the original target. A useful reduction, derivation, or research protocol can earn credit while the full task remains unresolved.

  1. 0No usable progress. Missing, unusable, or unsupported output.
  2. 1Framing or protocol. Useful synthesis or a proposed approach.
  3. 2Meaningful partial progress. Supported work with consequential obligations still open.
  4. 3Substantial target-level progress. Strong progress on the original task, short of full completion.
  5. 4Full-target success. The original target is met.

Fractional scores add resolution

An experimental quality index, Q, distinguishes work within levels 1–3. It preserves each task’s integer band and its full-target outcome. Level 0 stays at 0; level 4 stays at 4.

Fractional score = s + 0.95 × Q
for integer score s ∈ {1, 2, 3}
Supported achievement strength
30%
Original-obligation coverage proxy
20%
Verification and reproducibility
15%
Scientific support
15%
Validity surviving decisive gaps
10%
Scope honesty and content integrity
10%

Reading precision responsibly. The audit uses per-task review-sensitivity allowances of ±0.05, ±0.10, or ±0.20 according to review confidence. These are review-sensitivity ranges, not statistical confidence intervals for model means. The comparison report treats task-score differences within 0.05 as practical ties. The fractional index is experimental; decimal precision does not establish a decisive capability gap.

Artifact audit: Mistral Large 4

The supplied audit matched all 120 tasks to the run manifest and verified all 507 declared artifacts against their hashes. Of the 120 attempts, 118 contained solver-authored content; two repeated the input instead of providing an answer. Of 42 Python files, 40 passed syntax checks. A short independent execution pass recorded 35 successful exits, six nonzero exits, and one timeout. Five transport errors had recoverable paper content. Artifact integrity and successful code execution provide audit evidence; they do not certify the full scientific target.

This page publishes aggregate results only. Task statements, per-task judgments, and model submissions remain outside the public comparison.

Evaluate your model across STEM.

Work with Ulam on private evaluations, domain comparisons, and evidence that helps turn research failures into better training data.

Discuss an evaluation ↗