Training Lift Scorecards

Frontier Data for Frontier Models.

Measured post-training results for OlympiadNet, ArxivNet, and Verified Reasoning. Every comparison is framed against an untouched base checkpoint under the same evaluation contract—never the training dataset measured against itself.

  • Exact untouched control
  • Same evaluation items
  • Matched inference settings
  • Dataset-specific attribution
  • Explicit verification boundary

Verified Reasoning · A-SFT + V-SAO

Measured

Best ErdősBench result

Qwen3.6-27B · untouched base compared with ArxivNet SFT followed by the V-SAO treatment.

Untouched B02.155
A-SFT + V-SAO2.655
+23.2%+0.500 average points on the same 139-item ErdősBench evaluation.

OlympiadNet · O-RLVR

Measured

Small-model proof lift

Qwen3.5-4B-Base · aggregate SimoBench proof score across 126 problems and 882 available points.

Untouched base41.8%
O-RLVR44.8%
+3.0 pp369 / 882 → 395 / 882: 26 additional scored proof points.

ArxivNet · A-SFT

Measured

27B reaches near-120B research usefulness

Qwen3.6-27B after A-SFT is within 3.07 points of the GPT-OSS-120B untouched base—at roughly one quarter of the parameter count.

27B A-SFT33.23%
120B base36.30%
≈4× scale gapThe 27B A-SFT checkpoint also scores 2.518 overall versus 2.438 for the 120B base.
Evaluation contract

The lift thanks to the intervention.

Each result names the base, treatment route, and evaluation. OlympiadNet, ArxivNet, and Verified Reasoning are training interventions; AIME, GSM8K, SimoBench, and ErdősBench supply the measurement.

Post-trained checkpoint − exact untouched base

Same evaluation items and scoring policy. Where exact revisions, hashes, confidence intervals, or compute ledgers were not supplied in this result set, the page marks them as not yet published rather than inventing values.

base × treatment × evaluation
OlympiadNet → O-RLVR

Tests whether olympiad training transfers to AIME and SimoBench, with smaller-model gains and a neutral 9B aggregate result.

ArxivNet → A-SFT

Tests whether long-context research supervision improves answer completion, research search, and mathematical transfer.

Verified Reasoning → V-RLVR / V-SAO

Tests shaped verification rewards alone and after specialized SFT, separating proof-critique skill from broad completion reliability.

Measured dataset interactions

How each dataset changes the model.

The headline boxes show the outcome. The scorecards below show where the lift comes from, where it does not, and how the training route interacts with model scale and prior specialization.

3 scorecards

Percentage points, /4 averages, raw counts, and relative lifts are labeled explicitly and are not interchangeable.

Scorecard 01 · A-SFT

ArxivNet A-SFT

ArxivNet supervised training raises research-usefulness and usable-answer completion across Qwen 4B, 9B, and 27B. Transfer results show gains on GSM8K and AIME, with the AIME improvement driven substantially by better final-answer completion.

  • ErdosBench light
  • research-usefulness
  • usable-answer completion
  • GSM8K + AIME
+25.14 pp27B research-usefulness
+27.91 pp27B research-aptitude alone
+13.33 ppGPT-OSS AIME 2024+2025
+2.12 pp27B GSM8K
2.518 / 427B A-SFT ErdosBench score

Research-usefulness lift by exact base

ArxivNet A-SFT minus the named base, in percentage points.

lift point estimate

Usable-answer completion

The largest consistent A-SFT change is whether the model produces a usable answer at all.

Exact baseB0A-SFTLift
Qwen3.5-4B-Base16.7%62.2%+45.5 pp
Qwen3.5-9B-Base10.0%54.4%+44.4 pp
Qwen3.6-27B13.3%60.0%+46.7 pp

Scale comparison

Qwen3.6-27B A-SFT reaches 33.23% research-usefulness versus 36.30% for GPT-OSS-120B base. Its ErdosBench score is 2.518 / 4 versus 2.438 / 4 for GPT-OSS-120B base.

Transfer results

Separate benchmarks retain their own denominators and scoring contracts.

EvaluationModelB0A-SFTLift
GSM8KQwen3.6-27B94.77%96.89%+2.12 pp
AIME 2024+2025GPT-OSS-120B22 / 6030 / 60+13.33 pp

What changed—and what did not

A-SFT improves answer completion and research-search usefulness, but the improvement is not equivalent to pure reasoning gain.

Completion is a major part of the AIME gain

GPT-OSS AIME pass@1 rises by 13.33 points largely because long or unfinished reasoning is converted into final answers. The supplied analysis also reports new completed-answer errors, so the gain should not be described as pure reasoning improvement.

Additional ErdosBench light readout

Qwen3.6-27B rises from 33.98% to 48.02% overall, with a separately reported +27.91 percentage-point gain on research-aptitude alone, excluding proof-gap items.

Reported breakdown and interpretation

Research-usefulness

Qwen3.5-4B
7.98% → 17.18% · +9.20 pp
Qwen3.5-9B
5.53% → 22.51% · +16.98 pp
Qwen3.6-27B
8.09% → 33.23% · +25.14 pp
GPT-OSS-120B base
36.30% comparator

Usable answer

Qwen3.5-4B
16.7% → 62.2%
Qwen3.5-9B
10.0% → 54.4%
Qwen3.6-27B
13.3% → 60.0%

Transfer

GSM8K
Qwen3.6-27B: 94.77% → 96.89%
AIME 2024+2025
GPT-OSS-120B: 22 / 60 → 30 / 60

Best 27B readout

ErdosBench light
33.98% → 48.02%
Average score
2.518 / 4 after A-SFT
Research aptitude
+27.91 pp, excluding proof-gap items

Interpretation

The most consistent effect is better completion robustness and more useful research-search output. The effect grows with Qwen model size in the reported research-usefulness slice.

Boundary

The supplied numbers do not include uncertainty intervals or multi-seed aggregation. Each metric should be presented as the reported comparison, not as a universal capability claim.

Scorecard 02 · Verified reasoning

Verified Reasoning RLVR + SAO

Pure V-RLVR strengthens proof-gap diagnosis but remains uneven on broad research-progress completion. The strongest direction is specialized supervised training followed by verified reasoning, with A-SFT + V-SAO ranking first on the supplied ErdosBench slice.

  • V-RLVR
  • A-SFT + V-RLVR
  • A-SFT + V-SAO
  • proof-gap + progress
2.655 / 4A-SFT + V-SAO · rank 1
2.564 / 4A-SFT + V-RLVR · rank 2
2.219 / 4pure V-RLVR · rank 6
0.497 → 2.297research progress · almost 5×
3.415 / 4V-RLVR proof-gap detection

Average ErdosBench score by variant

All seven variants cover 139 / 139 rows. Bars use the native 0–4 score scale.

average / 4 observed score

Research-progress effect

Verified training is strongest when it follows specialized math SFT rather than replacing it.

VariantResearch progress / 4Readout
B00.497baseline
V-RLVR0.823+0.326
A-SFT + V-RLVR2.2974.62× B0

Best post-training direction

B0 → A-SFT + V-RLVR moves research progress from 0.497 to 2.297 and average score from 2.155 to 2.564. A-SFT + V-SAO then reaches 2.655.

Pure V-RLVR is a proof-critique specialist

It is the strongest proof-gap model in this slice at 3.415 / 4, but many research-progress rows are missing or placeholder-like, which makes its broad completion profile uneven.

Full ErdosBench ranking

A/B/C/D/F/M counts and length finishes are reported exactly as supplied.

RankVariantServed modelCoverageA / B / C / D / F / MAvg / 4Length finishes
1a-sft-v-sao-seed101a-sft-v-sao-101139 / 13935 / 70 / 10 / 23 / 0 / 12.65514
2a-plus-v-seed101arxivnet-plus-verified-101139 / 13935 / 66 / 6 / 28 / 1 / 32.56415
3a-sft-seed202arxivnet-sft-202139 / 13932 / 63 / 17 / 23 / 2 / 22.51824
4a-sft-seed303arxivnet-sft-303139 / 13931 / 68 / 8 / 25 / 1 / 62.49119
5a-sft-seed101arxivnet-sft-101139 / 13924 / 70 / 10 / 31 / 2 / 22.39617
6v-rlvr-seed101verified-reasoning-101139 / 13930 / 57 / 8 / 21 / 3 / 202.21935
7b0qwen36-27b139 / 13934 / 49 / 6 / 21 / 2 / 272.15538

Why V-SAO wins this slice

The clearest gain is specific proof-gap diagnosis: the model names missing lemmas, theorem-scope errors, or invalid transfer steps rather than only saying that a proof is incomplete. The supplied examples include a missing Kac–Rice variance/concentration step, a missing pseudorandomness-preservation lemma for tree packing, an invalid Benford-to-normality jump, and missing orbit-collision estimates for divisor-function maps.

V-SAO also recovers the SFT-style observation + finite check + literature-risk + next-lemma packet that pure V-RLVR often fails to complete. That combination explains why it beats both V-RLVR and the older A-SFT + V-RLVR run.

Specialist strengths, completion failures, and route choice

B0

Surprisingly decent on proof-gap detection and literature triage, but weak on completion robustness. Research progress is 0.497 / 4, with many length hits or placeholder JSON outputs.

V-RLVR

Research progress rises to 0.823 and the overall score to 2.219. Proof-gap detection reaches 3.415 / 4, but broad research-progress completion remains uneven.

A-SFT + V-RLVR

Research progress reaches 2.297—4.62× the B0 level—and the overall score reaches 2.564 / 4.

A-SFT + V-SAO

The best overall result at 2.655 / 4, with 14 length finishes and stronger specific gap diagnoses plus usable research-progress structure.

Route choice

For general models, the supplied evidence favors verified training after specialized mathematical SFT. Pure RLVR improves proof critique but can hurt broad completion reliability.

Boundary

The ranking is a single supplied 139-item slice. It supports relative comparison within this evaluation, not an unqualified claim about every research task.

Scorecard 03 · O-RLVR

OlympiadNet O-RLVR

Outcome RLVR produces a clear but modest SIMOBench gain for Qwen3.5-4B, driven by better literature/release triage, conservative labeling, and scope control. The Qwen3.5-9B overall result is neutral.

  • Qwen3.5-4B-Base
  • Qwen3.5-9B-Base
  • 882-point SIMOBench
  • AIME +6.67 pp
+264B raw SIMOBench points
69 → 504B scores at 0–2
2 → 34B median score
+6.67 pp4B AIME 2024+2025
50.2%9B before and after

SIMOBench score percentage

Each percentage is the reported score divided by 882.

score percentage observed result

Qwen3.5-4B change profile

The gain comes with a materially thinner weak tail, not more perfect-score outputs.

MeasureB0O-RLVRChange
Score369 / 882395 / 882+26
Percentage41.8%44.8%+3.0 pp
Mean / 72.933.13+0.20
Median23+1
Full 7s87−1
Scores ≤26950−19
Cap hits1241240

Full SIMOBench ranking

The two 9B runs tie overall; the 4B O-RLVR model moves into third place.

RankModelScorePercentageMeanMedianFull 7sScores ≤2Cap hits
1TQwen3.5-9B-O-RLVR443 / 88250.2%3.52 / 731042121
1TQwen3.5-9B-Base443 / 88250.2%3.52 / 731342122
3Qwen3.5-4B-O-RLVR395 / 88244.8%3.13 / 73750124
4Qwen3.5-4B-Base369 / 88241.8%2.93 / 72869124

What improves at 4B

The supplied error analysis attributes the gain primarily to better literature/release triage, conservative labeling, and scope control.

Small aggregate gain, meaningful tail change

The score rises by 26 points and the 0–2 tail falls from 69 to 50. The number of full 7s does not rise, and the combined finite-solving plus open-problem-progress score does not improve.

What happens at 9B

The overall effect is neutral at 443 / 882 before and after O-RLVR.

Offsetting movement

O-RLVR modestly improves research-progress yield at 9B but slightly weakens audit and release-risk calibration, leaving the total score unchanged.

AIME transfer

Qwen3.5-4B gains +6.67 percentage points on AIME 2024+2025 under O-RLVR.

Behavioral interpretation and evaluation boundary

4B aggregate

Score
369 / 882 → 395 / 882
Percentage
41.8% → 44.8%
Mean
2.93 / 7 → 3.13 / 7

4B tail

Median
2 → 3
Scores ≤2
69 → 50
Full 7s
8 → 7

9B aggregate

Score
443 / 882 → 443 / 882
Percentage
50.2% → 50.2%
Mean
3.52 / 7 → 3.52 / 7

4B interpretation

The improvement is mainly better triage and calibration rather than stronger finite solving or more open-problem progress.

9B interpretation

Research-progress yield improves modestly, while audit and release-risk calibration weaken slightly. The net result is neutral.

Boundary

SIMOBench and AIME results should remain separate. The supplied data does not support pooling them into a single O-RLVR quality score.

No matching scorecards Try a broader search or clear one of the filters.

Evidence note: the page uses only the supplied measurements. Percentage points, relative percentages, /4 scores, and raw-count changes are labeled separately.

Cross-scorecard readout

The strongest result comes from combining broad supervision with verified data.

The three programs improve different failure modes. The scorecards keep those differences visible instead of flattening them into one leaderboard number.

A

A-SFT broadens usable research output

Research-usefulness rises by +9.20, +16.98, and +25.14 points across Qwen 4B, 9B, and 27B, while usable-answer rates rise by more than 44 points in every reported Qwen size.

O

O-RLVR improves 4B triage more than solving

SIMOBench improves by +3.0 points and the weak tail shrinks sharply, but full 7s do not increase and the combined finite-solving plus open-problem-progress score is unchanged.

V

Verified Reasoning RL works best after specialized SFT

Pure V-RLVR is a strong proof-gap specialist. A-SFT + V-SAO is the best broad result at 2.655 / 4, 23.2% above B0 and ahead of every supplied SFT or RLVR alternative.

Frontier data matters when every training route is tied to an exact before-and-after result.

These scorecards preserve the model, benchmark, denominator, failure mode, and interpretation boundary behind each number—from raw SIMOBench points to ErdosBench research-progress and proof-gap ratings.