Verified Reasoning · A-SFT + V-SAO
MeasuredBest ErdősBench result
Qwen3.6-27B · untouched base compared with ArxivNet SFT followed by the V-SAO treatment.
Measured post-training results for OlympiadNet, ArxivNet, and Verified Reasoning. Every comparison is framed against an untouched base checkpoint under the same evaluation contract—never the training dataset measured against itself.
Verified Reasoning · A-SFT + V-SAO
MeasuredQwen3.6-27B · untouched base compared with ArxivNet SFT followed by the V-SAO treatment.
OlympiadNet · O-RLVR
MeasuredQwen3.5-4B-Base · aggregate SimoBench proof score across 126 problems and 882 available points.
ArxivNet · A-SFT
MeasuredQwen3.6-27B after A-SFT is within 3.07 points of the GPT-OSS-120B untouched base—at roughly one quarter of the parameter count.
Each result names the base, treatment route, and evaluation. OlympiadNet, ArxivNet, and Verified Reasoning are training interventions; AIME, GSM8K, SimoBench, and ErdősBench supply the measurement.
Same evaluation items and scoring policy. Where exact revisions, hashes, confidence intervals, or compute ledgers were not supplied in this result set, the page marks them as not yet published rather than inventing values.
Tests whether olympiad training transfers to AIME and SimoBench, with smaller-model gains and a neutral 9B aggregate result.
Tests whether long-context research supervision improves answer completion, research search, and mathematical transfer.
Tests shaped verification rewards alone and after specialized SFT, separating proof-critique skill from broad completion reliability.
The headline boxes show the outcome. The scorecards below show where the lift comes from, where it does not, and how the training route interacts with model scale and prior specialization.
3 scorecards
Percentage points, /4 averages, raw counts, and relative lifts are labeled explicitly and are not interchangeable.
ArxivNet supervised training raises research-usefulness and usable-answer completion across Qwen 4B, 9B, and 27B. Transfer results show gains on GSM8K and AIME, with the AIME improvement driven substantially by better final-answer completion.
ErdosBench light
The biggest visible A-SFT story is that provisional research-search usefulness rises across every exact Qwen base.
ErdosBench light
A-SFT converts many formerly non-usable outputs into answerable research responses, even before looking at deeper reasoning quality.
Transfer benchmarks
The AIME gain is the headline transfer result: under the archived contract, A-SFT pushes GPT-OSS-120B from 22 / 60 to 30 / 60, largely by converting unfinished reasoning into final answers.
ArxivNet A-SFT minus the named base, in percentage points.
The largest consistent A-SFT change is whether the model produces a usable answer at all.
| Exact base | B0 | A-SFT | Lift |
|---|---|---|---|
| Qwen3.5-4B-Base | 16.7% | 62.2% | +45.5 pp |
| Qwen3.5-9B-Base | 10.0% | 54.4% | +44.4 pp |
| Qwen3.6-27B | 13.3% | 60.0% | +46.7 pp |
Qwen3.6-27B A-SFT reaches 33.23% research-usefulness versus 36.30% for GPT-OSS-120B base. Its ErdosBench score is 2.518 / 4 versus 2.438 / 4 for GPT-OSS-120B base.
Separate benchmarks retain their own denominators and scoring contracts.
| Evaluation | Model | B0 | A-SFT | Lift |
|---|---|---|---|---|
| GSM8K | Qwen3.6-27B | 94.77% | 96.89% | +2.12 pp |
| AIME 2024+2025 | GPT-OSS-120B | 22 / 60 | 30 / 60 | +13.33 pp |
A-SFT improves answer completion and research-search usefulness, but the improvement is not equivalent to pure reasoning gain.
GPT-OSS AIME pass@1 rises by 13.33 points largely because long or unfinished reasoning is converted into final answers. The supplied analysis also reports new completed-answer errors, so the gain should not be described as pure reasoning improvement.
Qwen3.6-27B rises from 33.98% to 48.02% overall, with a separately reported +27.91 percentage-point gain on research-aptitude alone, excluding proof-gap items.
The most consistent effect is better completion robustness and more useful research-search output. The effect grows with Qwen model size in the reported research-usefulness slice.
The supplied numbers do not include uncertainty intervals or multi-seed aggregation. Each metric should be presented as the reported comparison, not as a universal capability claim.
Pure V-RLVR strengthens proof-gap diagnosis but remains uneven on broad research-progress completion. The strongest direction is specialized supervised training followed by verified reasoning, with A-SFT + V-SAO ranking first on the supplied ErdosBench slice.
ErdosBench
Pure V-RLVR helps proof critique, but the broad research-progress jump appears only once verified reasoning is layered on top of specialized SFT.
ErdosBench
The broad-score winner is A-SFT + V-SAO, which preserves proof-gap gains while restoring SFT-style completion structure.
What changes qualitatively
The clearest SAO gain is more specific proof-gap detection—naming the missing lemma, theorem-scope issue, or invalid transfer step—while also recovering usable research-progress packets that pure V-RLVR often leaves incomplete.
All seven variants cover 139 / 139 rows. Bars use the native 0–4 score scale.
Verified training is strongest when it follows specialized math SFT rather than replacing it.
| Variant | Research progress / 4 | Readout |
|---|---|---|
| B0 | 0.497 | baseline |
| V-RLVR | 0.823 | +0.326 |
| A-SFT + V-RLVR | 2.297 | 4.62× B0 |
B0 → A-SFT + V-RLVR moves research progress from 0.497 to 2.297 and average score from 2.155 to 2.564. A-SFT + V-SAO then reaches 2.655.
It is the strongest proof-gap model in this slice at 3.415 / 4, but many research-progress rows are missing or placeholder-like, which makes its broad completion profile uneven.
A/B/C/D/F/M counts and length finishes are reported exactly as supplied.
| Rank | Variant | Served model | Coverage | A / B / C / D / F / M | Avg / 4 | Length finishes |
|---|---|---|---|---|---|---|
| 1 | a-sft-v-sao-seed101 | a-sft-v-sao-101 | 139 / 139 | 35 / 70 / 10 / 23 / 0 / 1 | 2.655 | 14 |
| 2 | a-plus-v-seed101 | arxivnet-plus-verified-101 | 139 / 139 | 35 / 66 / 6 / 28 / 1 / 3 | 2.564 | 15 |
| 3 | a-sft-seed202 | arxivnet-sft-202 | 139 / 139 | 32 / 63 / 17 / 23 / 2 / 2 | 2.518 | 24 |
| 4 | a-sft-seed303 | arxivnet-sft-303 | 139 / 139 | 31 / 68 / 8 / 25 / 1 / 6 | 2.491 | 19 |
| 5 | a-sft-seed101 | arxivnet-sft-101 | 139 / 139 | 24 / 70 / 10 / 31 / 2 / 2 | 2.396 | 17 |
| 6 | v-rlvr-seed101 | verified-reasoning-101 | 139 / 139 | 30 / 57 / 8 / 21 / 3 / 20 | 2.219 | 35 |
| 7 | b0 | qwen36-27b | 139 / 139 | 34 / 49 / 6 / 21 / 2 / 27 | 2.155 | 38 |
The clearest gain is specific proof-gap diagnosis: the model names missing lemmas, theorem-scope errors, or invalid transfer steps rather than only saying that a proof is incomplete. The supplied examples include a missing Kac–Rice variance/concentration step, a missing pseudorandomness-preservation lemma for tree packing, an invalid Benford-to-normality jump, and missing orbit-collision estimates for divisor-function maps.
V-SAO also recovers the SFT-style observation + finite check + literature-risk + next-lemma packet that pure V-RLVR often fails to complete. That combination explains why it beats both V-RLVR and the older A-SFT + V-RLVR run.
Surprisingly decent on proof-gap detection and literature triage, but weak on completion robustness. Research progress is 0.497 / 4, with many length hits or placeholder JSON outputs.
Research progress rises to 0.823 and the overall score to 2.219. Proof-gap detection reaches 3.415 / 4, but broad research-progress completion remains uneven.
Research progress reaches 2.297—4.62× the B0 level—and the overall score reaches 2.564 / 4.
The best overall result at 2.655 / 4, with 14 length finishes and stronger specific gap diagnoses plus usable research-progress structure.
For general models, the supplied evidence favors verified training after specialized mathematical SFT. Pure RLVR improves proof critique but can hurt broad completion reliability.
The ranking is a single supplied 139-item slice. It supports relative comparison within this evaluation, not an unqualified claim about every research task.
Outcome RLVR produces a clear but modest SIMOBench gain for Qwen3.5-4B, driven by better literature/release triage, conservative labeling, and scope control. The Qwen3.5-9B overall result is neutral.
SIMOBench outcome rate
The measured O-RLVR gain appears at 4B. At 9B, the overall percentage is unchanged even though the qualitative mix shifts slightly.
4B change profile
The 4B gain comes from literature/release triage, conservative labeling, and scope control—not from a broader finite-solving or open-problem-progress improvement.
Interpretation
Qwen3.5-4B-Base moves from 41.8% (369 / 882) to 44.8% (395 / 882) under O-RLVR. The overall effect at 9B is neutral, with both 9B models ending at 50.2%.
Each percentage is the reported score divided by 882.
The gain comes with a materially thinner weak tail, not more perfect-score outputs.
| Measure | B0 | O-RLVR | Change |
|---|---|---|---|
| Score | 369 / 882 | 395 / 882 | +26 |
| Percentage | 41.8% | 44.8% | +3.0 pp |
| Mean / 7 | 2.93 | 3.13 | +0.20 |
| Median | 2 | 3 | +1 |
| Full 7s | 8 | 7 | −1 |
| Scores ≤2 | 69 | 50 | −19 |
| Cap hits | 124 | 124 | 0 |
The two 9B runs tie overall; the 4B O-RLVR model moves into third place.
| Rank | Model | Score | Percentage | Mean | Median | Full 7s | Scores ≤2 | Cap hits |
|---|---|---|---|---|---|---|---|---|
| 1T | Qwen3.5-9B-O-RLVR | 443 / 882 | 50.2% | 3.52 / 7 | 3 | 10 | 42 | 121 |
| 1T | Qwen3.5-9B-Base | 443 / 882 | 50.2% | 3.52 / 7 | 3 | 13 | 42 | 122 |
| 3 | Qwen3.5-4B-O-RLVR | 395 / 882 | 44.8% | 3.13 / 7 | 3 | 7 | 50 | 124 |
| 4 | Qwen3.5-4B-Base | 369 / 882 | 41.8% | 2.93 / 7 | 2 | 8 | 69 | 124 |
The supplied error analysis attributes the gain primarily to better literature/release triage, conservative labeling, and scope control.
The score rises by 26 points and the 0–2 tail falls from 69 to 50. The number of full 7s does not rise, and the combined finite-solving plus open-problem-progress score does not improve.
The overall effect is neutral at 443 / 882 before and after O-RLVR.
O-RLVR modestly improves research-progress yield at 9B but slightly weakens audit and release-risk calibration, leaving the total score unchanged.
Qwen3.5-4B gains +6.67 percentage points on AIME 2024+2025 under O-RLVR.
The improvement is mainly better triage and calibration rather than stronger finite solving or more open-problem progress.
Research-progress yield improves modestly, while audit and release-risk calibration weaken slightly. The net result is neutral.
SIMOBench and AIME results should remain separate. The supplied data does not support pooling them into a single O-RLVR quality score.
Evidence note: the page uses only the supplied measurements. Percentage points, relative percentages, /4 scores, and raw-count changes are labeled separately.
The three programs improve different failure modes. The scorecards keep those differences visible instead of flattening them into one leaderboard number.
Research-usefulness rises by +9.20, +16.98, and +25.14 points across Qwen 4B, 9B, and 27B, while usable-answer rates rise by more than 44 points in every reported Qwen size.
SIMOBench improves by +3.0 points and the weak tail shrinks sharply, but full 7s do not increase and the combined finite-solving plus open-problem-progress score is unchanged.
Pure V-RLVR is a strong proof-gap specialist. A-SFT + V-SAO is the best broad result at 2.655 / 4, 23.2% above B0 and ahead of every supplied SFT or RLVR alternative.
These scorecards preserve the model, benchmark, denominator, failure mode, and interpretation boundary behind each number—from raw SIMOBench points to ErdosBench research-progress and proof-gap ratings.