ULAM BLOG · 2026
Hillclimbing AI Benchmarks with RL Tasks
How clean-room, verifier-backed RL task families turn benchmark failure modes into targeted training curricula while preserving sealed evaluation boundaries.
Published September 14, 2026
Research notes, technical releases, and analysis from Ulam—covering machine reasoning, mathematical discovery, verification, model efficiency, benchmarks, and practical AI systems.
14 articles
Research notes, releases, and essays from Ulam.
ULAM BLOG · 2026
How clean-room, verifier-backed RL task families turn benchmark failure modes into targeted training curricula while preserving sealed evaluation boundaries.
Published September 14, 2026
ULAM BLOG · 2026
A theory of adaptation across changing environments, and why more varied experience does not automatically produce better reinforcement learning.
Published September 7, 2026
ULAM BLOG · 2026
A 3.086B mathematical reasoning model selected for research-style problem solving, transparent evaluation, and practical local deployment.
Published August 20, 2026
ULAM BLOG · 2026
Private research-level mathematics problems, executable verifier-backed tasks, and audited solving trajectories for frontier evaluation, SFT, RLVR, and long-horizon agents.
Published August 18, 2026
ULAM BLOG · 2026
An arXiv-like public record for AI-generated research, with explicit verification scope, downloadable source, and a foundation for building better reasoning datasets.
Published July 31, 2026
ULAM BLOG · 2026
Public benchmarks are useful, but they are not release gates. Private evaluations reveal whether a model can reason reliably on the capabilities that matter - and what to train next.
Published July 21, 2026
ULAM BLOG · 2026
RAVE introduces router-aware virtual experts—a middle path between pruning and merging MoE experts. In an internal SU-01 proof-judging pilot, RAVE-64 led the aggregate leaderboard while using a 50% expert-centroid budget.
Published June 6, 2026
ULAM BLOG · 2026
We are building a low-contamination math benchmark from synthetic Erdős-style research problems and IMO-like proof problems—testing not just final answers, but rigor, partial progress, and honesty about gaps.
Published May 26, 2026
ULAM BLOG · 2026
We used GPT-5.4 Pro, Opus 4.6, and GPT-5.2 to attack open problems in mathematics—and one of them is now fully settled, with a machine-checked proof in Lean 4.
Published March 18, 2026
ULAM BLOG · 2026
We're releasing UlamAI Prover, a truth-first CLI that combines LLM-guided reasoning with Lean 4 verification to produce machine-checked proofs—no hallucinations, no trust required.
Published February 14, 2026
ULAM BLOG · 2026
We're releasing a research repository that uses LLMs to verify Gaitsgory's proof of the Geometric Langlands Conjecture and transfer its methods to ℓ-adic and p-adic settings—opening a new paradigm for testing AI mathematical reasoning on theory-building, not just problem-solving.
Published February 9, 2026
ULAM BLOG · 2026
We're releasing UnsolvedMath, a curated collection of 1,146 open mathematical problems designed to benchmark AI reasoning capabilities on problems that humanity hasn't yet solved.
Published February 3, 2026
ULAM BLOG · 2026
We're building an AI research lab focused on inference economics—methods that make modern language models cheaper, faster, and more reliable in the real world.
Published January 5, 2026
ULAM BLOG · 2017
Few words about the blog. Welcome to the ULAM blog where we share our thoughts on artificial intelligence, machine learning, and their practical applications in business.
Our articles explore both theoretical aspects of AI research and real-world implementations across various industries.
Published September 15, 2017
Try a broader search or a different year.