LLM Math Benchmark Leaderboard: Which LLM Is Best at Math?

Updated August 21, 2026 · Current benchmark data: Vals AI ProofBench v1.1 and Epoch AI FrontierMath v2

Large language models have become so good at standard math benchmarks that some of the tests used to rank them only a year ago are no longer useful at the frontier. AIME, MATH 500 and MGSM have all been archived by Vals AI after performance approached saturation. Vals now lists ProofBench v1.1 as its active mathematics benchmark, reflecting a shift away from simply getting the right numerical answer toward producing mathematical reasoning that can be formally verified.

At the same time, Epoch AI’s FrontierMath v2 pushes models in another direction: exceptionally difficult mathematical problems that can require hours or days of work from specialists. That changes what an LLM math benchmark needs to measure. The important question is no longer:

Can an LLM solve a grade-school word problem? It is increasingly:

Can an LLM solve unfamiliar, expert-level mathematics, maintain a valid chain of reasoning and produce an answer or proof that survives objective verification? This updated LLM math benchmark leaderboard therefore separates current frontier benchmarks from legacy benchmarks. We use:

  • ProofBench v1.1 for formally verified mathematical proofs;
  • FrontierMath v2 for extremely difficult mathematical problem solving;
  • AIME, MATH 500, GSM8K and MGSM as historical and diagnostic benchmarks;
  • MathVista as a separate signal for visual and multimodal mathematical reasoning.

So which is the best LLM for math? There is no defensible single winner across every mathematical task. A model that excels at formal proofs may not be the same model you would choose for visual geometry, multilingual word problems or ordinary tutoring. The current leaderboards below show which models lead the hardest evaluations and, just as importantly, what those scores actually mean.

LLM Math Benchmark Leaderboard: Current Results

ProofBench v1.1 — Best LLMs for Formal Math Proofs

Vals AI updated ProofBench v1.1 on August 19, 2026. It is currently the only active benchmark listed in Vals’ Math category; AIME, MATH 500 and MGSM are archived. The current top five are:

LLM Math Benchmark Leaderboard: ProofBench v1.1 Results

RankModel ProofBench v1.1 Accuracy
1 Claude Opus 599%
2 Claude Fable 595%
3 Kimi K387%
4Aristotle86%
5 GPT-5.6 Sol83%

Source: Vals AI ProofBench v1.1, updated August 19, 2026. Current leader: Claude Opus 5 leads ProofBench v1.1 with 99% accuracy, followed by Claude Fable 5 at 95%. That does not mean Claude Opus 5 has been proven to be the best AI for every type of mathematics.

ProofBench measures something unusually specific and valuable: the ability to construct machine-checkable Lean 4 proofs. That makes it especially relevant to formal mathematics and proof verification rather than everyday arithmetic or homework assistance.

There is also an interesting efficiency story behind the scores. Vals reports that Claude Opus 5 and Claude Fable 5 use far fewer tool calls per problem than much of the leaderboard. Kimi K3, meanwhile, costs less per task but takes considerably longer on average. In other words, accuracy, price and speed are three different questions.

FrontierMath Tier 4 Leaderboard: Research-Level Mathematics

ProofBench asks whether a proof can be formally verified. Epoch AI’s FrontierMath Tier 4 v2 asks a different question:

Can the model solve extremely difficult research-level mathematics in the first place? FrontierMath contains original problems written and reviewed by expert mathematicians. According to Epoch, a typical problem can take a relevant researcher several hours, while the hardest may require multiple days. Current FrontierMath Tier 4 v2 results

LLM Math Benchmark Leaderboard: FrontierMath Tier 4 v2 Results

RankSystemOrganizationAccuracy
1Claude Fable 5 (max)Anthropic87.8% ± 5.2%
2GPT-5.6 Sol (max)OpenAI82.9% ± 5.9%
3GPT-5.6 Sol (pro, max)OpenAI80.5% ± 6.3%
4GPT-5.5 Pro (xhigh)OpenAI78.0% ± 6.5%
5 AI co-mathematicianGoogle DeepMind75.6% ± 6.7%
6Claude Opus 5 (max)Anthropic73.2% ± 7.0%
7GPT-5.5 (xhigh)OpenAI72.5% ± 7.1%
8GPT-5.6 Terra (max)OpenAI70.7% ± 7.2%
9GPT-5.6 Luna (max)OpenAI61.0% ± 7.7%
10GPT-5.4 Pro (xhigh)OpenAI58.5% ± 7.8%

Source: Epoch AI, FrontierMath Tier 4 v2.

Claude Fable 5 currently has the highest score among the general-purpose models shown in Epoch’s Tier 4 v2 table. But notice the uncertainty ranges. Claude Fable 5 scores 87.8% ± 5.2%, while GPT-5.6 Sol scores 82.9% ± 5.9%. That means a five-point leaderboard difference should not automatically be interpreted as proof of a large underlying capability gap.

This is one reason BinaryVerse does not turn mathematically different benchmarks into a made-up composite score.

Why FrontierMath v2 Matters More Than the Old FrontierMath Scores

There is a crucial versioning issue here. On June 12, 2026, Epoch released FrontierMath v2 after a major review found problems in the earlier dataset. Epoch says the update addressed errors in 42% of problems. The current benchmark contains:

  • 295 problems in Tiers 1–3;
  • 43 problems in Tier 4;
  • 338 problems total.

Only twelve are public: ten from Tiers 1–3 and two from Tier 4. Unless otherwise stated, Epoch’s reported results use the private sets. For this reason, this article does not mix FrontierMath v1 and v2 scores. A model’s impressive score on an older benchmark revision should not simply be dropped into a new leaderboard as though the tests were identical.

What Is the Best LLM for Math?

The answer depends on what you mean by “math.”

LLM Math Benchmark Guide: Best Benchmark for Each Math Task

Mathematical TaskBenchmark to ConsultCurrent Takeaway
Formal mathematical proofsProofBench v1.1Claude Opus 5 currently leads
Research-level math problemsFrontierMath Tier 4 v2Claude Fable 5 currently leads general-purpose models
Advanced undergraduate / graduate mathFrontierMath Tiers 1–3 v2More useful than saturated school benchmarks
Competition mathematicsAIMEFrontier models have largely saturated it
High-school contest mathMATH 500Archived by Vals after saturation
Grade-school word problemsGSM8K / MGSMUseful historically, weak frontier discriminator
Multilingual math reasoningMGSMUseful diagnostically but saturated at the top
Visual mathematical reasoningMathVistaSeparate multimodal capability

If your definition of best LLM for math is “best at producing formally valid mathematical proofs,” the current Vals data favors Claude Opus 5. If your definition is “best at extremely difficult research-style problem solving,” Epoch’s current FrontierMath Tier 4 v2 leaderboard puts Claude Fable 5 first among the general-purpose models shown. For ordinary school-level arithmetic, however, these comparisons become less useful because many frontier systems already perform close to the ceiling.

Which LLM Math Benchmarks Still Matter?

Not all benchmark percentages mean the same thing. A 95% score on one test can represent a much harder achievement than 99% on another.

LLM Math Benchmark Comparison: What Each Benchmark Measures

BenchmarkMain CapabilityCurrent Role
ProofBench v1.1Formal Lean 4 proofsActive frontier benchmark
FrontierMath Tier 4 v2Research-level mathematicsActive frontier benchmark
FrontierMath Tiers 1–3 v2Advanced undergrad through early researchActive frontier benchmark
AIMEHigh-level competition mathematicsArchived by Vals
MATH 500High-school contest mathematicsArchived by Vals
GSM8KGrade-school word problemsHistorical / saturated
MGSMMultilingual grade-school mathArchived by Vals
MathVistaVisual mathematical reasoningSeparate multimodal benchmark

The evolution itself tells the story of progress. A few years ago, solving GSM8K reliably was notable. Then the focus shifted toward MATH and AIME. Today, leading models are being evaluated on research-level problems and machine-verifiable proofs because the older tests no longer provide enough separation.

How ProofBench Tests Formal Mathematical Reasoning

Infographic showing how the LLM Math Benchmark ProofBench evaluation verifies formal Lean 4 proofs
Infographic showing how the LLM Math Benchmark ProofBench evaluation verifies formal Lean 4 proofs

ProofBench is fundamentally different from a benchmark that checks whether the final answer is 42. Each task gives the model:

  1. a natural-language mathematical statement; and
  2. a corresponding formal statement written in Lean 4.

The model must then construct a proof accepted by the Lean proof checker. The benchmark covers areas including:

  • probability and stochastic analysis;
  • measure theory;
  • real and functional analysis;
  • algebra and commutative algebra;
  • algebraic geometry;
  • number theory;
  • set theory;
  • logic;
  • model theory.

These are drawn from advanced undergraduate and graduate-level sources, including textbooks and qualifying examinations.

The tools models receive

Within the ProofBench environment, models can use:

ToolPurpose
lean_loogleSearch Mathlib
lean_run_codeExecute Lean code and receive feedback
submit_proofSubmit the final proof

Models may search for existing lemmas, test partial code and revise their proof before submission. They receive up to 40 interaction turns per problem.

Correct means formally correct

There is no score for a response that merely sounds convincing. If Lean accepts the proof, it counts.

If it does not, it fails. Vals also checks accepted submissions for disallowed axioms. Constructs such as sorry and admit cannot be used to bypass the proof obligation. This is a major advantage over ordinary natural-language grading. A long explanation can hide:

  • an unsupported assumption;
  • a missing case;
  • an invalid inference;
  • a circular argument;
  • a subtle algebraic error.

Formal verification makes those mistakes much harder to hide.

Public and private test sets

ProofBench contains:

  • 100 public problems
  • 100 private test problems

The main results use the private test split, helping reduce the contamination problem that affects many older public benchmarks. Vals also explicitly warns that ProofBench v1.1 scores are not comparable with scores from earlier revisions. That is why benchmark version numbers matter.

FrontierMath: When Olympiad Problems Are No Longer Hard Enough

FrontierMath tackles another weakness of traditional LLM math leaderboards: benchmark saturation. The problems are original and span major branches of mathematics, ranging from computational number theory and real analysis to areas such as algebraic geometry and category theory. Epoch divides the current v2 benchmark into two groups.

FrontierMath Tiers 1–3

The 295-problem base set covers difficulty from advanced undergraduate mathematics through exploratory problems suitable for advanced graduate students and early-career researchers.

FrontierMath Tier 4

The 43 Tier 4 problems are the hardest. Epoch describes Tier 4 as research-level mathematics. That distinction makes FrontierMath much harder to saturate than older contests aimed at school students.

Models can use Python

FrontierMath is not a closed-book mental-arithmetic test. The evaluation allows models to reason and execute Python code. The model eventually submits a Python answer() function that returns its final answer. Correct answers receive one point and incorrect or missing submissions receive zero. That means FrontierMath evaluates a modern reasoning system operating with computational tools rather than pretending that tool use does not exist.

GSM8K Leaderboard: Why the Famous Math Benchmark Is Now Saturated

GSM8K was one of the defining early benchmarks for mathematical reasoning in language models. It contains roughly 8,500 grade-school math word problems requiring multi-step reasoning. MGSM was later derived from GSM8K to test multilingual performance.

GSM8K mattered because early language models frequently failed even when the required arithmetic was elementary. A model might understand every word in a problem yet lose track of the sequence of operations needed to solve it. That is no longer where the frontier lies. Modern reasoning models routinely achieve very high scores on GSM8K-like problems, making small differences between top systems difficult to interpret. This is why someone searching for a GSM8K leaderboard today should be careful. A score close to 100% does not demonstrate that the model can:

  • construct a rigorous proof;
  • solve unfamiliar research mathematics;
  • reason correctly about a complex visual diagram;
  • generalize to a different mathematical domain.

GSM8K remains historically important, but it should no longer determine which frontier model is “best at math.”

MGSM: The Multilingual Version of GSM8K

MGSM extends the GSM8K idea across languages. It contains a human-translated subset of 250 GSM8K problems in ten languages, including languages such as Bengali, Telugu and Swahili. That makes it useful for asking whether mathematical reasoning survives changes in language.

Vals’ final MGSM results were tightly clustered. Claude Opus 4.5 Thinking finished first at 95.2%, but Vals concluded that differences among the leading models were small enough that they were likely not statistically significant. Vals also found that models generally performed better in English than in other languages. The benchmark has now been archived because performance became too saturated to remain a useful frontier discriminator. That does not make MGSM useless. It simply changes the question it can answer.

MGSM is more useful for investigating multilingual consistency than for deciding which 2026 frontier model has the strongest overall mathematical reasoning.

Archived 2025 LLM Math Leaderboard: MATH 500 and MGSM

The table below preserves the May 2025 Vals results originally published on this page. It is retained intentionally because it provides a historical snapshot of how quickly LLM math performance changed and remains useful for readers searching older GSM8K, MGSM and MATH benchmark results. It should not be treated as the current frontier leaderboard.

LLM Math Benchmark: Archived 2025 MATH 500 and MGSM Leaderboard

RankModelMATH 500MGSM
1 Gemini 2.5 Pro Exp95.2%92.2%
2 ChatGPT o394.6%91.4%
3Qwen 3 235B94.6%92.7%
4Grok 3 Mini Fast High94.2%90.3%
5GPT-4 Mini94.2%87.9%
6DeepSeek R192.2%92.4%
7Gemini 2.5 Flash Preview Thinking91.8%90.0%
8ChatGPT o3 Mini91.8%91.6%
9Claude 3.7 Sonnet Thinking91.6%92.8%
10Gemini 2.5 Flash Preview91.6%89.8%

These figures come from the original May 2025 version of this BinaryVerse article. The final Vals MATH 500 leaderboard later reached an even higher ceiling. Gemini 3 Pro finished at 96.4%, and Vals reported that most recent frontier models were consistently above 90%. Vals subsequently archived the benchmark because the results had become too saturated and the public dataset also creates a substantial contamination risk.

This is the central lesson of the old table:

the benchmark did not become useless because models became worse. It became less useful because models became too good at it.

AIME: From Elite Math Contest to Saturated LLM Benchmark

The American Invitational Mathematics Examination is not an easy human exam. AIME is aimed at high-performing secondary-school mathematics students and contains 15 problems with integer answers from 000 to 999. Vals evaluated the 2024 and 2025 exams, giving it 60 questions total, and ran each model eight times to reduce variance.

Yet by April 2026, Gemini 3.1 Pro Preview had reached 98.12% in the Vals evaluation. Vals archived AIME because performance had saturated. There is another complication. AIME questions and solutions are public. Vals found that models tended to perform better on the older 2024 questions than the 2025 set, raising concerns about exposure during training. AIME remains useful for understanding competition-style mathematical reasoning, but a near-perfect AIME score is no longer enough to distinguish the strongest frontier models.

MathVista Leaderboard: Testing Visual Math Reasoning

Most traditional math benchmarks are text-only. Real mathematical reasoning often is not. A geometry problem may depend on a diagram. Scientific reasoning may require interpreting a graph. A statistics question may rely on a chart. Mathematical information can be encoded visually rather than written explicitly in the prompt.

MathVista was designed to evaluate this intersection between mathematics and visual understanding. The complete benchmark contains 6,141 examples drawn from 28 existing multimodal datasets plus three newly created datasets: IQTest, FunctionQA and PaperQA. Its official test split contains 5,141 examples with private ground truth, while the testmini subset contains 1,000 examples. MathVista measures areas including:

  • algebraic reasoning;
  • arithmetic;
  • geometry;
  • logical reasoning;
  • numerical reasoning;
  • scientific reasoning;
  • statistical reasoning;
  • visual question answering.

This makes MathVista valuable, but it should remain separate from ProofBench or FrontierMath. A multimodal model interpreting a geometric diagram is solving a different problem from a system formalizing a theorem in Lean. For the same reason, BinaryVerse does not average MathVista accuracy into an overall math score.

Best LLM for Math by Use Case

Best LLM for Math Proofs

For formal proof construction, the strongest current evidence in this comparison comes from ProofBench v1.1. Claude Opus 5 leads at 99%, followed by Claude Fable 5 at 95%, Kimi K3 at 87%, Aristotle at 86% and GPT-5.6 Sol at 83%. The critical advantage of this benchmark is that successful proofs are checked by Lean rather than accepted because they sound convincing. If your work involves theorem formalization or machine-verifiable mathematics, ProofBench is more informative than GSM8K or AIME.

Best LLM for Difficult Math Problems

For extremely difficult research-style problems, FrontierMath Tier 4 v2 is the more appropriate benchmark. Claude Fable 5 currently leads the general-purpose models shown by Epoch with 87.8% ± 5.2%, while GPT-5.6 Sol follows at 82.9% ± 5.9%. The error bars mean those numbers should not be interpreted as an exact ranking of intrinsic intelligence. They are measurements from a finite benchmark under specific evaluation conditions.

Best LLM for Competition Math

AIME was once an excellent separator. It no longer is. Multiple frontier systems now operate close to the ceiling, and Vals has retired the benchmark from ongoing model testing. For comparing today’s strongest models, FrontierMath gives substantially more room between success and saturation.

Best LLM for Multilingual Math

MGSM remains useful for examining how mathematical reasoning changes across languages. But the final Vals scores became so tightly clustered that the benchmark was archived. If multilingual consistency is your priority, inspect performance by language rather than relying only on the overall MGSM average.

Best LLM for Visual Math Problems

For mathematics involving diagrams, plots, charts and other visual information, MathVista is conceptually better suited than text-only GSM8K or MATH. However, benchmark version and model freshness matter. An old MathVista leaderboard should not be used to declare the best 2026 multimodal model simply because newer systems have not been evaluated under the identical official setup.

How We Rank the Best LLMs for Math

BinaryVerse does not create a single composite math score by averaging unrelated benchmarks. ProofBench, FrontierMath, AIME, MATH 500, GSM8K, MGSM and MathVista test different abilities, use different datasets and evaluation procedures, and operate at very different levels of saturation. Our current methodology therefore uses:

ProofBench v1.1 as the primary current benchmark for formally verified mathematical reasoning. FrontierMath v2 as the primary current benchmark for extremely difficult mathematical problem solving.

Legacy benchmarks such as AIME, MATH 500, GSM8K and MGSM are retained for historical comparison and search context, but they do not determine the current frontier ranking because leading models now perform near the ceiling on many of them. We report benchmark scores separately rather than combining percentages into an artificial overall number.

Accuracy, cost and latency are separate

A mathematically stronger model is not automatically the cheapest. The cheapest is not automatically the fastest. Vals explicitly reports:

  • accuracy;
  • latency;
  • cost;
  • tool-use behavior;
  • qualitative error patterns

as separate pieces of information. BinaryVerse follows the same principle.

We preserve statistical uncertainty

Where a benchmark publisher reports error bars, we preserve them. Vals reports standard errors to reflect statistical uncertainty and distinguishes between single-run and multiple-run evaluations. It also warns that these error bars do not capture all variation caused by prompts, random seeds, deployment settings or stochastic generation.

Therefore:

87.8% is not automatically meaningfully better than 82.9% simply because the number is larger. The size of the benchmark and the associated uncertainty matter.

We lock benchmark versions

Results from different benchmark revisions are not mixed unless the publisher explicitly indicates that they are comparable. This is especially important for:

  • ProofBench v1.1;
  • FrontierMath v2.

ProofBench explicitly states that v1.1 results cannot be compared directly with earlier revisions. Epoch’s June 2026 FrontierMath revision corrected a substantial portion of the old dataset, making version labels equally important there.

Why LLM Math Benchmark Scores Can Be Misleading

A leaderboard looks objective. Interpretation is much harder.

1. Benchmark saturation

When nearly every frontier model scores above 90%, a benchmark loses its ability to distinguish among them. This happened to MATH 500, MGSM and AIME on Vals. A ceiling score can mean:

the model is extraordinary or:

the test is no longer hard enough. Often it is some combination of both.

2. Training-data contamination

Public benchmarks are easy to study. They are also easier to accidentally or deliberately include in training corpora. MATH and AIME have been public for years. Vals specifically highlights contamination risk for both. Private and newly created test sets reduce this problem, although they do not eliminate every form of benchmark overfitting.

3. Tool access matters

A model using Python, Lean or a search tool is not being tested under the same conditions as a model forced to answer from its internal computation alone. But banning tools can also create an unrealistic evaluation. Modern AI systems are increasingly designed to use tools. The important thing is therefore not pretending tool use does not exist. It is documenting the evaluation environment clearly.

4. Reasoning effort matters

Some evaluations allow much larger reasoning budgets than others. FrontierMath’s current harness can permit extremely long trajectories before final submission. A “max” or “xhigh” reasoning setting should therefore not automatically be compared with a low-effort configuration as though the computational budgets were identical.

5. Formal proofs and numerical answers are different skills

A model may discover the correct numerical answer without being capable of writing a formal proof. Another may be excellent at Lean formalization but not be the model you would choose to interpret a graph or explain algebra to a beginner. There is no reason to expect one benchmark to measure every mathematical ability.

6. Small leaderboard differences may be noise

FrontierMath Tier 4 contains only 43 problems in v2. That is intentionally difficult, but it also means each problem carries substantial weight. This is why Epoch’s uncertainty ranges around Tier 4 scores are relatively wide. Treat close results as a performance tier, not a horse race decided by one decimal point.

Why Math Benchmarks Keep Getting Harder

LLM Math Benchmark infographic showing six stages of rising math benchmark difficulty
LLM Math Benchmark infographic showing six stages of rising math benchmark difficulty

The history of LLM math evaluation is partly the history of benchmarks becoming obsolete.

Stage 1: GSM8K

Could language models reliably follow a sequence of elementary arithmetic steps? For early systems, often not.

Stage 2: MATH

Researchers moved toward harder high-school competition mathematics requiring more complex algebra, probability, geometry and trigonometry. Eventually top models moved above 90%.

Stage 3: AIME

Competition mathematics raised the difficulty again. By 2026, the leading Vals model was above 98%.

Stage 4: FrontierMath

Instead of asking models to imitate elite high-school competitors, FrontierMath pushes into advanced undergraduate, graduate and research mathematics.

Stage 5: Formal verification

ProofBench changes the success criterion. It is no longer enough to generate a persuasive argument. The proof has to survive Lean.

Stage 6: Open research problems

Epoch has gone another step with FrontierMath: Open Problems, a collection of significant problems that were unsolved by mathematicians when selected. By July 31, 2026, Epoch reported that the collection had expanded to 50 open problems and that AI systems had solved three. This is qualitatively different from benchmark optimization. If AI systems can increasingly make correct progress on genuine open mathematics, the question will eventually move beyond:

How well does AI score on a test?

toward:

Can AI expand mathematical knowledge?

That remains a much higher bar.

The Bottom Line: What Is the Best LLM for Math?

There is no single best LLM for math across every mathematical task. As of August 21, 2026:

  • Claude Opus 5 leads Vals AI’s ProofBench v1.1 at 99%, making it the strongest model in this comparison for formally verified Lean proofs.
  • Claude Fable 5 leads the general-purpose models shown on Epoch’s FrontierMath Tier 4 v2 leaderboard at 87.8% ± 5.2%, with GPT-5.6 Sol close behind at 82.9% ± 5.9%.
  • AIME, MATH 500 and MGSM should no longer decide the frontier winner, because Vals has archived all three after performance saturated.

The deeper lesson is that the definition of “good at math” keeps moving. First, models had to stop making elementary arithmetic mistakes. Then they had to solve competition problems. Now they are being asked to solve research-level mathematics and produce machine-verifiable proofs. That is why the most useful LLM math benchmark leaderboard is not the one with the highest percentages.

It is the one that remains difficult enough to tell us something new.

What is the difference between GSM8K and MGSM?

GSM8K contains roughly 8,500 English grade-school math word problems. MGSM uses a 250-problem subset translated by human annotators into ten languages, allowing researchers to study both mathematical reasoning and multilingual consistency.

Why was MATH 500 archived by Vals AI?

Vals archived MATH 500 because performance had saturated. The leading models were consistently above 90%, while the public nature of the dataset also creates a significant risk of training-data contamination. Vals concluded that the benchmark was becoming less useful for distinguishing the mathematical capabilities of new frontier models.

Is AIME still useful for testing AI math ability?

AIME remains a difficult human mathematics competition and is still useful historically, but it is no longer a strong frontier discriminator. Vals’ top model reached 98.12%, and Vals subsequently archived the benchmark because performance had saturated.

What is MathVista?

MathVista is a multimodal mathematics benchmark designed to test reasoning when mathematical information is presented visually. It contains 6,141 examples involving areas such as geometry, plots, scientific figures, arithmetic and logical reasoning. Its official test set contains 5,141 examples with private ground truth.

Is ChatGPT good at math?

Current OpenAI reasoning models are extremely strong on difficult mathematics, but “ChatGPT” is a product rather than a single fixed model and benchmark scores depend on which underlying model and reasoning setting is used. For example, GPT-5.6 Sol scores 83% on Vals ProofBench v1.1 and 82.9% ± 5.9% on Epoch’s FrontierMath Tier 4 v2 under the reported evaluation settings.

Can LLMs solve previously unsolved mathematics problems?

There is now evidence that AI systems can make genuine progress on some open mathematical problems, but this is far harder than solving standard benchmarks. Epoch AI’s FrontierMath Open Problems collection expanded to 50 significant unsolved problems in July 2026, and Epoch reported that AI had solved three of them at that point. Each claimed solution still requires careful mathematical verification.

Why do different LLM math leaderboards disagree?

Different leaderboards test different skills and use different evaluation settings. ProofBench measures Lean proof success, FrontierMath measures extremely difficult problem solving, MathVista includes visual reasoning, and GSM8K focuses on grade-school word problems. Models may also receive different reasoning budgets or tool access. A leaderboard should therefore be interpreted in the context of what the benchmark actually measures rather than treated as a universal intelligence ranking.

Are higher LLM math benchmark scores always better?

Higher scores are better within the same benchmark version and evaluation setup, but small score differences may not be statistically meaningful. Benchmark size, error bars, tool access, reasoning effort, contamination risk and saturation all matter. Scores from different benchmarks or benchmark versions should not be averaged or compared as though they measure exactly the same ability.

Which math benchmark should I trust most?

For current frontier models, use the benchmark that matches your task. ProofBench v1.1 is particularly useful for formal mathematical proofs, while FrontierMath v2 is better for extremely difficult mathematical problem solving. GSM8K, MGSM, AIME and MATH remain valuable historical benchmarks but are less useful for separating today’s strongest models.