Pith. sign in

REVIEW 2 major objections 2 minor 40 references

Benchmarks in Leipzig

T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read A dataset of 100 research-level math questions with known answers leaves only two unsolved after staged LLM evaluations.

desk verdict The Leipzig paper gives us a new set of 100 math questions compiled by 49 people, but the claim that LLMs are now impressive at research math rests on thin evidence about contamination controls. read the letter →

arxiv 2606.05818 v1 pith:7UY2KIUM submitted 2026-06-04 math.HO cs.AImath.AGmath.COmath.RT

classification math.HOcs.AImath.AGmath.COmath.RT
keywords LLMmathematicalreasoningresearch-levelmathbenchmarksAIevaluationmathematicsdatasetslargelanguagemodelsproblemsolving
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper assembles 100 questions drawn from research mathematics, each with a verified answer, through a collaborative effort by 49 mathematicians. These questions undergo testing first with five leading LLMs in single attempts, then with three models across 20 runs each, and finally with two advanced models in three runs. The process reduces the number of unsolved questions from 41 after the first stage to 16 after the second and to 2 after the third. The authors present this outcome as direct evidence that current LLMs exhibit strong mathematical reasoning on problems at the level of recent or ongoing research.

What carries the argument

The 100-question dataset itself, compiled during a workshop and evaluated in three successive stages of LLM interaction with answer verification.

What would settle it

Demonstration that a substantial fraction of the 100 questions appeared in the training corpora of the evaluated LLMs or that the questions do not match the difficulty of typical unsolved research problems.

Watch

Extended reading notes

Core claim

The central claim is that a newly compiled collection of 100 research-level mathematics questions, assembled with known answers, can be solved by state-of-the-art LLMs in all but two cases after multiple evaluation stages that include repeated sampling and verification.

Load-bearing premise

The questions are representative of genuine research-level mathematics and have not appeared in the training data of the tested models.

Editorial extensions

If this is right

  • LLMs can now address mathematical problems that previously required specialized human expertise.
  • Repeated sampling and verification protocols allow models to reach solutions on a large majority of such questions.
  • The remaining unsolved questions provide a focused target for further model improvement.
  • Benchmarks of this type will require regular updates as model performance advances.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The benchmark could serve as a continuing yardstick to measure progress in LLM mathematical capability over time.
  • Success on these questions may indicate readiness for LLMs to contribute to open mathematical research problems.
  • The staged evaluation method highlights the value of multiple attempts and verification in assessing reasoning rather than one-shot performance.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper describes the compilation of a 100-question dataset of research-level mathematics problems with known answers by 49 mathematicians between April 1 and May 15, 2026 (primarily during a 3-day workshop in Leipzig). It reports the results of three staged LLM evaluations: a single attempt by five models (41 unsolved), a 20-run evaluation with three models (16 unsolved), and a 3-run evaluation with two heavy-thinking models (2 unsolved), concluding that LLM mathematical reasoning capabilities are becoming impressive.

Significance. If the questions are verifiably post-training-cutoff and the evaluation protocol includes rigorous controls against contamination, the work would supply concrete evidence of rapid progress on hard, previously unsolved mathematical problems. The multi-mathematician curation and staged evaluation design are methodological strengths that could make the benchmark useful for future comparisons.

major comments (2)
  1. [Abstract] Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching.
  2. [Abstract] Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability.
minor comments (2)
  1. The abstract could state the topical distribution or difficulty stratification of the 100 questions to allow readers to judge representativeness.
  2. Clarify whether the 'known answers' were independently re-derived by the curators or taken from existing literature, and how any discrepancies were resolved.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their detailed and constructive comments, which highlight important aspects of transparency needed for the benchmark's claims. We agree that the manuscript would benefit from additional methodological details and will revise accordingly. Our point-by-point responses follow.

read point-by-point responses
  1. Referee: [Abstract] Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching.

    Authors: We agree the current manuscript lacks an explicit protocol. All 100 questions were newly formulated as original research-level problems by the 49 mathematicians during the 2026 workshop (post-dating all model training cutoffs referenced). We will add a subsection detailing the generation process, including that questions were created de novo without reference to existing public corpora and cross-verified for novelty by multiple participants against arXiv, MathOverflow, and MathStackExchange. This revision will support the reasoning-over-memorization interpretation. revision: yes

  2. Referee: [Abstract] Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability.

    Authors: We concur that these elements are missing from the manuscript. The revised version will include a new section on verification: each solution was confirmed by its originating mathematician and independently reviewed by at least two additional experts, with consensus required. For the multi-run stages, we will report per-run solve counts, observed variance, and the precise criteria for labeling a question 'solved' (correct final answer provided). This will enable readers to evaluate reliability. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical benchmark report with no derivations or self-referential reductions

full rationale

The paper compiles and evaluates a dataset of 100 research-level math questions against LLMs across three stages, reporting unsolved counts dropping from 41 to 2. No equations, fitted parameters, predictions, or derivations are present. The central claim is an empirical observation from external model runs on a newly compiled dataset; it does not reduce to any input by construction, self-citation, or renaming. No load-bearing steps match any of the enumerated circularity patterns.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The paper is an empirical benchmark report with no mathematical derivations, fitted parameters, or new theoretical entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Benchmarks in Leipzig." pith.science (2026). https://pith.science/paper/7UY2KIUM

@misc{pith2026260605818,
  author       = {Pith},
  title        = {Pith review of: Benchmarks in Leipzig},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7UY2KIUM}},
  note         = {Machine review of arXiv:2606.05818}
}
read the original abstract

Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100 questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 3 canonical work pages

  1. [1]

    Ifn= 1, thens(1) = 1

  2. [2]

    This operationsis often called the stack sorting

    Ifp=LnR, that isLis the substring on the left of the entryn, thens(p) =s(L)s(R)n. This operationsis often called the stack sorting. LetB n,k be the set of permutationspof lengthnfor which all of the following hold

  3. [3]

    The permutations(p)avoids the patterns2314and3124

  4. [4]

    The number of descents ofpisk−1. Ablock moveis an operation that turns the permutation p=p 1p2 · · ·p i pi+1 · · ·p j pj+1 · · ·p k pk+1 · · ·p ℓ · · ·p n into the permutation p′ =p 1p2 · · ·p i pk+1 · · ·p ℓ pj+1 · · ·p k pi+1 · · ·p j pℓ+1 · · ·p n. That is, a block move interchanges the string pi+1 · · ·p j and the string pk+1 · · ·p ℓ, for some number...

  5. [5]

    reversing bounce path

    bothpandπare of length30, and 2.p∈B 30,15, and 3.π∈D p(10), that is,πis at distance10fromp? Question 013solved in Stage 1 Fix two positive integers r and m. Consider the family of subsets Ci ={2i+ 1, ...,2i+r} where i ranges from 0 to m−1 . Let M be the matroid of rank r on n= 2(m−1) +r elements whose non-bases are exactly the m subsets C0, ..., Cm−1. Usi...

  6. [6]

    What is the largestksuch thatX k is normal? Question 040solved in Stage 1 Fix seven general points in the real projective plane

    For a natural number k, let Xk be the projective toric variety parametrized by the degree k monomials in Z (insideP(Z k)). What is the largestksuch thatX k is normal? Question 040solved in Stage 1 Fix seven general points in the real projective plane. Consider the 21 lines spanned by any two points, the 21 conics spanned by any five points, and the 7 cubi...

  7. [7]

    and let T= (C ∗)3 15 Benchmarks in Leipzig act on X via the induced action of the multiplication in the domain of this parametrization. A subdivision of a fan consists of picking a maximal cone of the fan and a hyperplane and replacing the cone by the two cones obtained by intersecting the initial cone with the two half-spaces of the hyperplane. What is t...

  8. [8]

    The minimal resolution ofX Γ has exceptional intersection graphE 8

Show all 40 references
  1. [9]

    Equivalently, C[u, v]Γ has primitive homogeneous generators of degrees 12,20,30 , satisfying one weighted- homogeneous relation of degree60

  2. [10]

    The degree-12primitive invariant is normalized as follows. It lies in the one-parameter family Fa(u, v) =u 11v+a u 6v6 −uv 11, anda∈Cis chosen so that, if Ha = Hess(Fa) = det   ∂2Fa ∂u2 ∂2Fa ∂u∂v ∂2Fa ∂v∂u ∂2Fa ∂v2   , and Ta = Jac(Fa, Ha) = det   ∂Fa ∂u ∂Fa ∂v ∂Ha ∂u...

  3. [11]

    the hidden groupΓ, up to conjugacy inSL 2(C)

  4. [12]

    the admissible value or values ofa

  5. [13]

    the exact Fisher informationI= R P1 ρ2 dσ

  6. [14]

    the exact threshold constantC= 2√ I

  7. [15]

    checksum

    the limiting log-Bayes factor, in favor of the hidden quotient-substrate sky over the continuum sky, assuming equal prior odds. Question 047solved in Stage 1 Let X be the variety in (C∗)5 cut out by x3x4 +x 1 −1, x 5 +x 2x3 −1, x 4x5 +x 2x3x4 +x 1x2 −1 . Consider the form Ω = ...

  8. [16]

    Ruhr University Bochum, Bochum, Germany

  9. [17]

    Max Planck Institute for Mathematics in the Sciences, Leipzig, Germany

  10. [18]

    TU Berlin, Berlin, Germany

  11. [19]

    University of Florida, Gainesville, FL, USA

  12. [20]

    Ecole Normale Superieure (ENS), PSL University, Paris, France

  13. [21]

    University of California, Davis, USA

  14. [22]

    University of Barcelona, Barcelona, Spain

  15. [23]

    University of California, Berkeley, USA

  16. [24]

    ETH Zurich, Zurich, Switzerland

  17. [25]

    Universidad de Talca, Talca, Chile

  18. [26]

    TU Dresden, Dresden, Germany

  19. [27]

    University of Bonn, Bonn, Germany

  20. [28]

    Université du Québec à Montréal (UQAM), Montreal, Canada

  21. [29]

    University of Southern California, Los Angeles, USA

  22. [30]

    University of Luxembourg, Luxembourg

  23. [31]

    Amherst College, Amherst, Massachusetts, USA

  24. [32]

    Bielefeld University, Bielefeld, Germany

  25. [33]

    Technische Universität Braunschweig, Braunschweig, Germany

  26. [34]

    University of Leipzig, Leipzig, Germany

  27. [35]

    University of Copenhagen, Copenhagen, Denmark

  28. [36]

    KTH Royal Institute of Technology, Stockholm, Sweden

  29. [37]

    Université de Montréal, Montréal, Canada

  30. [38]

    Georg-August-Universität Göttingen, Göttingen, Germany

  31. [39]

    Goethe University Frankfurt, Frankfurt am Main, Germany

  32. [40]

    University of Texas at Dallas, Richardson, Texas, USA 29

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.