REVIEW 2 major objections 2 minor 40 references
Benchmarks in Leipzig
T0 review · 2 major / 2 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read A dataset of 100 research-level math questions with known answers leaves only two unsolved after staged LLM evaluations.
desk verdict The Leipzig paper gives us a new set of 100 math questions compiled by 49 people, but the claim that LLMs are now impressive at research math rests on thin evidence about contamination controls. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The 100-question dataset itself, compiled during a workshop and evaluated in three successive stages of LLM interaction with answer verification.
What would settle it
Demonstration that a substantial fraction of the 100 questions appeared in the training corpora of the evaluated LLMs or that the questions do not match the difficulty of typical unsolved research problems.
Extended reading notes
Core claim
The central claim is that a newly compiled collection of 100 research-level mathematics questions, assembled with known answers, can be solved by state-of-the-art LLMs in all but two cases after multiple evaluation stages that include repeated sampling and verification.
Load-bearing premise
The questions are representative of genuine research-level mathematics and have not appeared in the training data of the tested models.
Editorial extensions
If this is right
- LLMs can now address mathematical problems that previously required specialized human expertise.
- Repeated sampling and verification protocols allow models to reach solutions on a large majority of such questions.
- The remaining unsolved questions provide a focused target for further model improvement.
- Benchmarks of this type will require regular updates as model performance advances.
Reading between the lines
- The benchmark could serve as a continuing yardstick to measure progress in LLM mathematical capability over time.
- Success on these questions may indicate readiness for LLMs to contribute to open mathematical research problems.
- The staged evaluation method highlights the value of multiple attempts and verification in assessing reasoning rather than one-shot performance.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the compilation of a 100-question dataset of research-level mathematics problems with known answers by 49 mathematicians between April 1 and May 15, 2026 (primarily during a 3-day workshop in Leipzig). It reports the results of three staged LLM evaluations: a single attempt by five models (41 unsolved), a 20-run evaluation with three models (16 unsolved), and a 3-run evaluation with two heavy-thinking models (2 unsolved), concluding that LLM mathematical reasoning capabilities are becoming impressive.
Significance. If the questions are verifiably post-training-cutoff and the evaluation protocol includes rigorous controls against contamination, the work would supply concrete evidence of rapid progress on hard, previously unsolved mathematical problems. The multi-mathematician curation and staged evaluation design are methodological strengths that could make the benchmark useful for future comparisons.
major comments (2)
- [Abstract] Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching.
- [Abstract] Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability.
minor comments (2)
- The abstract could state the topical distribution or difficulty stratification of the 100 questions to allow readers to judge representativeness.
- Clarify whether the 'known answers' were independently re-derived by the curators or taken from existing literature, and how any discrepancies were resolved.
Simulated Author's Rebuttal
We thank the referee for their detailed and constructive comments, which highlight important aspects of transparency needed for the benchmark's claims. We agree that the manuscript would benefit from additional methodological details and will revise accordingly. Our point-by-point responses follow.
read point-by-point responses
-
Referee: [Abstract] Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching.
Authors: We agree the current manuscript lacks an explicit protocol. All 100 questions were newly formulated as original research-level problems by the 49 mathematicians during the 2026 workshop (post-dating all model training cutoffs referenced). We will add a subsection detailing the generation process, including that questions were created de novo without reference to existing public corpora and cross-verified for novelty by multiple participants against arXiv, MathOverflow, and MathStackExchange. This revision will support the reasoning-over-memorization interpretation. revision: yes
-
Referee: [Abstract] Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability.
Authors: We concur that these elements are missing from the manuscript. The revised version will include a new section on verification: each solution was confirmed by its originating mathematician and independently reviewed by at least two additional experts, with consensus required. For the multi-run stages, we will report per-run solve counts, observed variance, and the precise criteria for labeling a question 'solved' (correct final answer provided). This will enable readers to evaluate reliability. revision: yes
Circularity Check
No circularity: empirical benchmark report with no derivations or self-referential reductions
full rationale
The paper compiles and evaluates a dataset of 100 research-level math questions against LLMs across three stages, reporting unsolved counts dropping from 41 to 2. No equations, fitted parameters, predictions, or derivations are present. The central claim is an empirical observation from external model runs on a newly compiled dataset; it does not reduce to any input by construction, self-citation, or renaming. No load-bearing steps match any of the enumerated circularity patterns.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Benchmarks in Leipzig." pith.science (2026). https://pith.science/paper/7UY2KIUM
@misc{pith2026260605818,
author = {Pith},
title = {Pith review of: Benchmarks in Leipzig},
year = {2026},
howpublished = {\url{https://pith.science/paper/7UY2KIUM}},
note = {Machine review of arXiv:2606.05818}
}
read the original abstract
Between April 1 and May 15, 2026, a group of 49 mathematicians compiled a dataset of research-level mathematics questions with known answers. Most of the work was done during the 3-day workshop *Benchmarks in Leipzig* with 35 participants at the Max Planck Institute for Mathematics in the Sciences in Leipzig, Germany. We present the resulting collection of 100 questions. We evaluated these questions in three stages: a single attempt by five state-of-the-art LLMs, followed by a 20-runs-per-model evaluation with three of these models, and finally a 3-run attempt with two heavy-thinking models. After Stage 1, 41 questions remained completely unsolved; after Stage 2, this count dropped to 16; and we concluded Stage 3 with only 2 unsolved questions. This demonstrates that the mathematical reasoning capabilities of LLMs are becoming impressive.
Reference graph
Works this paper leans on
-
[1]
Ifn= 1, thens(1) = 1
-
[2]
This operationsis often called the stack sorting
Ifp=LnR, that isLis the substring on the left of the entryn, thens(p) =s(L)s(R)n. This operationsis often called the stack sorting. LetB n,k be the set of permutationspof lengthnfor which all of the following hold
-
[3]
The permutations(p)avoids the patterns2314and3124
-
[4]
The number of descents ofpisk−1. Ablock moveis an operation that turns the permutation p=p 1p2 · · ·p i pi+1 · · ·p j pj+1 · · ·p k pk+1 · · ·p ℓ · · ·p n into the permutation p′ =p 1p2 · · ·p i pk+1 · · ·p ℓ pj+1 · · ·p k pi+1 · · ·p j pℓ+1 · · ·p n. That is, a block move interchanges the string pi+1 · · ·p j and the string pk+1 · · ·p ℓ, for some number...
-
[5]
bothpandπare of length30, and 2.p∈B 30,15, and 3.π∈D p(10), that is,πis at distance10fromp? Question 013solved in Stage 1 Fix two positive integers r and m. Consider the family of subsets Ci ={2i+ 1, ...,2i+r} where i ranges from 0 to m−1 . Let M be the matroid of rank r on n= 2(m−1) +r elements whose non-bases are exactly the m subsets C0, ..., Cm−1. Usi...
-
[6]
What is the largestksuch thatX k is normal? Question 040solved in Stage 1 Fix seven general points in the real projective plane
For a natural number k, let Xk be the projective toric variety parametrized by the degree k monomials in Z (insideP(Z k)). What is the largestksuch thatX k is normal? Question 040solved in Stage 1 Fix seven general points in the real projective plane. Consider the 21 lines spanned by any two points, the 21 conics spanned by any five points, and the 7 cubi...
-
[7]
and let T= (C ∗)3 15 Benchmarks in Leipzig act on X via the induced action of the multiplication in the domain of this parametrization. A subdivision of a fan consists of picking a maximal cone of the fan and a hyperplane and replacing the cone by the two cones obtained by intersecting the initial cone with the two half-spaces of the hyperplane. What is t...
-
[8]
The minimal resolution ofX Γ has exceptional intersection graphE 8
Show all 40 references
-
[9]
Equivalently, C[u, v]Γ has primitive homogeneous generators of degrees 12,20,30 , satisfying one weighted- homogeneous relation of degree60
-
[10]
The degree-12primitive invariant is normalized as follows. It lies in the one-parameter family Fa(u, v) =u 11v+a u 6v6 −uv 11, anda∈Cis chosen so that, if Ha = Hess(Fa) = det ∂2Fa ∂u2 ∂2Fa ∂u∂v ∂2Fa ∂v∂u ∂2Fa ∂v2 , and Ta = Jac(Fa, Ha) = det ∂Fa ∂u ∂Fa ∂v ∂Ha ∂u...
-
[11]
the hidden groupΓ, up to conjugacy inSL 2(C)
-
[12]
the admissible value or values ofa
-
[13]
the exact Fisher informationI= R P1 ρ2 dσ
-
[14]
the exact threshold constantC= 2√ I
-
[15]
checksum
the limiting log-Bayes factor, in favor of the hidden quotient-substrate sky over the continuum sky, assuming equal prior odds. Question 047solved in Stage 1 Let X be the variety in (C∗)5 cut out by x3x4 +x 1 −1, x 5 +x 2x3 −1, x 4x5 +x 2x3x4 +x 1x2 −1 . Consider the form Ω = ...
2001
-
[16]
Ruhr University Bochum, Bochum, Germany
-
[17]
Max Planck Institute for Mathematics in the Sciences, Leipzig, Germany
-
[18]
TU Berlin, Berlin, Germany
-
[19]
University of Florida, Gainesville, FL, USA
-
[20]
Ecole Normale Superieure (ENS), PSL University, Paris, France
-
[21]
University of California, Davis, USA
-
[22]
University of Barcelona, Barcelona, Spain
-
[23]
University of California, Berkeley, USA
-
[24]
ETH Zurich, Zurich, Switzerland
-
[25]
Universidad de Talca, Talca, Chile
-
[26]
TU Dresden, Dresden, Germany
-
[27]
University of Bonn, Bonn, Germany
-
[28]
Université du Québec à Montréal (UQAM), Montreal, Canada
-
[29]
University of Southern California, Los Angeles, USA
-
[30]
University of Luxembourg, Luxembourg
-
[31]
Amherst College, Amherst, Massachusetts, USA
-
[32]
Bielefeld University, Bielefeld, Germany
-
[33]
Technische Universität Braunschweig, Braunschweig, Germany
-
[34]
University of Leipzig, Leipzig, Germany
-
[35]
University of Copenhagen, Copenhagen, Denmark
-
[36]
KTH Royal Institute of Technology, Stockholm, Sweden
-
[37]
Université de Montréal, Montréal, Canada
-
[38]
Georg-August-Universität Göttingen, Göttingen, Germany
-
[39]
Goethe University Frankfurt, Frankfurt am Main, Germany
-
[40]
University of Texas at Dallas, Richardson, Texas, USA 29
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.