{"id":"acd08e25-cef4-44ef-a162-979988f7c383","arxiv_id":"2606.05818","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Mathematicians compiled 100 research-level questions and evaluated state-of-the-art LLMs across multiple stages, leaving only 2 unsolved.","lead":"A workshop of 49 mathematicians created a set of 100 research-level math questions with known answers and tested them on current LLMs in three stages of increasing effort. The results show only two questions remained unsolved, suggesting rapid progress in AI mathematical reasoning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"High solve rate may reflect data leakage or non-novel questions rather than reasoning on representative research-level problems","rationale":"This directly matches the reader's identified weakest assumption. The full text (per the prompt) does not appear to supply the missing leakage controls or selection criteria that would secure the claim, so the reader's UNVERDICTED stance remains appropriate pending the concrete check.","tokens_in":1635,"tokens_out":315,"duration_ms":15191,"concrete_test":"Release the full 100-question list with provenance and origin dates; run each question through a web search and against known pre-2026 training corpora; if any question or its solution appears in public sources before the evaluation window, recompute the unsolved count excluding those items.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that the 100 questions test genuine mathematical reasoning on research-level problems whose solutions are not derivable from public training data. The paper states the questions were compiled April–May 2026 with known answers and that only 2 remained unsolved after staged LLM evaluations, but provides no explicit protocol for (a) confirming questions post-date all model training cutoffs, (b) ensuring no overlap with arXiv/math.stackexchange/etc. sources, or (c) documenting that the 49 mathematicians did not inadvertently select problems whose solutions appear in LLM corpora. Without these, the drop from 41 to 2 unsolved questions is consistent with memorization or pattern completion rather than the claimed advance in reasoning.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper describes the compilation of a 100-question dataset of research-level mathematics problems with known answers by 49 mathematicians between April 1 and May 15, 2026 (primarily during a 3-day workshop in Leipzig). It reports the results of three staged LLM evaluations: a single attempt by five models (41 unsolved), a 20-run evaluation with three models (16 unsolved), and a 3-run evaluation with two heavy-thinking models (2 unsolved), concluding that LLM mathematical reasoning capabilities are becoming impressive.","tokens_in":1769,"tokens_out":507,"duration_ms":17025,"significance":"If the questions are verifiably post-training-cutoff and the evaluation protocol includes rigorous controls against contamination, the work would supply concrete evidence of rapid progress on hard, previously unsolved mathematical problems. The multi-mathematician curation and staged evaluation design are methodological strengths that could make the benchmark useful for future comparisons.","major_comments":[{"comment":"Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching.","section":"Abstract"},{"comment":"Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract could state the topical distribution or difficulty stratification of the 100 questions to allow readers to judge representativeness.","section":null},{"comment":"Clarify whether the 'known answers' were independently re-derived by the curators or taken from existing literature, and how any discrepancies were resolved.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is placed in math.HO; its empirical benchmark character may fit better in an AI or computational-mathematics venue, though that is a scope rather than a technical issue."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their detailed and constructive comments, which highlight important aspects of transparency needed for the benchmark's claims. We agree that the manuscript would benefit from additional methodological details and will revise accordingly. Our point-by-point responses follow.","responses":[{"response":"We agree the current manuscript lacks an explicit protocol. All 100 questions were newly formulated as original research-level problems by the 49 mathematicians during the 2026 workshop (post-dating all model training cutoffs referenced). We will add a subsection detailing the generation process, including that questions were created de novo without reference to existing public corpora and cross-verified for novelty by multiple participants against arXiv, MathOverflow, and MathStackExchange. This revision will support the reasoning-over-memorization interpretation.","revision_made":"yes","referee_comment":"[Abstract] Abstract (dataset compilation paragraph): the manuscript supplies no protocol for confirming that the 100 questions post-date all model training cutoffs, have no overlap with arXiv/MathOverflow/MathStackExchange content, or were selected without inadvertent inclusion of problems whose solutions appear in public LLM corpora. This information is load-bearing for the claim that the reduction from 41 to 2 unsolved questions demonstrates reasoning rather than memorization or pattern matching."},{"response":"We concur that these elements are missing from the manuscript. The revised version will include a new section on verification: each solution was confirmed by its originating mathematician and independently reviewed by at least two additional experts, with consensus required. For the multi-run stages, we will report per-run solve counts, observed variance, and the precise criteria for labeling a question 'solved' (correct final answer provided). This will enable readers to evaluate reliability.","revision_made":"yes","referee_comment":"[Abstract] Abstract (evaluation stages): no description is given of the answer-verification procedure, inter-rater reliability among the 49 mathematicians, or statistical controls (e.g., variance across the 20 runs or false-positive rates for 'solved' labels). Without these, the reported counts cannot be assessed for reliability."}],"tokens_in":1297,"tokens_out":456,"duration_ms":24466,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The new thing here is the dataset itself: 100 questions assembled by a large group of mathematicians in a 2026 workshop, with most work done in three days at the Max Planck Institute. They ran the questions through staged LLM evaluations and report only two left unsolved at the end. That specific collection and the participation numbers are the concrete addition.\n\nThe collaboration is a plus. Getting 49 people to contribute and verify questions with known answers is real work, and the staged testing (single pass, then 20 runs, then heavy models) shows they tried to push the models harder rather than stop at the first failure.\n\nThe main weakness is the missing protocol details. The abstract says the questions post-date some cutoff and have known answers, but it gives no account of how they checked for overlap with arXiv, MathOverflow, or other public sources that likely sit in training data. It also does not describe how they confirmed LLM outputs were genuine solutions rather than memorized text. Without those steps, the jump from 41 to 2 unsolved is consistent with pattern matching on familiar material.\n\nThis paper is aimed at people who build or use math benchmarks for LLMs. A reader looking for a fresh test set can pull the questions and run their own checks. The work is coherent on its own terms and shows honest effort to create something current, so it is worth sending to referees who can examine the curation and verification methods in the full text.","headline":"The Leipzig paper gives us a new set of 100 math questions compiled by 49 people, but the claim that LLMs are now impressive at research math rests on thin evidence about contamination controls.","tokens_in":2490,"tokens_out":378,"would_cite":false,"duration_ms":9856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A dataset of 100 research-level math questions with known answers leaves only two unsolved after staged LLM evaluations.","keywords":["LLM mathematical reasoning","research-level math benchmarks","AI evaluation","mathematics datasets","large language models","problem solving"],"falsifier":"Demonstration that a substantial fraction of the 100 questions appeared in the training corpora of the evaluated LLMs or that the questions do not match the difficulty of typical unsolved research problems.","tokens_in":2537,"feed_emoji":"🧮","tokens_out":570,"duration_ms":13224,"temperature":0.7,"pith_summary":"The paper assembles 100 questions drawn from research mathematics, each with a verified answer, through a collaborative effort by 49 mathematicians. These questions undergo testing first with five leading LLMs in single attempts, then with three models across 20 runs each, and finally with two advanced models in three runs. The process reduces the number of unsolved questions from 41 after the first stage to 16 after the second and to 2 after the third. The authors present this outcome as direct evidence that current LLMs exhibit strong mathematical reasoning on problems at the level of recent or ongoing research.","feed_headline":"LLMs solve 98 of 100 research math questions","feed_subtitle":"New collection of 100 questions with known answers shows top models succeeding after repeated evaluation stages.","key_machinery":"The 100-question dataset itself, compiled during a workshop and evaluated in three successive stages of LLM interaction with answer verification.","core_discovery":"The central claim is that a newly compiled collection of 100 research-level mathematics questions, assembled with known answers, can be solved by state-of-the-art LLMs in all but two cases after multiple evaluation stages that include repeated sampling and verification.","pith_inferences":["The benchmark could serve as a continuing yardstick to measure progress in LLM mathematical capability over time.","Success on these questions may indicate readiness for LLMs to contribute to open mathematical research problems.","The staged evaluation method highlights the value of multiple attempts and verification in assessing reasoning rather than one-shot performance."],"forward_implications":["LLMs can now address mathematical problems that previously required specialized human expertise.","Repeated sampling and verification protocols allow models to reach solutions on a large majority of such questions.","The remaining unsolved questions provide a focused target for further model improvement.","Benchmarks of this type will require regular updates as model performance advances."],"fun_headline_variants":["LLMs solve 98 of 100 math questions","98 research math questions solved by LLMs","Only two math questions left unsolved by LLMs","LLMs handle 98 of 100 new math questions"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The questions are representative of genuine research-level mathematics and have not appeared in the training data of the tested models.","fun_headline_variants_meta":{"raw":{"variants":["LLMs solve 98 of 100 math questions","98 research math questions solved by LLMs","Only two math questions left unsolved by LLMs","LLMs handle 98 of 100 new math questions"]},"model":"grok-4.3","cost_usd":0.004715,"raw_usage":{"total_tokens":2274,"prompt_tokens":561,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":47149500,"prompt_tokens_details":{"text_tokens":561,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1654,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":561,"tokens_out":59,"duration_ms":12272,"temperature":1.0,"reasoning_tokens":1654,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T22:48:27.712995+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Demonstration that a substantial fraction of the 100 questions appeared in the training corpora of the evaluated LLMs or that the questions do not match the difficulty of typical unsolved research problems.","supporting_citations":[],"review_version":1}