Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:05:57.763677Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 6 inbound Pith citation observations for arXiv:2602.24173.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T20:05:57.763677Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:57.472922Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
33 of 33 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation b4a44422-80b4-49ec-9128-de16d7991e5a · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics First proof
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e302ccb0-f7e8-4004-967b-440bdb69e627 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Fel’s conjecture on syzygies of numerical semigroups.arXiv preprint arXiv:2602.03716,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae5b453-b8a4-4110-bdcf-2ee28a0d2051 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cb223fe-f231-4995-872c-325d810d418e · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Mathematical research with GPT-5: a Malliavin-Stein experiment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96b6abe5-5624-4b03-9b09-ca2fc3f0c8e9 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Project aletheia: Verifier-guided distillation of back- tracking for small language models.arXiv preprint arXiv:2601.14290,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc79098-c4a3-4ac1-86f3-63d55d613da1 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Semi-autonomous mathematics discovery with gemini: A case study on the erd\h {o} s problems.arXiv preprint arXiv:2601.22401,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd39d68-61f9-42c5-a67e-6b7a1480586d · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b11f53-8652-454d-89ae-f51e74600799 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4df45a2d-1699-4b58-90ae-65099674c7b2 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Measuring Mathematical Problem Solving With the MATH Dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 463fff06-97bf-4fc4-8513-52619fd05d02 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Gemini 2.5 pro capable of winning gold at imo 2025.arXiv preprint arXiv:2507.15855,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9d4597-38d2-4e48-b861-593460e336d3 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8222b21f-bb37-4562-80e6-bb10c20254ca · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FIMO: A Challenge Formal Dataset for Automated Theorem Proving
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b06253-5a78-4153-b850-e996b538acd9 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c299cee-ab77-4d19-ba3b-1aeaa3a436e9 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics A New Approach Towards Autoformalization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44806c9e-e9ac-4a7b-916d-e5f39cb499c4 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Magistral
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 420d03a2-9d74-441d-a403-17f7909c3de9 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Extremal descendant integrals on moduli spaces of curves: An inequality discov- ered and proved in collaboration with ai.arXiv preprint arXiv:2512.14575,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6a5a56-523d-4b27-ab0d-e3d05e2dcf9b · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Resolution of erd \h {o} s problem# 728: a writeup of aristotle’s lean proof.arXiv preprint arXiv:2601.07421,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0843425-eb41-4a5c-8872-974e20d21a22 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cb2768c-e514-4e27-9729-23ed8f151c19 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Aristotle: IMO-level Automated Theorem Proving
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe28aa3a-fe9e-4d03-8dbe-4ebbf99dc756 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Uniqueness of the canonical reciprocal cost.arXiv preprint arXiv:2602.05753,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ccfd0b-1ad1-45a5-8e21-17c30ff436fd · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Benchmarking Benchmark Leakage in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eac8e4e-5ca0-47f2-8189-2041801c181c · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Realmath: A continuous benchmark for evaluating language models on research-level mathematics.arXiv preprint arXiv:2505.12575,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb11242e-4088-4b52-95ec-85022a328c70 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f2f82f-4e04-4cea-ab55-ced4f98ea4b1 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics In natural language, one can cite for instance GSM8k Cobbe et al
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 273c339d-f90d-4ed3-804a-2eb088a8f409 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics These benchmarks were often used to evaluate LLM capa- bilities to the point of being the staple of evaluation for mathematical abilities
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8406b6-1513-4047-8231-f648600ab3d5 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics A very large dataset of competition problems was created by the Numina project LI et al
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b16f78f-ae04-4493-a4d2-ee2ff5932a2f · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb152b7-01a2-4982-a54b-3e17cb0c6bb4 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics It is worth mentioning that other benchmarks exist in formal language, also focusing on middle school to undergraduate level problems such as miniF2F Zheng et al
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c804fa9-6c49-48b6-acbf-b8220346b0f3 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Objective-Function Free Multi-Objective Optimization: Rate of Convergence and Performance of an Adagrad-like algorithm
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d33a85-59b1-4d0d-8c67-c684348360f9 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Numina-lean-agent: An open and general agentic reasoning system for formal mathematics.arXiv preprint arXiv:2601.14027,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77057cf-3be0-4168-88c8-5e257064ae1c · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f178d5-08d6-4a96-8228-bbc78a47f290 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics PatternBoost: Constructions in Mathematics with a Little Help from AI
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc9ff95-a058-4bab-8b31-82b8715a7c66 · outbound
LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Semantic search over 9 million mathematical theorems.arXiv preprint arXiv:2602.05216,
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6390744e-bf5f-4833-b8f9-ae9a27e484e0 · inbound
Matlas: A Semantic Search Engine for Mathematics LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 885b33b1-fbbf-463d-8208-520b4b720156 · inbound
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bfbab266-10e3-4576-9ce2-eaa731f6d5b0 · inbound
Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2bdf5c23-3566-46c9-9948-3e184d546ea8 · inbound
Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09bfc06c-4615-4380-9c3b-14c7f4713fad · inbound
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bfab0a4-9441-483f-8d3e-91a279737295 · inbound
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.