Pith. sign in

Paper Citation Record · LEDGER

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

As of 21 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 6 inbound Pith citation observations for arXiv:2602.24173.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.24173 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:05:57.763677Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:18:57.472922Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b4a44422-80b4-49ec-9128-de16d7991e5a · outbound

This paper cites First proof.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics First proof

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.148361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.148361Z digest=sha256:f1674a1c8e1996a434c31a1a64df280f2eb31e4841e313f06b85f502839b4c1b

Observation e302ccb0-f7e8-4004-967b-440bdb69e627 · outbound

This paper cites Fel’s conjecture on syzygies of numerical semigroups.arXiv preprint arXiv:2602.03716,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Fel’s conjecture on syzygies of numerical semigroups.arXiv preprint arXiv:2602.03716,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.615415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.615415Z digest=sha256:dd91fa25a47fc2bbf94af6e395ca452f7bbc76003bfc68dd7e5041fd48c3f7d4

Observation 8ae5b453-b8a4-4110-bdcf-2ee28a0d2051 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.728815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.728815Z digest=sha256:fa04966e5256afbdf77ea194e1aecbbfe06810bbee6831651ad7b54d45811a0e

Observation 5cb223fe-f231-4995-872c-325d810d418e · outbound

This paper cites Mathematical research with GPT-5: a Malliavin-Stein experiment.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Mathematical research with GPT-5: a Malliavin-Stein experiment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.919446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.919446Z digest=sha256:22b5768c0e96251b4d39e43f02998361e52d1f7712c2aa48a119eb73cd976f3e

Observation 96b6abe5-5624-4b03-9b09-ca2fc3f0c8e9 · outbound

This paper cites Project aletheia: Verifier-guided distillation of back- tracking for small language models.arXiv preprint arXiv:2601.14290,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Project aletheia: Verifier-guided distillation of back- tracking for small language models.arXiv preprint arXiv:2601.14290,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.031583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.031583Z digest=sha256:d750d13f89ed539be1d478cf8069af41c9853bdb82c696754a7e3da15134d6f5

Observation 2cc79098-c4a3-4ac1-86f3-63d55d613da1 · outbound

This paper cites Semi-autonomous mathematics discovery with gemini: A case study on the erd\h {o} s problems.arXiv preprint arXiv:2601.22401,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Semi-autonomous mathematics discovery with gemini: A case study on the erd\h {o} s problems.arXiv preprint arXiv:2601.22401,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.105644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.105644Z digest=sha256:f4424acc5cb8c634ef59bc8165aa7accb6ff6851dc4cc524c7532722b2f3b5a5

Observation 8dd39d68-61f9-42c5-a67e-6b7a1480586d · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.217637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.217637Z digest=sha256:bd26811c2c49dab1b0a21dfeab09c02fdbaabc0d15a0dd17b62b89fb7ad1ae31

Observation 32b11f53-8652-454d-89ae-f51e74600799 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.402050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.402050Z digest=sha256:21cf66cbda6c219807c051ecbee00ce84e88a2b124b54a49f2d24e575588c94b

Observation 4df45a2d-1699-4b58-90ae-65099674c7b2 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.471045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.471045Z digest=sha256:80ca730da8230c8dbe7528091e622873e9e4cff0ae6c6f0cf2bdafcb8b902173

Observation 463fff06-97bf-4fc4-8513-52619fd05d02 · outbound

This paper cites Gemini 2.5 pro capable of winning gold at imo 2025.arXiv preprint arXiv:2507.15855,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Gemini 2.5 pro capable of winning gold at imo 2025.arXiv preprint arXiv:2507.15855,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.544138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.544138Z digest=sha256:89694ecc4e7c01f35e5f2daf43459cadbcec9c8b0b6cf433bf70ac268204844e

Observation 9e9d4597-38d2-4e48-b861-593460e336d3 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.621707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.621707Z digest=sha256:582e33f55d5b4a85d9fc08da0fc91ac0fdf939b9ee7a1263ce0f046b5d150066

Observation 8222b21f-bb37-4562-80e6-bb10c20254ca · outbound

This paper cites FIMO: A Challenge Formal Dataset for Automated Theorem Proving.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics FIMO: A Challenge Formal Dataset for Automated Theorem Proving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.734724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.734724Z digest=sha256:d0927281ffbebb35d8e3fb4140a73d192c86af1a8599edd9e30e52fb2af8533e

Observation d8b06253-5a78-4153-b850-e996b538acd9 · outbound

This paper cites EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics EternalMath: A Living Benchmark of Frontier Mathematics that Evolves with Human Discovery

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.850393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.850393Z digest=sha256:6d22bb2a32835a5c14474bbe665ccd202833205da16d4cf18f0d160418f79fc6

Observation 8c299cee-ab77-4d19-ba3b-1aeaa3a436e9 · outbound

This paper cites A New Approach Towards Autoformalization.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics A New Approach Towards Autoformalization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.925031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.925031Z digest=sha256:557ea5f9e25ca52c0b268cbd4bab898cb7b845f0427b046005e26fe85a2d2534

Observation 44806c9e-e9ac-4a7b-916d-e5f39cb499c4 · outbound

This paper cites Magistral.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Magistral

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.999962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.999962Z digest=sha256:0660beb60cda89a0801f1036ddde13537c4fba7e0e84cfcfe07c060c781d0280

Observation 420d03a2-9d74-441d-a403-17f7909c3de9 · outbound

This paper cites Extremal descendant integrals on moduli spaces of curves: An inequality discov- ered and proved in collaboration with ai.arXiv preprint arXiv:2512.14575,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Extremal descendant integrals on moduli spaces of curves: An inequality discov- ered and proved in collaboration with ai.arXiv preprint arXiv:2512.14575,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.082986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.082986Z digest=sha256:33bbab6dddc33a3a3b96f771cd8aeb5a8135eeff0311aad8f71b87f3cdb0b3f9

Observation 0e6a5a56-523d-4b27-ab0d-e3d05e2dcf9b · outbound

This paper cites Resolution of erd \h {o} s problem# 728: a writeup of aristotle’s lean proof.arXiv preprint arXiv:2601.07421,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Resolution of erd \h {o} s problem# 728: a writeup of aristotle’s lean proof.arXiv preprint arXiv:2601.07421,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.153967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.153967Z digest=sha256:54aacd8d8fae6d0dbc8454afb6d916a1f4ce3566fb2fbe6913c678e6a2aa31ca

Observation d0843425-eb41-4a5c-8872-974e20d21a22 · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.230480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.230480Z digest=sha256:84a7f148affde429c9ba26c8faeb067645f533a67459b60783a4b7623cf25166

Observation 2cb2768c-e514-4e27-9729-23ed8f151c19 · outbound

This paper cites Aristotle: IMO-level Automated Theorem Proving.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Aristotle: IMO-level Automated Theorem Proving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.347754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.347754Z digest=sha256:4c5a892b69a900545376d1bb95e40524c20c8d9b67b5735de4b05b141441da63

Observation fe28aa3a-fe9e-4d03-8dbe-4ebbf99dc756 · outbound

This paper cites Uniqueness of the canonical reciprocal cost.arXiv preprint arXiv:2602.05753,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Uniqueness of the canonical reciprocal cost.arXiv preprint arXiv:2602.05753,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.389635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.389635Z digest=sha256:25fa97b9de6f81996f5dfe070c57f526b56c5e6f8e1a106645bbe7baac594752

Observation d9ccfd0b-1ad1-45a5-8e21-17c30ff436fd · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Benchmarking Benchmark Leakage in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.424441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.424441Z digest=sha256:e914e3efaade02a169d16ddd19e7e4b728825213df767df1c8cd5a89dd18dba7

Observation 5eac8e4e-5ca0-47f2-8189-2041801c181c · outbound

This paper cites Realmath: A continuous benchmark for evaluating language models on research-level mathematics.arXiv preprint arXiv:2505.12575,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Realmath: A continuous benchmark for evaluating language models on research-level mathematics.arXiv preprint arXiv:2505.12575,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.481014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.481014Z digest=sha256:af92dafe6b52a40e2d011386e2f779cd3eb8c2742582ae9560c52934062f05f3

Observation cb11242e-4088-4b52-95ec-85022a328c70 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.525719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.525719Z digest=sha256:4b0599248f6ecf600b68a587d209780fd397074416bc0e460359076938a43ec7

Observation 62f2f82f-4e04-4cea-ab55-ced4f98ea4b1 · outbound

This paper cites In natural language, one can cite for instance GSM8k Cobbe et al.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics In natural language, one can cite for instance GSM8k Cobbe et al

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.581966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.581966Z digest=sha256:0d9b8e4569885ae16899df9caa04e7df441dca6c6e46d3fc5898fa44d055c5c5

Observation 273c339d-f90d-4ed3-804a-2eb088a8f409 · outbound

This paper cites These benchmarks were often used to evaluate LLM capa- bilities to the point of being the staple of evaluation for mathematical abilities.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics These benchmarks were often used to evaluate LLM capa- bilities to the point of being the staple of evaluation for mathematical abilities

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.652609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.652609Z digest=sha256:63220d3829e24020629696c0e25266f78b02a480b4937638b5b33e62351c2d12

Observation cf8406b6-1513-4047-8231-f648600ab3d5 · outbound

This paper cites A very large dataset of competition problems was created by the Numina project LI et al.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics A very large dataset of competition problems was created by the Numina project LI et al

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.696600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.696600Z digest=sha256:a43dcf53271e8bfa85b93867dd5daa6d71a9fa0dd3aebbcd1e923801fc4d8fdd

Observation 6b16f78f-ae04-4493-a4d2-ee2ff5932a2f · outbound

This paper cites an unresolved cited work.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.740526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.740526Z digest=sha256:48f7002d62ec7ae8fa141777190ddbd9d6696652c531027d84ed15a5894fcd37

Observation bfb152b7-01a2-4982-a54b-3e17cb0c6bb4 · outbound

This paper cites It is worth mentioning that other benchmarks exist in formal language, also focusing on middle school to undergraduate level problems such as miniF2F Zheng et al.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics It is worth mentioning that other benchmarks exist in formal language, also focusing on middle school to undergraduate level problems such as miniF2F Zheng et al

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:57.763677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:57.763677Z digest=sha256:ceee3ec6d2d9de184c622de995c2c392bc1b9a1db64c45b91980936a6d9a2a6a

Observation 9c804fa9-6c49-48b6-acbf-b8220346b0f3 · outbound

This paper cites Objective-Function Free Multi-Objective Optimization: Rate of Convergence and Performance of an Adagrad-like algorithm.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Objective-Function Free Multi-Objective Optimization: Rate of Convergence and Performance of an Adagrad-like algorithm

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.775362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.775362Z digest=sha256:451ed1cf8ae66711535a6f38734e605506773fa00669326b8dfe4eb9cbe7541f

Observation b4d33a85-59b1-4d0d-8c67-c684348360f9 · outbound

This paper cites Numina-lean-agent: An open and general agentic reasoning system for formal mathematics.arXiv preprint arXiv:2601.14027,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Numina-lean-agent: An open and general agentic reasoning system for formal mathematics.arXiv preprint arXiv:2601.14027,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:56.808051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:56.808051Z digest=sha256:94b9bd06cc147128d6f923448b5a6353e02783c9dd8e02b4f261928dcab9e3ac

Observation f77057cf-3be0-4168-88c8-5e257064ae1c · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.301028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.301028Z digest=sha256:3fcfd263ab83b24f977d6006ddb2d91f0b9f5291fffca46ba0072499e6b858df

Observation 48f178d5-08d6-4a96-8228-bbc78a47f290 · outbound

This paper cites PatternBoost: Constructions in Mathematics with a Little Help from AI.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics PatternBoost: Constructions in Mathematics with a Little Help from AI

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.466316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.466316Z digest=sha256:75bec2fe3e57ff71ffbdb3ff68d568b0008c2ac279315c1585faa7f790336f9f

Observation 4fc9ff95-a058-4bab-8b31-82b8715a7c66 · outbound

This paper cites Semantic search over 9 million mathematical theorems.arXiv preprint arXiv:2602.05216,.

LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics Semantic search over 9 million mathematical theorems.arXiv preprint arXiv:2602.05216,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T20:05:55.234189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:05:55.234189Z digest=sha256:2d2648548886d396bb58717223f4ed8bbb7e89d102abca08cce0d07438ba6d06

Pith citing papers

Observation 6390744e-bf5f-4833-b8f9-ae9a27e484e0 · inbound

Matlas: A Semantic Search Engine for Mathematics cites this paper.

Matlas: A Semantic Search Engine for Mathematics LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-23T04:13:44.884775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:32:30.216388Z digest=sha256:db71f403580ecb511996de81a965f34bfb084771a75c92fd27edda887a9c24bd

Observation 885b33b1-fbbf-463d-8208-520b4b720156 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-23T04:13:44.884775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:46:50.177357Z digest=sha256:aae42b256dbb3b49efa41c7220ed04884cb0a3e6b092616e3ea9c2f4d288a727

Observation bfbab266-10e3-4576-9ce2-eaa731f6d5b0 · inbound

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness cites this paper.

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-30T22:45:06.686133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:38:26.111517Z digest=sha256:4c0d1a9c193c93ce847069079c1697ed12a201059ab89daae43b1595af7f4084

Observation 2bdf5c23-3566-46c9-9948-3e184d546ea8 · inbound

Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory cites this paper.

Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T18:16:48.945820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:16:48.945820Z digest=sha256:b06bea85d79038555742826db01c5601845732ec6eede3d9a644a59016d5b8ac

Observation 09bfc06c-4615-4380-9c3b-14c7f4713fad · inbound

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability cites this paper.

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:30:26.831491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:30:26.831491Z digest=sha256:e3f12e86c1810c22e1bfd651dbec986af179508e97e4b8619535c0dac3c8e0f8

Observation 5bfab0a4-9441-483f-8d3e-91a279737295 · inbound

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability cites this paper.

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability LemmaBench: A Live, Research-Level Benchmark to Evaluate LLM Capabilities in Mathematics

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:18:57.472922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:18:57.472922Z digest=sha256:60be73fb5a167912d8ce815c870646c8b440b5c5166af84a036175b3dee8c1d8