Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.728402Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2502.00855.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.728402Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 46adc04e-6384-4ba7-8012-b5d099c8852a · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models A Comprehensive Overview of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d09e761d-77e3-4767-b373-dd821fc42455 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models A Survey of Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196b91ae-25b8-4233-9d40-4719144d5489 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Formal Mathematical Reasoning: A New Frontier in AI
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a792a89c-122f-401b-ac54-8d560c2bc5bc · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 216c546e-e74a-47a8-ad92-f6e0602a77da · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 724b3397-60c7-492f-8a31-f7ab8e55b027 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models LEGO-Prover: Neural Theorem Proving with Growing Libraries
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa175261-3913-4a31-b455-920faa242b10 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating Mathematical Reasoning Beyond Accuracy
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a96324a1-090c-460b-9c6f-6ba5004baa86 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models The lean 4 theorem prover and programming language
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b5dd55-cb6d-4b01-8b33-2a281145c171 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Isabelle: A generic theorem prover
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4719425c-dcbb-4594-b15b-5a0e8466dbe1 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models The coq proof assistant a tutorial
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c63c7c0-c915-40f5-959a-5cab0960a528 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Leandojo: Theorem proving with retrieval-augmented language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0fda7f-b129-4043-a78a-994dc9aa6a37 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Thor: Wielding hammers to integrate language models and automated theorem provers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dfd4d5c-5d14-4ec0-b538-ab945a6f4055 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Draft, Sketch, and Prove: Guiding Formal Theorem Provers with Informal Proofs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33370f08-92ce-4669-a42f-ae96a93c6fee · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Baldur: Whole-proof generation and repair with large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b55614-0113-486b-b881-35b0610bcf86 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 401c5ced-95d8-44dd-a157-a5ad13ebd946 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343d4693-79a4-4615-94f8-8ab3181eddb1 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169–186
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a350cd1-9f03-44ba-9458-0bbdcf69b8ce · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f8e8cb-e477-46cc-b372-20167b0f48ea · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models An Empirical Evaluation of LLMs for Solving Offensive Security Challenges
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d72be25-55bf-4215-8434-89bde6f7f85f · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Cladder: Assessing causal reasoning in language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f438116-dd13-4aac-bbbf-21fd26c5fcbc · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Using llms to facilitate formal verification of rtl
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e398a153-d0d9-43ab-8532-405143c50ea3 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad8c173-7105-47c4-b0e0-307126fad0a3 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Evaluating llms’ mathematical reasoning in financial document question answering
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68778ea7-cdff-46c2-b429-b893d4500153 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Position: AI Evaluation Should Learn from How We Test Humans
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2673de03-43d8-4b24-994f-5e7c47ea5911 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models tinyBenchmarks: evaluating LLMs with fewer examples
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87a1d1c7-2a9d-4572-9231-8c09383047a7 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78fc73ac-c379-4366-8b99-5625f369895e · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Psychometrics: an introduction
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 47d6942f-f917-4de4-a709-6d24bcb45c04 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Diagnostic measurement: Theory, methods, and applications
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85cbad2c-13ad-41cd-84cb-c09283918815 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models A brief introduction to evidence-centered design
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe1e5f72-840f-423a-ab35-605ab3140e3a · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models The basics of item response theory
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad463b42-6011-4d1c-b99e-f144a7832e5c · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Item response theory for psychologists, 2004
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1896d037-6f6a-41db-9854-1a75eea3dde7 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 700bd2d3-198a-429f-a4b0-079ebb0cdfff · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71011636-bb2e-40ab-8946-2254b05067d1 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6c94b34-ba09-4f4a-b1c9-104b90a3ff31 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 081c51c7-9493-4b03-b3ea-a873f45d064d · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Item response theory
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85011a5e-a23a-4f63-b776-710aa2781705 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models A tutorial on fisher information
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a57fdb7d-a90b-4869-9f32-dcfbd4540ea1 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e48c24e9-070d-45a2-8416-70530b903c9d · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Jiang, Jia Deng, Stella Biderman, and Sean Welleck
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a75fc9-f008-44f9-b42a-4806fb871dbc · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Theoremllama: Transforming general-purpose llms into lean4 experts, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a526a880-7541-4b2e-8e37-bb6dc2793119 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Code Llama: Open Foundation Models for Code
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e68a2c-1051-4db1-8559-8c665027fa83 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Qwen2 Technical Report
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34fdb993-c71c-49ec-b219-148b407bb726 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Qwen2.5-Coder Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b1b55d0-395b-4a56-b449-b2552450d9e8 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Lisa: Language models of isabelle proofs
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a8011d-b6bd-46a0-a65c-0c78c83290c4 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc6b863d-1272-4197-a4af-98b0f0d62e0b · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ccc95c8b-89dd-4762-b411-bbb3302c9cc5 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models A.2 Theorem Categories The theorems in miniF2F are classified using multiple criteria
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1305f129-dd40-4bd0-b677-d1563d2588b8 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84e96882-466d-4eba-9803-81063e4eb9da · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models This is because, in the miniF2F design, the test set is reserved for evaluation, while the validation set may have been used during model training [32]
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbc1cf20-f683-405d-934f-a7e8adb54c62 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models This is because competition problems tend to be more complex, but their higher complexity also leads to a lower discrimination
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73b796a4-de88-45db-8199-08e445c71525 · outbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
No inbound Pith citation observations are available.