Pith. sign in

Paper Citation Record · LEDGER

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

As of 7 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 1 inbound Pith citation observation for arXiv:2506.10903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10903 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:18:41.054671Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T09:29:45.021656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T09:30:48.177781Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a816238d-e1ab-4c7f-adf7-5d6b99775fc0 · outbound

This paper cites Draft, sketch, and prove: Guiding formal theorem provers with informal proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Draft, sketch, and prove: Guiding formal theorem provers with informal proofs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.195737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.195737Z digest=sha256:d414bfca40fdf2990f4ba29c6d2c588558b1b4d1b504a2687736a7a0ede5ced7

Observation 137b8a50-ca7c-482a-951d-c2fb60725cc0 · outbound

This paper cites Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ren, Junxiao Song, Zhihong Shao, Wanjia Zhao, Haocheng Wang, Bo Liu, Liyue Zhang, Xuan Lu, Qiushi Du, Wenjun Gao, Haowei Zhang, Qihao Zhu, Dejian Yang, Zhibin Gou, Z.F

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.397037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:32.275985Z digest=sha256:b163e0e10836d484768417ff240bba76c75da685a0e7d0e8335ebc0ae611754e

Observation 91dfdb80-5796-4738-addf-3a8f02fb9094 · outbound

This paper cites Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.420330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.420330Z digest=sha256:c9427d4a561bd95a195933eeec8a93f0c02cd3773fe74f3330b917934d8afd22

Observation d272895d-4865-4805-93a2-6e13de144499 · outbound

This paper cites Improver: Agent-based auto- mated proof optimization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improver: Agent-based auto- mated proof optimization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:50.191854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:32.532344Z digest=sha256:2f3cc9703e0ababdf3c0c15812f7eb1d84772f378fdb1767c61d3f42636bcd05

Observation 75d131e4-b583-4ee6-a264-8d4735dd8f7a · outbound

This paper cites Learning formal mathematics from intrinsic motivation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Learning formal mathematics from intrinsic motivation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.950599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:32.672032Z digest=sha256:1fa1f90b8b8adee88644b2c00159fb01c41a1e0b235c3048ff38675a945d788e

Observation 35677c87-0245-4d30-b063-71530db8b180 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.811609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.811609Z digest=sha256:7f9edde5c177f40a9d7cdc9a515fa5d57cccb988bb078ac774d9bbbfe7e464e9

Observation ed735183-5cc2-4980-b111-503812ddb3c8 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formal Mathematical Reasoning: A New Frontier in AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:32.934134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:32.934134Z digest=sha256:75c5f1c8339bc211d6fdd104093df54af8780cd11cca629fb966b23973ffb6b3

Observation f4062c19-fdea-48f6-8940-82595d81160d · outbound

This paper cites Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Logic-LM: Empowering large language models with symbolic solvers for faithful logical reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.046527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.046527Z digest=sha256:759963c04b8d053b8a5dd092b95884c688feacdc765dff6cc45f3fcdca795822

Observation 18ad6655-73c1-46cd-81ce-1d69c186e9d8 · outbound

This paper cites Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Faithful and Robust LLM-Driven Theorem Proving for NLI Explanations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.164701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.164701Z digest=sha256:ae311b945a38e96cc7ff975d0b53c609e8501575ff88a5dfc7ada5d5dc202091

Observation dc537e8e-e477-42a3-9675-eb46c219e13b · outbound

This paper cites Dennis, and Andre Freitas.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Dennis, and Andre Freitas

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.647213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:33.300555Z digest=sha256:c40f95ab017098ce34e4ea95f0c72f8aba81ddc6bdafb2b2982cbb57d62991c2

Observation 102261bd-0257-4fae-a262-88dc703e0259 · outbound

This paper cites Autoformalization with large language models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalization with large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.409341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:33.622889Z digest=sha256:1012905fe79c809288755c742578625cfe42b49ddb41b97f8284b2a5f0dc4d95

Observation adf526b7-37f6-45d8-bf35-408ab94636e4 · outbound

This paper cites Consistent autoformalization for constructing mathematical libraries.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Consistent autoformalization for constructing mathematical libraries

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:49.157181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:33.790467Z digest=sha256:feefa6d105aed110d6cf44d31667fbdaf7a5eb86cccef537bfcdd232f407d9fa

Observation 8e6ae0de-335e-4f1c-819a-132ccf2cc1fb · outbound

This paper cites Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Formalizing complex mathematical statements with llms: A study on mathematical definitions, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.907155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:33.924912Z digest=sha256:07f4d6920e8b517dea40ede5b84724c72e8b48904f28be529a07cb5469fe8d69

Observation 001367c0-3118-4ae0-b04d-ac25c0afdc2b · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:48.600638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:34.098846Z digest=sha256:b22ee7b2ad1ee78d93777e07bb33b7b598401f4498f8190068beed106442fd01

Observation 63f2c65f-bb2f-4d83-afa7-582da3a4167a · outbound

This paper cites The lean theorem prover (system description).

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning The lean theorem prover (system description)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.397535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:34.249390Z digest=sha256:1cbc9ee8cca7b0e233eddcfbe2c9814a62918c0d9bf690d024f5f5fce55f2b47

Observation 1571152a-433d-4204-9f1e-cd191b04572b · outbound

This paper cites Gonzalez, and Ion Stoica.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Gonzalez, and Ion Stoica

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:48.145267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:34.432097Z digest=sha256:e4f721f6cc9dd0b129b20a27efcda8ebc1ccbc4b8928a493d4390c2e1a8ec896

Observation 74021821-a3e1-4a4a-bbd9-23a8a17e825f · outbound

This paper cites minif2f: a cross-system benchmark for formal olympiad-level mathematics.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning minif2f: a cross-system benchmark for formal olympiad-level mathematics

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:34.562778Z digest=sha256:e8ebb6fde6620cb305627b08ef3d6a42374847c273d9e80bfc20611305e09e34

Observation 47aac0b0-bdfe-42a6-9f1b-9f7864f60632 · outbound

This paper cites Ayers, Dragomir Radev, and Jeremy Avigad.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Ayers, Dragomir Radev, and Jeremy Avigad

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:34.919040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:34.919040Z digest=sha256:919858f5fa03b3a5dd4555a3e4c2640f095343163cb830095e38d7bd7429a6cd

Observation 37032d3a-37ac-4c07-9d44-6cc6ad0ac01f · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Bleu: a method for automatic evaluation of machine translation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.048175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.048175Z digest=sha256:b1da87c69d84f9c3915eeb4d3934daa8b43197c1094d6e7529d023b3860bba13

Observation ae2b1546-3a2c-4ffd-9afa-c804c94c7b7d · outbound

This paper cites chrF: character n-gram F-score for automatic MT evaluation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning chrF: character n-gram F-score for automatic MT evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.147923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.147923Z digest=sha256:865c5242cee2a47ebe772320301c5d2af600d6ed974cff7b88607e8286b2d866

Observation a9f4a789-209f-490d-9a2b-81fab12a788e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.273746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.273746Z digest=sha256:f339fc3c6c66f8913f0e11d08deed00ce81bae74466c9547388de041812a39e9

Observation 2dbd116f-a8d9-41e4-aba6-5a76d2fb2468 · outbound

This paper cites GPT-4 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.413820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.413820Z digest=sha256:5d63178cf9546e612e5daff5baa69a5034e86ebd718a233f85a4e83c20dad71f

Observation b52d544c-1183-4060-9c2d-f7634d36decd · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.595478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.595478Z digest=sha256:96eaf3a511489315ecb3379a018dea5848720f711827787bc3bb37131aec172f

Observation 3bea1d13-df8c-486c-bdb5-ece8dbb8bfed · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.713877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.713877Z digest=sha256:392ddce0819a41a5ff013cbfd1600ef8dc65aaf597f66517fa20aa602d14ac76

Observation 333f42d6-2093-4662-8ca7-c489eabde370 · outbound

This paper cites Qwen2.5 Technical Report.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:35.867509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:35.867509Z digest=sha256:4f346374ea1a44045ceed288cc7817f93c4ea398f70f35d99ff9f080a2dcd7b0

Observation e0644526-4e3a-4db7-b468-0c8f226d2275 · outbound

This paper cites Enhancing ethical explanations of large language models through iterative symbolic refinement.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Enhancing ethical explanations of large language models through iterative symbolic refinement

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.378180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:36.004769Z digest=sha256:60439349f7b796d2e62033270faa5c66e2e5fda827bb058970f20cf2790db40a

Observation ea25568f-b3a0-4d18-b498-36ba5a2aeae9 · outbound

This paper cites Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Daniel Raggi, Wenda Li, and Mateja Jamnik

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.811020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:36.305644Z digest=sha256:977689242bea5e454be3585629bb20aa000b998a2eca96cd33ab85e485382c05

Observation 21afc887-d59f-442c-82a1-ae126ab38535 · outbound

This paper cites Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Towards autoformalization of mathe- matics and code correctness: Experiments with elementary proofs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.491908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.491908Z digest=sha256:d056bba8b5c8aabf5be9e9dd0c399cab2f7fa14ab0f2eef9d568167f60f6aa1f

Observation beb74f61-cf34-4f7f-b3da-6085559f02c7 · outbound

This paper cites URL https://aclanthology.org/2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning URL https://aclanthology.org/2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:47.098812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:36.188422Z digest=sha256:eacdd1e46c7cfc41508dafd60c6304f80ddd0a7af685a5be5530b58a42fa41a0

Observation cd3f2379-5699-4520-b467-8e93aef866b3 · outbound

This paper cites Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Rethinking and improving autoformalization: towards a faithful metric and a dependency retrieval-based approach

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.618070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:36.813396Z digest=sha256:0aa7b9f7f0ff82f7aa96efac8463b9d8bae52307cf98dacdbb26d7bb1db60676

Observation c3cdeb5b-e500-453a-b856-ef46c5be2a5b · outbound

This paper cites Process-driven autoformalization in lean 4, 2024.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Process-driven autoformalization in lean 4, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.352503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:36.948080Z digest=sha256:0bd2b7810753d9fb593b564cf2efc9717a002cda65c3388db7a210fbbe25c558

Observation 1d0d0b32-4316-4192-aeff-b148aa2ea300 · outbound

This paper cites LeanDojo: Theorem Proving with Retrieval-Augmented Language Models.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LeanDojo: Theorem Proving with Retrieval-Augmented Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:36.682669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:36.682669Z digest=sha256:4c0ffe6c508c96da483e370f51ff38dcd31d8f691b05832ecde0c1b987e8ad00

Observation 6ff906c5-70ef-4cef-bbb5-3508baaec954 · outbound

This paper cites Jiang, Wenda Li, and Mateja Jamnik.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Jiang, Wenda Li, and Mateja Jamnik

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:46.159363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:37.243294Z digest=sha256:668069dc092bd3dce7c933626047f4408c6c6010875b2160ed5f603d46a4f095

Observation 4d3ae176-89da-43b3-b7e4-412ae99385df · outbound

This paper cites Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Atlas: Autoformalizing theorems through lifting, augmentation, and synthesis of data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.933592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:37.400350Z digest=sha256:65e95c617ae02cf54351a6afac588a1da85064639b6b29ab3275bcac54b68938

Observation 0c7db4fb-86ba-4368-8d0f-462518c4710f · outbound

This paper cites Autoformalize mathematical statements by symbolic equivalence and semantic consistency.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Autoformalize mathematical statements by symbolic equivalence and semantic consistency

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.081378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.081378Z digest=sha256:3773820d464401119fbd472a052cc8f57fd60f7f45db89d4e5c5a3e9b81c417d

Observation 1eaac6e5-82b5-4339-91d3-8e3d656f47f4 · outbound

This paper cites FormalAlign: Automated Alignment Evaluation for Autoformalization.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning FormalAlign: Automated Alignment Evaluation for Autoformalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.768435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.768435Z digest=sha256:254f317c512dcfb3674ac77802a5a777c690b376f01c20ed32b9f228fbe97bf5

Observation 6ad1e9c1-7d6d-4634-a860-c8a043088537 · outbound

This paper cites Branch-solve-merge improves large language model evaluation and generation.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Branch-solve-merge improves large language model evaluation and generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.937243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.937243Z digest=sha256:b569a7cdba1f65f8c42e69766a4c7b7768fab8732e7e328e65a702aa9aea3a94

Observation f6fc7ac3-3d55-4304-b3b2-e556f1ec9ee3 · outbound

This paper cites Self-Taught Evaluators.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Self-Taught Evaluators

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.084976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.084976Z digest=sha256:ff806f9bc4f1925636867601f3ff2113d0efc9c00c3797dd5aa1cc36e9ecaca7

Observation 60ff8893-e2a2-41f5-ae70-28be40973b6a · outbound

This paper cites Improving autoformaliza- tion using type checking, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Improving autoformaliza- tion using type checking, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.664687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.664687Z digest=sha256:4e05e6c72c0f661aa4b98eab18f8a8c443fee571b48738b406b668a919cf0919

Observation 1d1f828f-2576-4cfa-bf83-6ca7805c7c51 · outbound

This paper cites J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning J1: Incentivizing thinking in llm-as-a-judge via reinforcement learning, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.375950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.375950Z digest=sha256:3e6c5d8a169d08254080be9ceac470e06d6c29d5f465ecda2bc869dd0af8918d

Observation 15a2d426-3228-482a-9c22-0ce3cae0da50 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.572175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.572175Z digest=sha256:775bcaf638a6e24e212db1c4357f6a1e753fb552982eab5e3c2efca06cc71434

Observation dd5a74ee-294e-4314-a397-44d9d4c9fc11 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning A Survey on LLM-as-a-Judge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.698535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.698535Z digest=sha256:f44f6a69745bf7fafc891b62f39ac1ca8fd692dfb1dd1ad39472b30227f6dc80

Observation 24134fc3-59a4-4513-8ce4-0d407d1b94e2 · outbound

This paper cites Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.239873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.239873Z digest=sha256:d857d87a32c766a59a8a9c8bd4f97e5e26af4739afcbc243379ec421d1e90cd3

Observation 433988b7-2bc8-4b3d-8835-dbc4e74d7e36 · outbound

This paper cites DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.132366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.132366Z digest=sha256:18d1e81b52cc0ac02c329fb3bf1eae629f418cc9b88e30e26feff00fb2c4704c

Observation e654143c-fe28-4cd0-84c9-265a5007d2e5 · outbound

This paper cites Assessing judging bias in large reasoning models: An empirical study,.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing judging bias in large reasoning models: An empirical study,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:45.671838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:38.874413Z digest=sha256:0971abb6fc466f48b7ec047aa0e19708a24f7d19388f785156bbc658d09c62a8

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · outbound

This paper cites Assessing Judging Bias in Large Reasoning Models: An Empirical Study.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:9637e186ee41bea3ac58bd23a3a63d3ee2e73f8e3a5df3faa22170e34c212d51

Observation 4c1bd57e-0c22-4e90-a74f-5c29b3d838a0 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.443347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.261252Z digest=sha256:cd04e471e73f87de99187d74bab3a001d506d618b044c206528090b689312ab0

Observation 9a85e2f3-ae8e-401a-ba33-8a3d30b96e77 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:45.103398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.461393Z digest=sha256:af2c7b427684db8b5ffc38c261608521d7b61bb9f7f12814475839d137640fa5

Observation b7e0f182-6e81-4afe-bbf3-58e4cf7353f8 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.912106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.614636Z digest=sha256:e18289805f431045996d83acebbc7ba9b888ff4e81be04880f8a4805f8d29d04

Observation 781f4874-c6f8-4d24-9488-e7af0ab01019 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.668304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.767627Z digest=sha256:8ed409e32fa4183ffc924487bdd0a848087fb5104fde04cfa614ab4b7523e082

Observation 107593b5-e04c-4598-a95c-eeff9355911f · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.419887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.871380Z digest=sha256:9f4ab676699d622a884d664bb4c69a4b0686ddaee64b4a09901a3162a9ed2515

Observation 17c5e063-8659-4a54-88b3-01b9a2313b6c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:44.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:39.985638Z digest=sha256:67fb7b5f37ce643c9b483be47941745c34b462f25af0a39cd6d74c91f4c96d17

Observation 02ace0ff-3be6-4064-90f0-162cfda66110 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.942277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.098574Z digest=sha256:29d089be8f293f9dc740a4ffba3047b50611ece6f20f1c6e318ac72490b70bdc

Observation 6d05781f-064a-4e99-a2e0-8036bf1498e3 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.688713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.263334Z digest=sha256:45b77b4baa34822af3daa6cba275114c362ad088eb8c5f4c3bb552527967970c

Observation 261fb21b-acc0-4aff-a9c7-ddeb5b01092c · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.407700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.380673Z digest=sha256:c9695954d577a45c5d298094d63f90d973d8440739978ae8a9d5a87e7a4047dc

Observation 8118648f-d720-48bd-80e6-5e44868d4b3e · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:43.093832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.502132Z digest=sha256:0ded347c95d55e1e032fcb92b4d314c9328e815bc3a3df50340b4b7a56abf4a3

Observation f8b51ee0-2a3b-467c-b8e0-6b74b0b14180 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:42.853290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.606262Z digest=sha256:67674a33ed6789e089e2c37b096d5b056e7f4fb9375cae1d8d35e4ab1caf0479

Observation e6df49b0-4b5a-4383-b759-d4e3c3a6c432 · outbound

This paper cites 15 Purpose Content Basic You are an expert in formal language {formal_language}.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning 15 Purpose Content Basic You are an expert in formal language {formal_language}

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.549214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.724866Z digest=sha256:bae56b2663bd91db7a020dee267ba93a111689c4f855f902e7d720219d329cc3

Observation 79f0eca7-498c-42f9-a217-827f750abd28 · outbound

This paper cites True" or.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning True" or

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:42.266962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:40.886817Z digest=sha256:0aebf0a18e8b4bc4b725ded18ef241663fb780ea291d0b0fd4f65c74308ae058

Observation 5f14c5e3-8597-4858-8fe3-ed2c9e6a390f · outbound

This paper cites real ⇒ real.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning real ⇒ real

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:18:41.953448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:41.054671Z digest=sha256:5fd57f46b25a2d454db84112095142a908c50e5db95991436a31abebf5d143c2

Observation 49864cd8-4a36-4a36-a88a-5d12268d63a7 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:18:47.645468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:18:34.757295Z digest=sha256:ffc68bf6577edf668ebee1340fc5d1152cc71c4cfeea4a68d39c5b259053bd98

Observation 315b83f6-3c65-4b4b-84f5-c33ba707b30e · outbound

This paper cites doi: 10.18653/v1/2024.emnlp-main.172.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning doi: 10.18653/v1/2024.emnlp-main.172

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:33.473040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:33.473040Z digest=sha256:4adc813e120eae3f4e6e31388d95be30214f6da6340d67abfce5bbf599412c12

Observation d0d82ccf-bfb5-44be-bf8e-eb045006c155 · outbound

This paper cites an unresolved cited work.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:37.536531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:37.536531Z digest=sha256:5ddcedcdc8b47283217fe4692e17a78a15c2b799e5e194091fac2b9ee9e957b0

Pith citing papers

Observation dc4654cd-50d5-40c0-9878-6cbf9d206818 · inbound

Monotonic Reference-Free Refinement for Autoformalization cites this paper.

Monotonic Reference-Free Refinement for Autoformalization Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.179885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:29:45.021656Z digest=sha256:2e8a7c9e79163ddb01b594065a508c18fcec9991fa3d376fdd9ff2e526148de8