Pith. sign in

Paper Citation Record · LEDGER

Calibrating LLM-Based Evaluator

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.13308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.13308 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:05:54.291216Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8c32b91-489c-43ac-b4fa-8f140d97ad27 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Calibrating LLM-Based Evaluator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.204544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:127eaf8d1655240c3e380813e646d2fbc33475495995894d892af9a31556f37e

Observation 2784833e-994f-4b3d-b42f-4c80756c79fb · inbound

Towards unearthing neglected climate innovations from scientific literature using Large Language Models cites this paper.

Towards unearthing neglected climate innovations from scientific literature using Large Language Models Calibrating LLM-Based Evaluator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:05:54.291216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:05:54.291216Z digest=sha256:ffa4c95863f8798b51d15a85c6abe4f4715310836665ba7af9015bce89456c3d

Observation a57f28fd-7676-4364-9c5f-b975c931315e · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:877098cd0dc28eec03f60f75dd4a557ad0dd24b041ec85caec3e3b3de549d8d2

Observation ff2cbbb3-7af3-4f8b-927d-6c210e994160 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Calibrating LLM-Based Evaluator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:44.970534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:44.970534Z digest=sha256:1fc5f24447dd4bbc5e6f6a574f3803d2bdf22f8027717d6f1ba13a5ab45f81f0

Observation 4ecfa734-88d1-454a-9f7f-1d66f2b609a9 · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions Calibrating LLM-Based Evaluator

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:54.919951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:54.919951Z digest=sha256:18cb8d567cac4293587c029a160472db2e511843313311156bba838e70ec5452

Observation d7b9ddb2-75e4-42d3-903a-7ffacf660847 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Calibrating LLM-Based Evaluator

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.331051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:05bb80cb4b22d58a57fa6075a2a238ca4fbdb0747283773167535f77c2e34890

Observation 90b2acb0-4145-43dc-98de-32ca74b6b5e4 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations Calibrating LLM-Based Evaluator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.260875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.260875Z digest=sha256:967d42a7a64bba7eaf9628af80e588128c9181fa117f05de715a18d901b768a7

Observation 8a8ceceb-ed1b-4e88-874f-d08af2c85220 · inbound

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents cites this paper.

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents Calibrating LLM-Based Evaluator

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.894987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T09:36:50.323723Z digest=sha256:eee7a86d07c234eb2b7129d1a149f93a5d566cb5b918567d2d1e9d065b9e9476

Observation 5d5550e5-0e70-4a49-92f3-1286aa587f9c · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Calibrating LLM-Based Evaluator

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:55.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:55.567886Z digest=sha256:4d1a0c73c408caa655bdb1dbf95bd0a7e76d6572229ffa2a91385019e4c43ce5

Observation 3e8a5565-b587-45dd-90df-fa8a37b56243 · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Calibrating LLM-Based Evaluator

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.954717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:4b0fa77f217867cadfdcc4ce3c47c27b83daa2f935e2cf1d5217031cc46fcb63

Observation 3ae689bf-abfe-4479-937f-c253e9ab8569 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Calibrating LLM-Based Evaluator

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.496695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:427730aa640285e9d35e6892e9a6ba8ecdbf6c68345c71a9186f46d9b667ae17