Pith. sign in

Paper Citation Record · LEDGER

Calibrating LLM-Based Evaluator

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2309.13308.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.13308 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:05:54.291216Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8c32b91-489c-43ac-b4fa-8f140d97ad27 · inbound

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents cites this paper.

AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents Calibrating LLM-Based Evaluator

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:35:51.204544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-14T01:35:50.992477Z digest=sha256:701a752662faf659bb6b6c49fecef22bf9b1e654cde0e2adb708c5d458c41f2b

Observation 2784833e-994f-4b3d-b42f-4c80756c79fb · inbound

Towards unearthing neglected climate innovations from scientific literature using Large Language Models cites this paper.

Towards unearthing neglected climate innovations from scientific literature using Large Language Models Calibrating LLM-Based Evaluator

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:05:54.291216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:05:54.291216Z digest=sha256:ffa4c95863f8798b51d15a85c6abe4f4715310836665ba7af9015bce89456c3d

Observation a57f28fd-7676-4364-9c5f-b975c931315e · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Calibrating LLM-Based Evaluator

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.339313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.339313Z digest=sha256:877098cd0dc28eec03f60f75dd4a557ad0dd24b041ec85caec3e3b3de549d8d2

Observation ff2cbbb3-7af3-4f8b-927d-6c210e994160 · inbound

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator cites this paper.

Is my Meeting Summary Good? Estimating Quality with a Multi-LLM Evaluator Calibrating LLM-Based Evaluator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:15:44.970534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:15:44.970534Z digest=sha256:1fc5f24447dd4bbc5e6f6a574f3803d2bdf22f8027717d6f1ba13a5ab45f81f0

Observation 4ecfa734-88d1-454a-9f7f-1d66f2b609a9 · inbound

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions cites this paper.

A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions Calibrating LLM-Based Evaluator

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-11T20:37:54.919951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:37:54.919951Z digest=sha256:18cb8d567cac4293587c029a160472db2e511843313311156bba838e70ec5452

Observation d7b9ddb2-75e4-42d3-903a-7ffacf660847 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Calibrating LLM-Based Evaluator

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.331051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:fc305973b1a9198c95c01e015261019746123eae15fa30d377dd9bbb0644d2b6

Observation 90b2acb0-4145-43dc-98de-32ca74b6b5e4 · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations Calibrating LLM-Based Evaluator

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.260875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.260875Z digest=sha256:967d42a7a64bba7eaf9628af80e588128c9181fa117f05de715a18d901b768a7

Observation 8a8ceceb-ed1b-4e88-874f-d08af2c85220 · inbound

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents cites this paper.

VIDEE: Visual and Interactive Decomposition, Execution, and Evaluation of Text Analytics with Intelligent Agents Calibrating LLM-Based Evaluator

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:37:13.894987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T09:36:50.323723Z digest=sha256:88f3323d65c485a76620ba3c3e6fa156ef6e3f70598c5290da331c825d4234f2

Observation 5d5550e5-0e70-4a49-92f3-1286aa587f9c · inbound

Statutory Construction and Interpretation for Artificial Intelligence cites this paper.

Statutory Construction and Interpretation for Artificial Intelligence Calibrating LLM-Based Evaluator

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:52:55.567886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:52:55.567886Z digest=sha256:4d1a0c73c408caa655bdb1dbf95bd0a7e76d6572229ffa2a91385019e4c43ce5

Observation 3e8a5565-b587-45dd-90df-fa8a37b56243 · inbound

Calibrating Model-Based Evaluation Metrics for Summarization cites this paper.

Calibrating Model-Based Evaluation Metrics for Summarization Calibrating LLM-Based Evaluator

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.954717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T06:36:55.334742Z digest=sha256:b4e3f72905f6c0c5353c99f81f74c38dedce859a32c9bfd000b3c3c00abf9b2a

Observation 3ae689bf-abfe-4479-937f-c253e9ab8569 · inbound

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories cites this paper.

BabelJudge: Measuring LLM-as-a-Judge Reliability Across Languages and Agent Trajectories Calibrating LLM-Based Evaluator

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:49:41.496695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T11:06:07.303335Z digest=sha256:d96323e40d032098b2b249c4f2c47ce427adca59d3a8f9c74d06134551990136