Pith. sign in

Paper Citation Record · LEDGER

Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.13439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.13439 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:02:43.147194Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:31:44.412419Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9a89e36-d6db-4c10-b1b8-17b8bc92a592 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.338295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:40f3c9fe9a14a513207ca2841882c7c14b9ef5f820b08129fd0c2499595c18bb

Observation a0e01cf3-5634-4125-9be2-74bf851d986f · inbound

Tuning LLM Judge Design Decisions for 1/1000 of the Cost cites this paper.

Tuning LLM Judge Design Decisions for 1/1000 of the Cost Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T15:02:43.147194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:02:43.147194Z digest=sha256:3a73b6a5d3b9963a8cb309a8190e7bda7e7aad70706534034afea8ed1eaf5eaa

Observation f636348e-0aa3-4358-9ac0-529ed93c7a90 · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.022146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.022146Z digest=sha256:5b84ac5cc17fc3c62f424cabe9ca0d6f21941cc5df9c76044b321babff4e34f8

Observation 939eba42-85e9-4744-b545-6c91aa1c1a78 · inbound

OpenCoderRank: Personalized Technical Assessments with Generative AI cites this paper.

OpenCoderRank: Personalized Technical Assessments with Generative AI Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:31:44.414821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T18:27:54.737962Z digest=sha256:02c8086f04d03b13eaf595ddfff521c959c2b598503fc84d3d0d906ed2f89805

Observation 0b05fa48-7213-4b20-b8f9-e5db37b70f35 · inbound

OpenCoderRank: Personalized Technical Assessments with Generative AI cites this paper.

OpenCoderRank: Personalized Technical Assessments with Generative AI Finding Blind Spots in Evaluator LLMs with Interpretable Checklists

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:00.865165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:09:00.865165Z digest=sha256:c67c4141174afa7bd22c3120b72fe0897e34f3595bb8c268c432b78244a5b5f1