Pith. sign in

Paper Citation Record · LEDGER

Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2401.16788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.16788 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:47.132103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T23:25:53.828114Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 21ed5251-c58b-499b-b1f8-05b415085c39 · inbound

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment cites this paper.

The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T10:36:17.166384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:36:17.166384Z digest=sha256:ff955c27eae52445a472c7266a83550736fca4c049b2c139fd0a465b81b386e6

Observation f7f60efa-51c8-4524-9b71-0dd0d6cd0b71 · inbound

Benchmarking LLM-based Relevance Judgment Methods cites this paper.

Benchmarking LLM-based Relevance Judgment Methods Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:47.132103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:47.132103Z digest=sha256:51f9d1aa17d5ef73aacec1d2c9955ce2d7fbcf8f92dcc4c72a3117a8f2559851

Observation e320f7bf-b1a1-4f84-87b0-134252a3aa29 · inbound

Generative AI Act II: Test Time Scaling Drives Cognition Engineering cites this paper.

Generative AI Act II: Test Time Scaling Drives Cognition Engineering Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:02:42.784034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:02:42.784034Z digest=sha256:a7a09e3b75c4f60d9591f1ef3a369cb957057074eee5717582d1437fddf03c46

Observation fd5030dc-b7ad-4f50-91ec-f625b2d85ee4 · inbound

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data cites this paper.

ChemActor: Enhancing Automated Extraction of Chemical Synthesis Actions with LLM-Generated Data Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:07.304009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:07.304009Z digest=sha256:400d205e1165a938741b4d59e4ad1ae6790dbf27dda48247f5b7b68a73becbbb

Observation cd6a15b7-77aa-44f1-8b53-a9df7295ca61 · inbound

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate cites this paper.

CortexDebate: Debating Sparsely and Equally for Multi-Agent Debate Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:11:10.537701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:11:10.537701Z digest=sha256:7d25660d90648cf147a5c651dce09b4ccb948a470c1a714ef705f2a7804f53cc

Observation e6ee4100-cff8-4116-ad67-63afce0e9c73 · inbound

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials cites this paper.

Generative Artificial Intelligence Extracts Structure-Function Relationships from Plants for New Materials Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T22:56:20.064068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:56:20.064068Z digest=sha256:3baacaa0d0d9d4e10b3020cda3adb31bd9f85c9454a26ce771097a6f0ff84d3d

Observation 93a1ffad-dabd-437f-8aa7-8d98a91b72ca · inbound

Learning to Interrupt in Language-based Multi-agent Communication cites this paper.

Learning to Interrupt in Language-based Multi-agent Communication Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:25:53.830634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T19:08:45.818851Z digest=sha256:d34cb274120b7e272d45670f3428ba19dc372fcb6b0e0b69dd9274fda5545622