Pith. sign in

Paper Citation Record · LEDGER

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons

As of 19 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2506.03785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03785 v3

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:58:27.657384Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 331d1f76-3d4d-4d0a-9293-1f544e829147 · outbound

This paper cites an unresolved cited work.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:26.583069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:26.583069Z digest=sha256:7d9134d90df8ba64bbcdea0e270c585f01502d7ee2a679d27d982ac0e4dbe5b0

Observation 647f39a9-061e-4496-b608-7e577fe028fd · outbound

This paper cites SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:58:28.168041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T10:58:26.696473Z digest=sha256:3e03549117b3b3a0976ec170e392dbd57abfd7655de05d8f85b144f89758b1dd

Observation 85d3b424-5874-4216-83a8-654b2f96e3ad · outbound

This paper cites an unresolved cited work.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:26.789363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:26.789363Z digest=sha256:4518737398b237ae86e0faa707b220d8660bdd9aa349ea775aeecc724b5fe18d

Observation fceb344c-b0ce-4cbc-aa62-dcc7fff13356 · outbound

This paper cites an unresolved cited work.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:26.930200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:26.930200Z digest=sha256:8c183b612f18154c98e7fc4ec9c94c07f4e1d0ecd9b49176812b4052ccf63e20

Observation df6d0917-afb7-454d-b77f-f2a6d1a7b858 · outbound

This paper cites LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.025105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.025105Z digest=sha256:f0c31a08b0b737947f53b4b8fa814df95602550f4337e23c50a3c88b7a6236bb

Observation 3abebc9c-4799-4c0d-9afd-db3fab180062 · outbound

This paper cites Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.135754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.135754Z digest=sha256:d458df4845aad00c8f3fd09a388c74a92f296178f70969e026a14282cd5d5efb

Observation 333b0b5b-1c0a-4de9-9522-bc719485b433 · outbound

This paper cites Large Language Models are Biased Because They Are Large Language Models.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Large Language Models are Biased Because They Are Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.262021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.262021Z digest=sha256:0d658acbaf668a20ea664a86caa606f0ce76a82fe8c7cbb5a374b8bdf4c056c5

Observation d54c5bac-9273-48d6-b8c2-88b91f717d13 · outbound

This paper cites Is ChatGPT a Good NLG Evaluator? A Preliminary Study.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Is ChatGPT a Good NLG Evaluator? A Preliminary Study

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.365332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.365332Z digest=sha256:a74e399ac42445b8fa3105b0913536dbb1bed28f973394cdace9d796cd6bb032

Observation a9f5cabf-7b10-4ae3-8b80-50c79fa0b58e · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.445120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.445120Z digest=sha256:8301934d2ee9e0b94d4e3ca99acfef99175e13871ccf76a04ca57572b01ce585

Observation 5dc644c4-dcf5-48f5-ad19-b45acd21457d · outbound

This paper cites online" 'onlinestring :=.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons online" 'onlinestring :=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.548955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.548955Z digest=sha256:4b86fdc0bf74be9bd236bc59e7eacea0f716a9f27687a3380893b13ffa03ecbf

Observation 604b1408-b9d0-4fb3-970b-6e7158486d25 · outbound

This paper cites write newline.

Knockout LLM Assessment: Using Large Language Models for Evaluations through Iterative Pairwise Comparisons write newline

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:58:27.657384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:58:27.657384Z digest=sha256:dd7a0a5b42d48653d0532829639aa7aa92fe3af643092f3e05c3a2040212fee0

Pith citing papers

No inbound Pith citation observations are available.