Pith. sign in

Paper Citation Record · LEDGER

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation

As of 12 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2411.13212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13212 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:16.249314Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 719a5616-f964-4136-bae7-31952d4b207c · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.172964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.172964Z digest=sha256:f65dd3a7603d87b50987c170b15435e7fafb26c8872f1ad7bb8a4004f22d0bf5

Observation 23c5bd0d-aa97-4efe-8f17-2fa2e8b26ebb · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.181554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.181554Z digest=sha256:0ef078880da60541a4b546c2a5a3587c5c2f0343eb4c54e6da8f161c78da2946

Observation 11a01bd5-e960-49ab-af5e-b27b3506cbf1 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.188659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.188659Z digest=sha256:221cca58ee9bcc43e71c688af46f1a682e24b24092bb829a39496434bae2c260

Observation ba6f85e9-2379-493d-be2a-9cf6a11b8fc5 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:16.954953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:16.192143Z digest=sha256:7ef2ef183894ed11c5e242696f7cf4decc9df2204c872105994306dc9c39eb5a

Observation a319de42-8823-401c-ae81-93c3e7d9b326 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.195698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.195698Z digest=sha256:37d158f2ea14b85aa84e4944bfd941e3230adb74117c78559261387b0a8803a1

Observation 2877a261-d1ea-46e7-baae-f3e2868fecbd · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.199067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.199067Z digest=sha256:3959d19cd95613d3d9451fc1033329228be6d21c52bf7ce32636a79d0a2050cb

Observation 86fd2ccb-ce24-4e1e-8005-5574ed2bc5c5 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 8

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T16:45:16.286367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:16.202283Z digest=sha256:d9e7c8c8017e9a3f2e8d81b5e618ae62dc981f9fcd51e7b97df7a6dc80dd5caf

Observation 47810618-584b-478e-b3ab-0dc86d1b7c13 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:16.945776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:16.205442Z digest=sha256:a4278d9b49d33b8c45643e10328e0c60915b4e631793741768696321051fb827

Observation 63177dff-b5d9-4d3e-957b-3c3e97414cb9 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.208443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.208443Z digest=sha256:9d93d20c6ca96e9a39673c4e5ef0941360c0808cd28f0abbcd1660e67e501ad3

Observation 6cb2ca0e-e5b9-4f25-8638-1a7657acc8a2 · outbound

This paper cites Losada, and Álvaro Barreiro.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Losada, and Álvaro Barreiro

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.211432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.211432Z digest=sha256:722e1aa3c3209dc4a6913797100f7c50f015c019325a7c4f4c1ab53923247896

Observation 0f406067-1086-452a-9b58-e535f94cef88 · outbound

This paper cites Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.214392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.214392Z digest=sha256:16fee4c62e318c21ab2caf67af4fdc97c8b3ed637d5c240c3a7648a2333d3b34

Observation 1b0a63bc-0732-433b-b94f-57f2d74f53fb · outbound

This paper cites SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.217587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.217587Z digest=sha256:681b40aaa0248d1e6724bc6f9a462ea882ecea4f7b47f31db5014224bf4ef63d

Observation a75077f7-f433-44f9-95b8-7acf7330784d · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.220944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.220944Z digest=sha256:e221bbdab590a4141a8556ed087fe3e84c23ab2362bca013bee7dd010c829338

Observation 2d70b950-2328-4124-8d5b-f7841ce428da · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.224025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.224025Z digest=sha256:cb4c3131ca9bc4b6578219f3762b62d54373f36899ac20382427a2bc5e82143b

Observation 2d0f5d2f-ac9c-4096-81c6-035bc5fa079b · outbound

This paper cites LLMs Can Patch Up Missing Relevance Judgments in Evaluation.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation LLMs Can Patch Up Missing Relevance Judgments in Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.227040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.227040Z digest=sha256:cb20778cc5568400f80e2744ba8a1ca38551f53781b1f25dd4d14ffc9f655966

Observation d9a39eb1-883e-4ef9-8e2f-d267468b9c6c · outbound

This paper cites A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.230608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.230608Z digest=sha256:b96ac4646a6f37de8b3edcfda8e11423960e4f7b580f2bdb90c9baa857016dc5

Observation a28b4541-130e-4fb8-aad0-e16e78c42fa9 · outbound

This paper cites UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.233934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.233934Z digest=sha256:a7e6aa8caf808f9ae5e6e8ceb331356b9b84424db1c5cda83c763bcb7a38b71d

Observation 3173b6b2-a376-491c-b094-64c0585e645a · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.237208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.237208Z digest=sha256:3757e2064bf6310c4b971e8088e6bc9a2904e04f7aeb163091823f68c768f4e5

Observation fb359f50-d6b9-4044-a36c-19f311244491 · outbound

This paper cites Voorhees.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Voorhees

Reference 20

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T16:45:16.276630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:45:16.240158Z digest=sha256:d17e929d329e3c8b93cf34254df4acb4f893c95e74836dde0d0d57552ded711d

Observation 03e2af04-5850-4189-8f0e-3026fdc512b7 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.243268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.243268Z digest=sha256:71af9a7081cbe2eebc368f1a686d17db17a776503d712bcbadc7741742054709

Observation 239b7110-afd6-451b-a258-12c7eaa8d807 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.246309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.246309Z digest=sha256:66ed7e14275cdf4e484441bb68f5d4bc85661e5f5fafd24ac92f12385c90d0b6

Observation 183aa483-9e76-4d1a-9cda-f107feacfb56 · outbound

This paper cites Aslam, and Stephen Robertson.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Aslam, and Stephen Robertson

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.249314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.249314Z digest=sha256:51bcb36e64ce06556e49286757c62ef753709a94717a1ca37c006312eda13470

Observation 33d3d959-0248-4a3c-9cd5-4354eb5fa3d6 · outbound

This paper cites Can We Use Large Language Models to Fill Relevance Judgment Holes?.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Can We Use Large Language Models to Fill Relevance Judgment Holes?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.177722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.177722Z digest=sha256:64fbedd30a9408eb32f71f019a41ec0cad6ea1c363f992596e69d947dc1d66e8

Pith citing papers

No inbound Pith citation observations are available.