Pith. sign in

Paper Citation Record · LEDGER

Do Large Language Models Judge Error Severity Like Humans?

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.05142.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05142 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:54.652219Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23e28b88-8d3b-46b1-9b99-c3f0e1dba8f9 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.163265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.520546Z digest=sha256:b53dd3ba41905a24918d4e26d830c27b07d6d6d0de1b842801e6bea226b72d8a

Observation 1f6c0407-0538-41c2-8fab-18ab5bad2ca9 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.148449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.525813Z digest=sha256:042458f169ab637e15af9a046d49e085be64ba26c53c00caaefbba4670bf94bd

Observation 95096cb4-f49b-4eb4-a0b8-b9d2020dea69 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.133540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.530604Z digest=sha256:31d548a3e14df75f205135e06a6186ec5a14a1f6bed6518948f7c0e94688e9b3

Observation 489e0593-877f-44ed-9f2b-d35daac3b446 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.117544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.535909Z digest=sha256:2e0c5631272f732634bcb145ca870302167dc2098c181929a59b84e75fca4cc6

Observation 538c5280-8f31-4ead-84f0-46f36e653094 · outbound

This paper cites LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks.

Do Large Language Models Judge Error Severity Like Humans? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.541102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.541102Z digest=sha256:78b6e42027850574ab4fb3f008074acd6f63855fb9e154b7084841dad42c8415

Observation 8f3dcd97-fecf-4efb-a0c3-c93b30946d3d · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T10:31:54.713324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.546447Z digest=sha256:a022c1f97db6e16171d750df48cd554fe75fce20dfdf9f56d8ceb27c63f978ad

Observation 1c9d168f-68ac-4789-9e93-da2e676a5d45 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.101594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.551925Z digest=sha256:3d65cd72993f3d7692db0c010f0f637bb77f38dd4bc5f45cfc1ad21426cd0530

Observation d8f5341d-b763-4673-bef0-8667d56d22b9 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.085846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.556447Z digest=sha256:aaa5d6ee4498419a2d7ea1ee39885f808a8fd8b75922b94dc05f67f6d3b86bfe

Observation 8bced555-f814-4944-914d-a9153b1f8a7d · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.069817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.561031Z digest=sha256:51051e5ac3a37fcdf1a66505430864a30577df42d548996a0ea9c11da39e677d

Observation da06e200-e497-4967-9fdc-13a25827909f · outbound

This paper cites A Survey on LLM-as-a-Judge.

Do Large Language Models Judge Error Severity Like Humans? A Survey on LLM-as-a-Judge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.565468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.565468Z digest=sha256:284f7bad3028df4cd0425ea1b2835520c7a2a248aec9212c1be86f6409bc3391

Observation 35e5b12f-1318-459e-995a-535adbe21f63 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.053107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.570566Z digest=sha256:60504c4acfe15d19a000326756920e16237fd524b0c7157848238aa80d8be168

Observation 75c367c4-6f58-4df0-8fc7-ab9ba014645c · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Do Large Language Models Judge Error Severity Like Humans? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.575005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.575005Z digest=sha256:054f030cf37c578e69f4ea47f6b044e80a2eb86cd6da08681ac15780ad5df562

Observation 043b73b8-63ac-4819-8010-b3b72f9d51a1 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:55.036558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.580279Z digest=sha256:b614e9d9a1d8829dd09147099a67735971c3ad8099de1e0e54097bc925a70874

Observation 08a0e908-5509-4bdb-885d-04962e1e23d7 · outbound

This paper cites GPT-4o System Card.

Do Large Language Models Judge Error Severity Like Humans? GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.584906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.584906Z digest=sha256:e4cd8acc2f2ed327e55c655b5bce22aa163882c76e3427437b52ec9c95c411d2

Observation c85c8a38-8cc4-4434-964d-beb7e1f6fc9a · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.590438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.590438Z digest=sha256:baf23f76d7554d761f456d35282027c4172e7f960c7050e6a0e1c9a21577b238

Observation 1411bdf0-d6ab-4a8f-9155-261076f5ad36 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.594921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.594921Z digest=sha256:eab9b53caa6a5779a8d940bd281bd6793dac213163fd8168839bf333c43d80e6

Observation 5cfe1b8f-f9e2-4edc-8849-2b66467e5277 · outbound

This paper cites DeepSeek-V3 Technical Report.

Do Large Language Models Judge Error Severity Like Humans? DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.599156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.599156Z digest=sha256:189adfbdf06bd879c16d36ba00452f6fdb10caa7b2d7667d58ee3a59b8470532

Observation f55ecfef-b332-4b57-bb84-381e851bb831 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:54.998296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.603855Z digest=sha256:c72503b64c0e7dd2bc2a808742e3f1049968900816e403e5dced13a88af9e48c

Observation d3438b11-d351-49b1-bb24-5570e39c17f5 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.608201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.608201Z digest=sha256:24760d3d54661653e8e307b134e980df6bc373046059370345000ae3caf309cc

Observation d294be5a-72d8-4e99-a28e-75a2f8f1e959 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.612813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.612813Z digest=sha256:0d414a412b4e2747a273209cf957570d6f5aa6a88142f36289da5187d3cd3927

Observation 90f9e6ab-fe22-4fd1-8302-55a5eb06cef1 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.618689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.618689Z digest=sha256:2be321c47214de5e62c1d7d65fbc4111ad4a27efb63baf82f27125544cf28ee4

Observation 156ef62c-5135-4d10-87d7-414d4b2654cb · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:54.970791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.623265Z digest=sha256:da04765a3b4fd1f2a37e8ce0464df16420273749c28c11979537df27a94e903f

Observation dda6c5df-a67a-4818-a97c-d7e49b525460 · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:54.953141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.627635Z digest=sha256:2b496f647877ea84c4a49c2191e3a768650e0d0896a85abd14094284b3296453

Observation e2624fef-0de2-4ca7-a184-acbb150bed80 · outbound

This paper cites Room for improvement in automatic image description: an error analysis.

Do Large Language Models Judge Error Severity Like Humans? Room for improvement in automatic image description: an error analysis

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:54.734049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.632681Z digest=sha256:976a11849acb5219ed420dbe09c42b51317451122c28330c8fed5b1a33f41aef

Observation 19b432ca-5063-4c77-bf13-f61522a6e01d · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-07T10:31:54.688687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.637345Z digest=sha256:20faba129e6ae694fdc4e05e7b8c03df0ecb0fbac8a26e6462a171007f77f9b6

Observation e185a3f3-63a4-4659-a469-e68948136b7c · outbound

This paper cites an unresolved cited work.

Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:31:54.937319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T10:31:54.642016Z digest=sha256:953b731b4c28150c450c81ee56ca20167c9afd5740b669310a221f774a0e9a9c

Observation eb84a8ee-9524-4d72-ac2d-5a5add128bff · outbound

This paper cites online" 'onlinestring :=.

Do Large Language Models Judge Error Severity Like Humans? online" 'onlinestring :=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.646925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.646925Z digest=sha256:533396e21b923f8304882c8c1944dbaf6b2845e42a16d1f8569c9d9627725945

Observation 53314b21-949d-494c-b398-17b19163128d · outbound

This paper cites write newline.

Do Large Language Models Judge Error Severity Like Humans? write newline

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.652219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.652219Z digest=sha256:547e88b22e23a2703fea406fd11700836643b5f1c2bd3d087aac1e850beb6343

Pith citing papers

No inbound Pith citation observations are available.