Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:54.652219Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2506.05142.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:54.652219Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 23e28b88-8d3b-46b1-9b99-c3f0e1dba8f9 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f6c0407-0538-41c2-8fab-18ab5bad2ca9 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 95096cb4-f49b-4eb4-a0b8-b9d2020dea69 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 489e0593-877f-44ed-9f2b-d35daac3b446 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 538c5280-8f31-4ead-84f0-46f36e653094 · outbound
Do Large Language Models Judge Error Severity Like Humans? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f3dcd97-fecf-4efb-a0c3-c93b30946d3d · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c9d168f-68ac-4789-9e93-da2e676a5d45 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8f5341d-b763-4673-bef0-8667d56d22b9 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bced555-f814-4944-914d-a9153b1f8a7d · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da06e200-e497-4967-9fdc-13a25827909f · outbound
Do Large Language Models Judge Error Severity Like Humans? A Survey on LLM-as-a-Judge
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35e5b12f-1318-459e-995a-535adbe21f63 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75c367c4-6f58-4df0-8fc7-ab9ba014645c · outbound
Do Large Language Models Judge Error Severity Like Humans? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043b73b8-63ac-4819-8010-b3b72f9d51a1 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08a0e908-5509-4bdb-885d-04962e1e23d7 · outbound
Do Large Language Models Judge Error Severity Like Humans? GPT-4o System Card
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c85c8a38-8cc4-4434-964d-beb7e1f6fc9a · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1411bdf0-d6ab-4a8f-9155-261076f5ad36 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cfe1b8f-f9e2-4edc-8849-2b66467e5277 · outbound
Do Large Language Models Judge Error Severity Like Humans? DeepSeek-V3 Technical Report
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f55ecfef-b332-4b57-bb84-381e851bb831 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3438b11-d351-49b1-bb24-5570e39c17f5 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d294be5a-72d8-4e99-a28e-75a2f8f1e959 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90f9e6ab-fe22-4fd1-8302-55a5eb06cef1 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 156ef62c-5135-4d10-87d7-414d4b2654cb · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dda6c5df-a67a-4818-a97c-d7e49b525460 · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2624fef-0de2-4ca7-a184-acbb150bed80 · outbound
Do Large Language Models Judge Error Severity Like Humans? Room for improvement in automatic image description: an error analysis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19b432ca-5063-4c77-bf13-f61522a6e01d · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e185a3f3-63a4-4659-a469-e68948136b7c · outbound
Do Large Language Models Judge Error Severity Like Humans? Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb84a8ee-9524-4d72-ac2d-5a5add128bff · outbound
Do Large Language Models Judge Error Severity Like Humans? online" 'onlinestring :=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53314b21-949d-494c-b398-17b19163128d · outbound
Do Large Language Models Judge Error Severity Like Humans? write newline
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.