Pith. sign in

Paper Citation Record · LEDGER

Evaluating Scoring Bias in LLM-as-a-Judge

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2506.22316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22316 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T05:30:26.408835Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e86fff31-0402-458a-b5f1-7ee0d4b2b6fd · inbound

Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge cites this paper.

Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge Evaluating Scoring Bias in LLM-as-a-Judge

Reference 2024

Resolution
malformed identifier
no resolver link, observed 2026-08-03T05:30:26.408835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:30:26.408835Z digest=sha256:c57ea6c527bf02405a4a6f0d1c71444e156917bb9debee5082d6b6cf95bf2cc5

Observation c8863da1-2085-4c5a-99ea-67cc1df6bf4a · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Evaluating Scoring Bias in LLM-as-a-Judge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T15:55:53.399860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:55:53.399860Z digest=sha256:c7d3a667d0edebe9fc11adcc7a390c36d7896b0b12d3aa2f25a01f6e61bd8080

Observation 314e2b7f-a1f1-4022-8ecb-f6ff2ac85c63 · inbound

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents cites this paper.

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents Evaluating Scoring Bias in LLM-as-a-Judge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T17:09:18.065540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:09:18.065540Z digest=sha256:8993026f3646b802762ca5bde885c8319b13b1cc6babb9b90cdb15f03630ecd8

Observation 8102e053-c053-4b67-bd3d-d6dc4f24cfd6 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Evaluating Scoring Bias in LLM-as-a-Judge

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:e46c9abe47af3c51c30ea0258196e0914c76d757b66cd4225034b9c83810ee1f

Observation 4cf0e146-d824-47de-ac93-d3e625f6f18e · inbound

Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction cites this paper.

Token Statistics Reveal Conversational Drift in Multi-turn LLM Interaction Evaluating Scoring Bias in LLM-as-a-Judge

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T09:20:21.372218Z digest=sha256:c0b15e8da250bf6110683e6deff5a62454401f322604e4bf2c63b4ce8b512953

Observation f3cd20a8-15e1-4610-825e-173ec1d15614 · inbound

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges cites this paper.

Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges Evaluating Scoring Bias in LLM-as-a-Judge

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T10:23:58.198464Z digest=sha256:573efa300d6db9382812c431212025a1675e3d8349c105fe7171561d0c925be7

Observation 3c8b3fcc-0e11-4d23-913f-87bb037a98cc · inbound

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation cites this paper.

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation Evaluating Scoring Bias in LLM-as-a-Judge

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:41:42.003483Z digest=sha256:6189497a527daa7e972cc2e1b069132583b0f80a1594cc29ed2f8ef9eb96c3aa

Observation b878fea0-28cd-42e2-848a-88619164a651 · inbound

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization cites this paper.

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization Evaluating Scoring Bias in LLM-as-a-Judge

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-14T21:16:46.840884Z digest=sha256:127eb270a67b26f97af0d9c2b8226fceb029188b4e0c7d5e75c77ec3534cb540

Observation 14cdd7bf-7832-4798-a34f-ed0ea6395695 · inbound

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization cites this paper.

ODRPO: Ordinal Decompositions of Discrete Rewards for Robust Policy Optimization Evaluating Scoring Bias in LLM-as-a-Judge

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-19T14:33:31.538969Z digest=sha256:e5d8f77439a27d2ee22d167b776a993fbc605495d2e557e0c079c6b1c1711830

Observation 07e8dd86-edba-4760-8a85-8345e866e236 · inbound

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring cites this paper.

Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoring Evaluating Scoring Bias in LLM-as-a-Judge

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:03:47.800935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T22:38:17.939704Z digest=sha256:ad07d69dd792b4ee93a8790cb269bc5f8b8d467d4632f74a1decf7e1ecd4c20f

Observation 94748b3a-64b5-4728-a08c-749747f8d9a1 · inbound

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research cites this paper.

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research Evaluating Scoring Bias in LLM-as-a-Judge

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-06-28T12:12:07.826015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T12:08:10.552789Z digest=sha256:37b810b30fd61580f6c166107185f3a1f7dbc5f7f62450c0e103cbe9236b661d

Observation c2d22f34-103d-4995-b826-237f8dec133e · inbound

Towards Fast Domain Adaptation and Fine-Grained User Simulation for Evaluating Conversational Recommender Systems cites this paper.

Towards Fast Domain Adaptation and Fine-Grained User Simulation for Evaluating Conversational Recommender Systems Evaluating Scoring Bias in LLM-as-a-Judge

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:09:48.818870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T07:13:21.198407Z digest=sha256:6666135aa517bf4064505e2cb2440361715f7718a1e79e12ea9cc28aadcb03e0

Observation f808fac0-18f9-4a75-be1a-ba4060ee7ead · inbound

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation cites this paper.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evaluating Scoring Bias in LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.637387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.637387Z digest=sha256:d483463869bb0c705452462be09331d49e5151e0897f54e51f32a75f54156d2b