Pith. sign in

Paper Citation Record · LEDGER

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2504.09946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09946 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:30:44.834444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:46:58.798974Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f023a88-ad66-4856-bf90-666a5247da9f · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:26.281619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:26.281619Z digest=sha256:b7955ea3cb6d4e3a2259b0d364dbee88b222dc2bc175d07faca63de5b8fc1e40

Observation 4602b950-ac16-45f8-9cdb-edb20ef4be25 · inbound

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading cites this paper.

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:11.379436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:57:11.379436Z digest=sha256:8a8e9f75a2581e808721371610078872044ad5b09875a98032e019851665447c

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · inbound

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning cites this paper.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:4d5b8a8b0504260d6fd4455258d204e0f41a4d42d56cb9a0a7b7b7ba91031d97

Observation 5c18f31d-decc-49ed-9033-a5fc6ef24b86 · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:39.513701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:39.513701Z digest=sha256:304c3808ae34c17a85fce600b1b2fe2a764cadb6501d8f91cd4dac671360284c

Observation 8bdcb376-98e8-4552-9ed3-189f805dad09 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.098978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.098978Z digest=sha256:e5a5e2e1baa0e555f77eeb450292f7e5ca6bf88d39f2fa64f8a0cee73ea8d682

Observation 38f9930e-ce24-4312-8ac3-6a5e0fb3a48c · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:09.271466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:09.271466Z digest=sha256:c8fc190bc3ad8fc4eda66d4935db0f14d1bccbb4e98579e34cacfad136de1296

Observation 079480e8-7c06-46c0-a94e-dc321e49ed16 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.228860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:76557c2176c0219154ad7b1a5a9cd82b2e69ee6615b6b95f4ded2d2a89157d87

Observation 8a2caa14-9a8c-4dca-b0f2-00ba14ba491e · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.880554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:86ea43bc0cc5be2dcf3594e85049d580e469a7bc42c9283a312fe1326b983950

Observation c37bc6d8-22b6-47f0-9ab2-6df4471574f2 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.715805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:19597bde00e84754ed4e41215cda8c10bf5daf51d550a37b92cf21a891743a5a

Observation 1fdf6cc4-d633-4575-a4fd-0dbcdededa35 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:28.071739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:fa574156bb106cd2f08bb962037a80dc342107b98d79c7b7fffc4e1279d2a549

Observation 04ef2b08-c858-4f54-97a1-7758616c055c · inbound

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents cites this paper.

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T09:53:00.677605Z digest=sha256:d89095b64c1161dc156d2faab9220aa6e889f20e14e4930716e95c9887383311

Observation de7ab540-249c-4c37-88ba-c09b73a8695a · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.979242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:c4063a76139a32abb3cf2517b2951c5f46d94367745a82d27859a97f8bbd6a57

Observation 5c98fb3c-8520-4226-b3ec-997ec28c0470 · inbound

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents cites this paper.

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.117102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T09:27:48.400371Z digest=sha256:6c831a0b30593134339b3687c2826033c547720d1cf54eef7d66ce58e2b8e25c

Observation 2e2cdb51-dc98-4b39-ac21-03865044b9a3 · inbound

A Mechanistic View of Authority Hierarchy in LLM Sycophancy cites this paper.

A Mechanistic View of Authority Hierarchy in LLM Sycophancy Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.800444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-02T13:41:18.279864Z digest=sha256:1e315b66287d39fcb8be7e4496f1c414ffe9387c7434080c3543755cb29a95bc

Observation 0f63028f-c491-4286-899d-89e202045863 · inbound

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs cites this paper.

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:30:44.834444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:30:44.834444Z digest=sha256:1a280716d2d46806ebb9a7e5a5335857e399152ea8e9119c14cdce5439ff65a0