Pith. sign in

Paper Citation Record · LEDGER

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2504.09946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.09946 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:26.281619Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:46:58.798974Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6f023a88-ad66-4856-bf90-666a5247da9f · inbound

Strategic Reflectivism In Intelligent Systems cites this paper.

Strategic Reflectivism In Intelligent Systems Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:26.281619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:26.281619Z digest=sha256:906a3ae81731f5198e072835b8e19df03995efd5046326e0e9503a32c3227e1f

Observation 4602b950-ac16-45f8-9cdb-edb20ef4be25 · inbound

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading cites this paper.

PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:57:11.379436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:57:11.379436Z digest=sha256:e8107c0957abb3ec493cd508c6cad8d6bae8989be967e4f7b79b100a3fa488cd

Observation 9c5711f3-5794-460f-a6df-54f652c3d180 · inbound

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning cites this paper.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:39.012600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:39.012600Z digest=sha256:cea7363eaef10d8a2a62a88d8ed986d87c272b0b7b08483953deb6559182a0f8

Observation 5c18f31d-decc-49ed-9033-a5fc6ef24b86 · inbound

One Token to Fool LLM-as-a-Judge cites this paper.

One Token to Fool LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:15:39.513701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:15:39.513701Z digest=sha256:33af4b9969f6270df20a6d1bb38cb6baf84e7ae04833d12d56c37dc5a852106d

Observation 8bdcb376-98e8-4552-9ed3-189f805dad09 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 177

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.098978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.098978Z digest=sha256:57c90190557b83ee335193718257d02ac815e4f72bea1159560612db7622176e

Observation 38f9930e-ce24-4312-8ac3-6a5e0fb3a48c · inbound

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models cites this paper.

Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T09:37:09.271466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:37:09.271466Z digest=sha256:feffdc542d27dc02d2fcf12aab7b6ee91fdd81c1f66f288827707314328c77db

Observation 079480e8-7c06-46c0-a94e-dc321e49ed16 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.228860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:1800c73f1db4031246cbbc074f3f6a80adc35bf5fd57bd6ac5e4ee5e640982d9

Observation 8a2caa14-9a8c-4dca-b0f2-00ba14ba491e · inbound

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models cites this paper.

Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:16.880554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:29:31.354743Z digest=sha256:9d305b518c9c0140b6d920936978c7efece6de619b8daafffc04b41253b2661b

Observation c37bc6d8-22b6-47f0-9ab2-6df4471574f2 · inbound

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge cites this paper.

Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:51.715805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T19:38:11.595077Z digest=sha256:e4faa871252f567dabd6a9625858c956a98a4fff6d076a68ef33faa45e2f1e3a

Observation 1fdf6cc4-d633-4575-a4fd-0dbcdededa35 · inbound

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding cites this paper.

LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:56:28.071739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:39:55.256518Z digest=sha256:bd0c56431ef5b0100f83ded1670c276f3838b663af2d57309d3bcaf10ed61778

Observation 04ef2b08-c858-4f54-97a1-7758616c055c · inbound

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents cites this paper.

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.764776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T09:53:00.677605Z digest=sha256:8498d631986644364c9b919f292f240a1127b5dd5a89615e00ba63962dff2912

Observation de7ab540-249c-4c37-88ba-c09b73a8695a · inbound

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models cites this paper.

An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:14.979242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T16:45:33.046568Z digest=sha256:4ad5825d3527b89465ce2446f3c4b6eb5f40a488acc2d1b7139fdb04e87a8576

Observation 5c98fb3c-8520-4226-b3ec-997ec28c0470 · inbound

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents cites this paper.

The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:06:35.117102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T09:27:48.400371Z digest=sha256:d9cd06dbe77072cdad1918adb50d7baab80f46dbada5a5df52454be4ca16e7ba

Observation 2e2cdb51-dc98-4b39-ac21-03865044b9a3 · inbound

A Mechanistic View of Authority Hierarchy in LLM Sycophancy cites this paper.

A Mechanistic View of Authority Hierarchy in LLM Sycophancy Assessing Judging Bias in Large Reasoning Models: An Empirical Study

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.800444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T13:41:18.279864Z digest=sha256:c2eb10f555172e16d42060ee14cc5551af3f4ad75f1e8a89cc29d871b1d3fef3