Pith. sign in

Paper Citation Record · LEDGER

Humans or LLMs as the Judge? A Study on Judgement Biases

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2402.10669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.10669 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:24.048215Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4027fdd0-17f8-4fbb-868a-7ff070657a16 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 255

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:44:49.863860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:5f0aa081b1b34e3a143e49f535743c0bc9b755713ef0635a65ff3d44de735d70

Observation 7dd4cb82-3d90-4cf0-9336-09fe5bc730bc · inbound

ShieldGemma: Generative AI Content Moderation Based on Gemma cites this paper.

ShieldGemma: Generative AI Content Moderation Based on Gemma Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:17:39.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T13:17:39.444002Z digest=sha256:52d230d590e5b553badd23ae535127f6dbd399747dcb4a0f410c2b544b64cd65

Observation 75fc38b4-ebb7-48a1-9090-e55fc7015d55 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.821090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:684a86b3ba3666c3221414e208b843a8e8cb601d7db5e076bb6be3cd98e1606b

Observation fd6d3586-5beb-4c5b-9938-bd3790d629df · inbound

Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash cites this paper.

Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:42:17.212896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:42:17.212896Z digest=sha256:4584da6179bae324b492d7dfff6d7af4878e9691bb2a3d0879fc835f67ea5354

Observation 7c236961-24bc-41a8-bca3-0b64bd74db0e · inbound

Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties cites this paper.

Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:11:30.660277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:11:30.660277Z digest=sha256:cc76bef8ee9ae22467f6b6a01ab761768347673e3ae591469ddcec965a42399e

Observation d8e46ea7-78de-4d65-986d-b6179b43eb1d · inbound

Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems cites this paper.

Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:49:20.375535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:49:20.375535Z digest=sha256:aa45d8d4fffc031e9899a45d42aaf434c17f897f37719c1d21e40ec1db767951

Observation 41d3d734-3661-4946-82f2-a4826d009f37 · inbound

The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges cites this paper.

The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:25:13.739893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:25:13.739893Z digest=sha256:879634dfcc143b4342b8b84549092c498103bea453fce5bf0e82d79a18b6751b

Observation d797e876-a248-4128-84ec-4e6e856441d3 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.180203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:6dbb1190409e0277f829232091393155fb80b6141801da5b7670966987583baa

Observation d2fa95c7-8c31-498e-8bdf-dec0558784c8 · inbound

Engineering AI Judge Systems cites this paper.

Engineering AI Judge Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T11:59:18.228507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:59:18.228507Z digest=sha256:fd0fce5437ec044a52dce20851e0aa936252a6b867e12aac414bf641746df5f8

Observation 7099f953-cf2c-41a1-b988-756f77c873d8 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.091045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:b8b264a670148a970b4383fc1f3db3b3f8e11b93f134ec1e8aaa8c564be9b888

Observation 2054bfaf-0f8b-4756-894b-d4bd25d69240 · inbound

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge cites this paper.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.132660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.132660Z digest=sha256:33758d569509c51bd2ea03fb139394dec36fc425d952624a63977df7bbc6ed18

Observation 4413d568-657e-476b-bb0b-b89f89467812 · inbound

Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition cites this paper.

Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:02:07.447822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:02:07.447822Z digest=sha256:cffba0afb380996342487a5118690a37ed4db598fd88f9c88e9fe9526a6cef15

Observation 02db9a5c-1e71-4426-b6cc-829e4d8d7069 · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.379357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.379357Z digest=sha256:806a064b6649134341f2735dcda0b73e0336ae987ca97ae61b7c2c127f36560f

Observation ec22f9b3-d273-4408-82d9-5e49f8bdd9a2 · inbound

Anger Speaks Louder? Exploring the Effects of AI Nonverbal Emotional Cues on Human Decision Certainty in Moral Dilemmas cites this paper.

Anger Speaks Louder? Exploring the Effects of AI Nonverbal Emotional Cues on Human Decision Certainty in Moral Dilemmas Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:07:48.629049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:07:48.629049Z digest=sha256:db7bf3c6e8d47f7399a761060902bc23df59c47c88da8d6bb1a630e339ec3a13

Observation 92da7c97-23d3-46dc-96ee-c9e6b810c311 · inbound

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions cites this paper.

Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 214

Resolution
unresolved
no resolver link, observed 2026-08-10T21:37:05.136174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:37:05.136174Z digest=sha256:a8776e7c4e9f2cf40a9bca3b1eed268973f4d65368bd16ecc8c10ab3b0632271

Observation c879e031-6704-409d-8004-5c04e4578781 · inbound

Unmasking Conversational Bias in AI Multiagent Systems cites this paper.

Unmasking Conversational Bias in AI Multiagent Systems Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:18:21.788231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:18:21.788231Z digest=sha256:2ebe2f4f2087995947cc38d542ba6bff453b3afae02f985abb1f0d7ae9624433

Observation 2e4776e4-96de-4374-989f-efc659d284da · inbound

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition cites this paper.

ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:33:24.048215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:33:24.048215Z digest=sha256:0220b9e4cd698ec94d57f9178a9da808bbc416bd2ce5fcf7d7c2c8aec07f5887

Observation b96c30df-bd6e-4c43-a40b-99e1dd633ef7 · inbound

Explainable AI in Usable Privacy and Security: Challenges and Opportunities cites this paper.

Explainable AI in Usable Privacy and Security: Challenges and Opportunities Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T12:21:21.268668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:21:21.268668Z digest=sha256:a5d768f0574d0d9052a45a25fed2fd901d6d35a4661bb4780c05bff429c9345f

Observation 3f7353c8-ed5d-4574-87f6-0df17b641fc1 · inbound

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents cites this paper.

A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:47:11.126977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:47:11.126977Z digest=sha256:98e6e44c7abf96bbc9944af479f886bec03a8e4cc98c4b0a22e66e3fba079386

Observation 2c6cb8ac-ddde-46e9-ae57-ccda4e56453f · inbound

Real-World Gaps in AI Governance Research cites this paper.

Real-World Gaps in AI Governance Research Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:53:11.571819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:53:11.571819Z digest=sha256:606cdcdd23a260baf432373ca6e2b27955682f4bb4d5b2c8a6d43dc59b5b2cfd

Observation d6ad02f7-fe5e-4985-83e8-0247a5af8f93 · inbound

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) cites this paper.

The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T21:10:19.784175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:10:19.784175Z digest=sha256:c660dc5c9158a81c7aa282f0e4853dfb0e8499f263d19d639198195b45572a2d

Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:01.891765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:01.891765Z digest=sha256:17df3a0f8e7251780d027cb50db33a75021b8ad85c527c2fd5f978d196d0bde6

Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · inbound

How Benchmark Prediction from Fewer Data Misses the Mark cites this paper.

How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:34:26.647558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:34:26.647558Z digest=sha256:fe98586f6d620b07165c29843f005e86d221281c2086b0b6cdd3fd1804693cb2

Observation 3212f0dd-80e7-48af-93e1-3ae1de774f2c · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:22:01.345621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:24810a0f0796c9a5a5d6777a94cce7fe4cef4353685c19e63409894507cdd946

Observation adde8bc6-e443-4c2b-9b0a-7bb1bc90f9fa · inbound

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support cites this paper.

Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T23:04:56.860420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:04:56.860420Z digest=sha256:21254e2f4a6e07c009a8c849c43cc05610dfa73b97267fe23d67d6f9ba2a6784

Observation 1133009d-cbdc-4171-9003-eefcab3affe3 · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:55:56.852022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:55:56.852022Z digest=sha256:aa828b8ecd27452296d7d6ed6a876409ab6d6af5dd1f05ef020436b2af2031b2

Observation 1e3af1f4-8c56-4cf6-8c82-b198eb344214 · inbound

Can LLMs Make (Personalized) Access Control Decisions? cites this paper.

Can LLMs Make (Personalized) Access Control Decisions? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:34:05.015121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T05:33:18.457421Z digest=sha256:ef91454eeefcdcc164a821c2b5a37ccd7df5830ee2859f65171f51b73ac6fa29

Observation b0c0797a-c760-4f0d-a1b4-dd183d554350 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:47:15.241556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:5eef081cc622ebd11f8e0b59fbcf501bf71e2722aa77060b16ab4b8dc9780fcc

Observation bdc06a1f-3152-4c53-b325-23af5498a4cf · inbound

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis cites this paper.

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:09:53.360123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T09:06:33.531027Z digest=sha256:54beaa59cc1e62f1a4f0e3d38f51c79454454f89de69feaef630677cda7397ef

Observation e11eef20-585c-4cd1-9f30-cc82530e31ea · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.610523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:1bac0d6497781533b5bcb4e14114d9fea64fa8b00d6c21230844881e7e3b5f73

Observation 55cad9a5-2ce2-4824-b350-a07331f019f6 · inbound

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring cites this paper.

Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:51:03.617526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T04:31:53.825854Z digest=sha256:d992d8c80b8dedb110eea16bdb7bd7154ece03385ec330f9951a36dbbb08921a

Observation fa516d12-dbb4-4f71-b29a-ef868771e54e · inbound

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards cites this paper.

Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:21:07.359059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T21:58:05.584559Z digest=sha256:fad18fbc8aa980578a54adbf4765d4b14e681892db6d9afc1d99ab3a2fc01515

Observation d52acf4d-da07-433f-a3f5-212bb7de8ab5 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.136347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:00d99fe08b40c62296074f67a4b756553fbdd7ebca49167d4b412671a6461f38

Observation 2d1153fc-2934-442f-9d40-a63336ac8edf · inbound

TRUST: A Framework for Decentralized AI Service v.0.1 cites this paper.

TRUST: A Framework for Decentralized AI Service v.0.1 Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:01:28.726224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-07T08:22:14.239443Z digest=sha256:d4208d8f442465dad78d568f6ac7110bc0ced4cc2ebdd91b4cd4d9a747b739bf

Observation 4fca8ff2-f0e2-4b94-a225-46b6c1dc4946 · inbound

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents cites this paper.

RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:29:09.209885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T22:27:16.974169Z digest=sha256:7f1774889e72b50239216f5263289e9766615e5c877de055640947b42623dc9e

Observation a9aa0af7-0d17-4bf2-942b-87235affd6c8 · inbound

Are LLMs Bad at Moral Reasoning? cites this paper.

Are LLMs Bad at Moral Reasoning? Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.640230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:20:24.251540Z digest=sha256:e0da1eb71568f87d39f87d9f0519395bff8376719d51feb0dac808e4d30f655f

Observation 302284f3-6fdf-4743-b37b-05ebf3e098ae · inbound

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering cites this paper.

Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 236

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:55:06.221097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T22:54:03.054871Z digest=sha256:60c48b465b10e48e4dfd64e8d4e446ff0468357fd135d357f07a7a614690a87f

Observation 182049fe-ebd1-4305-9170-b82122fce22a · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Humans or LLMs as the Judge? A Study on Judgement Biases

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.097000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:c03d7ab5ce2cd7e78185e61a404b7c5a4ce5a834b6062de434f90d95112ee337