Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2402.10669.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:33:24.048215Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
9
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 4027fdd0-17f8-4fbb-868a-7ff070657a16 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 255
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7dd4cb82-3d90-4cf0-9336-09fe5bc730bc · inbound
ShieldGemma: Generative AI Content Moderation Based on Gemma Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 75fc38b4-ebb7-48a1-9090-e55fc7015d55 · inbound
From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fd6d3586-5beb-4c5b-9938-bd3790d629df · inbound
Evaluating Creativity and Deception in Large Language Models: A Simulation Framework for Multi-Agent Balderdash Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c236961-24bc-41a8-bca3-0b64bd74db0e · inbound
Dialectal Toxicity Detection: Evaluating LLM-as-a-Judge Consistency Across Language Varieties Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8e46ea7-78de-4d65-986d-b6179b43eb1d · inbound
Writing Style Matters: An Examination of Bias and Fairness in Information Retrieval Systems Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d3d734-3661-4946-82f2-a4826d009f37 · inbound
The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d797e876-a248-4128-84ec-4e6e856441d3 · inbound
A Survey on LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d2fa95c7-8c31-498e-8bdf-dec0558784c8 · inbound
Engineering AI Judge Systems Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7099f953-cf2c-41a1-b988-756f77c873d8 · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2054bfaf-0f8b-4756-894b-d4bd25d69240 · inbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4413d568-657e-476b-bb0b-b89f89467812 · inbound
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02db9a5c-1e71-4426-b6cc-829e4d8d7069 · inbound
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec22f9b3-d273-4408-82d9-5e49f8bdd9a2 · inbound
Anger Speaks Louder? Exploring the Effects of AI Nonverbal Emotional Cues on Human Decision Certainty in Moral Dilemmas Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92da7c97-23d3-46dc-96ee-c9e6b810c311 · inbound
Integrating LLMs with ITS: Recent Advances, Potentials, Challenges, and Future Directions Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 214
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c879e031-6704-409d-8004-5c04e4578781 · inbound
Unmasking Conversational Bias in AI Multiagent Systems Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4776e4-96de-4374-989f-efc659d284da · inbound
ZeroSumEval: Scaling LLM Evaluation with Inter-Model Competition Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b96c30df-bd6e-4c43-a40b-99e1dd633ef7 · inbound
Explainable AI in Usable Privacy and Security: Challenges and Opportunities Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f7353c8-ed5d-4574-87f6-0df17b641fc1 · inbound
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c6cb8ac-ddde-46e9-ae57-ccda4e56453f · inbound
Real-World Gaps in AI Governance Research Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ad02f7-fe5e-4985-83e8-0247a5af8f93 · inbound
The Hitchhikers Guide to Production-ready Trustworthy Foundation Model powered Software (FMware) Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a704e88-0a57-490d-aa55-a99138c00bb1 · inbound
Beyond the Surface: Measuring Self-Preference in LLM Judgments Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0696c877-6cb1-450c-898d-563b3cd76d5b · inbound
How Benchmark Prediction from Fewer Data Misses the Mark Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3212f0dd-80e7-48af-93e1-3ae1de774f2c · inbound
League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation adde8bc6-e443-4c2b-9b0a-7bb1bc90f9fa · inbound
Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1133009d-cbdc-4171-9003-eefcab3affe3 · inbound
Can You Trick the Grader? Adversarial Persuasion of LLM Judges Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3af1f4-8c56-4cf6-8c82-b198eb344214 · inbound
Can LLMs Make (Personalized) Access Control Decisions? Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b0c0797a-c760-4f0d-a1b4-dd183d554350 · inbound
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bdc06a1f-3152-4c53-b325-23af5498a4cf · inbound
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e11eef20-585c-4cd1-9f30-cc82530e31ea · inbound
Pioneer Agent: Continual Improvement of Small Language Models in Production Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 55cad9a5-2ce2-4824-b350-a07331f019f6 · inbound
Semantic Needles in Document Haystacks: Sensitivity Testing of LLM-as-a-Judge Similarity Scoring Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fa516d12-dbb4-4f71-b29a-ef868771e54e · inbound
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d52acf4d-da07-433f-a3f5-212bb7de8ab5 · inbound
Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 2d1153fc-2934-442f-9d40-a63336ac8edf · inbound
TRUST: A Framework for Decentralized AI Service v.0.1 Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4fca8ff2-f0e2-4b94-a225-46b6c1dc4946 · inbound
RecoAtlas: From Semantic Plausibility to Set-Level Utility in LLM Recommendation Agents Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a9aa0af7-0d17-4bf2-942b-87235affd6c8 · inbound
Are LLMs Bad at Moral Reasoning? Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 302284f3-6fdf-4743-b37b-05ebf3e098ae · inbound
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 236
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 182049fe-ebd1-4305-9170-b82122fce22a · inbound
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning Humans or LLMs as the Judge? A Study on Judgement Biases
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.