Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2406.18403.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:21:21.253366Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 754be9ac-cab6-4a3e-9497-30cb7c453b1c · inbound
Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5495ec06-07b3-4dc0-b3cf-22ea57457e7e · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bdfe1bc-f1d0-4fad-9705-568971ba4c66 · inbound
On Limitations of LLM as Annotator for Low Resource Languages LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d05559fc-20c9-41ae-863b-2a6a34e2294c · inbound
LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 159f5c08-43d9-4c24-97e3-04217a3c7b97 · inbound
JuStRank: Benchmarking LLM Judges for System Ranking LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65fb1b27-013c-4bab-9972-429cb03127ec · inbound
Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c92bdb5-2795-40cf-8fef-06fc07aecfda · inbound
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f25c9310-ceb0-45a6-b316-6773f4780eae · inbound
Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5321d1f0-b718-43c9-9449-a56be0d54f13 · inbound
LLMs can be easily Confused by Instructional Distractions LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f24b30-7e2e-45d4-8b78-2e170cd8f140 · inbound
Aligning Black-box Language Models with Human Judgments LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce7ba2cf-2306-42a0-b4fd-c37869d47b85 · inbound
Explainable AI in Usable Privacy and Security: Challenges and Opportunities LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3705a84-0292-452c-a1b8-e246cb5003cc · inbound
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76a27338-92c0-4937-b367-6759d065cae1 · inbound
What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17799c0e-3d5c-4d19-95bb-09f1777eaf04 · inbound
IberBench: LLM Evaluation on Iberian Languages LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc7a9f1d-e85b-46ff-aeaf-ba4d4a536bed · inbound
Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 910af266-d3ef-4788-97c6-e44e8c4e72fc · inbound
Patterns and Mechanisms of Contrastive Activation Engineering LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b54337a7-2f13-4943-81bd-a9ddc7108038 · inbound
Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538c5280-8f31-4ead-84f0-46f36e653094 · inbound
Do Large Language Models Judge Error Severity Like Humans? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179b2d20-9325-4ff1-b447-3b8fc0cc03b0 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a2d426-3228-482a-9c22-0ce3cae0da50 · inbound
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6320319-de2b-45dd-8b78-e569c378aecc · inbound
Hatevolution: What Static Benchmarks Don't Tell Us LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69b0178-882e-4c9b-beba-e99237dd2353 · inbound
Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ddfcbc4-d589-44a4-a627-a6805ccc57bc · inbound
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d057c26-ba66-41aa-8fc7-1033fdc3a767 · inbound
IMPACT: Inflectional Morphology Probes Across Complex Typologies LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb83d160-535e-4026-9385-50edf7de0464 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1472de57-2950-4def-b719-a5b691e572ed · inbound
Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 198
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d20c4c79-101e-4bc7-bbb9-d1b2020c5a7e · inbound
Real-World Summarization: When Evaluation Reaches Its Limits LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c29cf48-4c00-4b26-80f0-b9c83b4cd5ec · inbound
Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143bd1cf-07e1-414b-95d2-f560bfb658f2 · inbound
Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5e49d5-924a-4dee-9172-01983c6a4b12 · inbound
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e10f27-aa64-468c-b9c8-fd9a0f02af6a · inbound
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5c050c-6db0-4d80-b492-286daf836952 · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d20dbc68-7d9a-4ff9-9261-76849587813b · inbound
Guidelines for Empirical Studies in Software Engineering involving Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a8d66146-810a-4fb2-bea0-7ec6e93c4326 · inbound
AI Propaganda factories with language models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b2cd72e-8192-4832-a563-6d696062a4a8 · inbound
E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b781d0d5-4c57-4c73-9a6c-66a0c0c80a2c · inbound
Personalized AI Practice Replicates Learning Rate Regularity at Scale LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 894f4d4a-d92d-4f67-8db3-f8d7c631899c · inbound
Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 14dd8efc-1881-4428-b542-536b9c5c141b · inbound
Mixed response geometry and critical crossover in the Ising model LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67dceee5-e2ae-4fd2-835b-083e57a810d2 · inbound
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 221
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b4a85437-cc68-4617-a7b3-c9a9735460ec · inbound
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 221
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b62dad26-4a6c-4138-b313-3daa2778c68a · inbound
Instructions Shape Production of Language, not Processing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7dd30433-18a3-46c7-8688-f37703eb8a72 · inbound
Instructions Shape Production of Language, not Processing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4d5897e8-6118-4e7c-ad90-2babf4c7b44b · inbound
Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7eb2240-5837-4610-9687-d8cf96f1b79f · inbound
Natural Language-Focused Software Engineering via Code-Documentation Equivalence LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7aeec3ce-04c7-45d7-b70f-9361993fc159 · inbound
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.