Pith. sign in

Paper Citation Record · LEDGER

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2406.18403.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.18403 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:21:21.253366Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 754be9ac-cab6-4a3e-9497-30cb7c453b1c · inbound

Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet? cites this paper.

Large Language Models as Robust Data Generators in Software Analytics: Are We There Yet? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:37:58.331238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:37:58.331238Z digest=sha256:185eadd6872e690358db1d9b66dff3c852cf2615051efb0e834eeb2cd578d83e

Observation 5495ec06-07b3-4dc0-b3cf-22ea57457e7e · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:29.387291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:29.387291Z digest=sha256:5d1bc03711664882a56d0cdb94c7e96347036ad96d0b29a848b02ce317a0d522

Observation 7bdfe1bc-f1d0-4fad-9705-568971ba4c66 · inbound

On Limitations of LLM as Annotator for Low Resource Languages cites this paper.

On Limitations of LLM as Annotator for Low Resource Languages LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:26.259206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:56:26.259206Z digest=sha256:02ec5031ec19095173aa59a9a71f266f6cd0ad0b0f0144f18e81f1d87667e9ab

Observation d05559fc-20c9-41ae-863b-2a6a34e2294c · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:36.025848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:a8098ec830f55614641ef88315a0bc2ba56f34aec0988ee13282c4b76da0a40e

Observation 159f5c08-43d9-4c24-97e3-04217a3c7b97 · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.430351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.430351Z digest=sha256:b76fa3ca9c12f6fe67fa8b77034ee19958635b931c0c58276d792bd10b8b8edd

Observation 65fb1b27-013c-4bab-9972-429cb03127ec · inbound

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change cites this paper.

Enhancing LLMs for Governance with Human Oversight: Evaluating and Aligning LLMs on Expert Classification of Climate Misinformation for Detecting False or Misleading Claims about Climate Change LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:39:33.294297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:39:33.294297Z digest=sha256:f3d81eb2ec0e6c60640cea7eca323c09bdd4cef72211681d2f56d354e6d1b9dd

Observation 2c92bdb5-2795-40cf-8fef-06fc07aecfda · inbound

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs cites this paper.

A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T10:44:41.083648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:44:41.083648Z digest=sha256:e313dc2deead578a472450c63114c091456744fce8bf89f5911d4e5a205a0349

Observation f25c9310-ceb0-45a6-b316-6773f4780eae · inbound

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences cites this paper.

Powering LLM Regulation through Data: Bridging the Gap from Compute Thresholds to Customer Experiences LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:52.906859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:51:52.906859Z digest=sha256:92688b53e71f51fbe0e7a599a2732dad6b221979f3a902731b96cb4378e7dc75

Observation 5321d1f0-b718-43c9-9449-a56be0d54f13 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.136791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.136791Z digest=sha256:8b11514f16b0ccab14ad7a79f480ba778e8fbd39978010e4dfb0afb3a12d7de7

Observation d9f24b30-7e2e-45d4-8b78-2e170cd8f140 · inbound

Aligning Black-box Language Models with Human Judgments cites this paper.

Aligning Black-box Language Models with Human Judgments LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T20:45:46.264867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:45:46.264867Z digest=sha256:a7596ea23fa24ed6b61d5b4963716ef68b06382b9c12911c4018262a7647b1b9

Observation ce7ba2cf-2306-42a0-b4fd-c37869d47b85 · inbound

Explainable AI in Usable Privacy and Security: Challenges and Opportunities cites this paper.

Explainable AI in Usable Privacy and Security: Challenges and Opportunities LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:21:21.253366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:21:21.253366Z digest=sha256:0ae5803576ef54dd55ff22544f2e5fae95241292331555dec73198c3b411c47a

Observation d3705a84-0292-452c-a1b8-e246cb5003cc · inbound

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results cites this paper.

Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:09:30.947387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:09:30.947387Z digest=sha256:1392080e9e14258783fd1095f43ce51c23a88db9e6069dfa7a34b32c267ceaa5

Observation 76a27338-92c0-4937-b367-6759d065cae1 · inbound

What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns cites this paper.

What's the Difference? Supporting Users in Identifying the Effects of Prompt and Model Changes Through Token Patterns LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T11:21:29.092702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:21:29.092702Z digest=sha256:a4c44ce45253f63b2a37fa40eaff65504ee52285ce02753bcec635a0eac0c602

Observation 17799c0e-3d5c-4d19-95bb-09f1777eaf04 · inbound

IberBench: LLM Evaluation on Iberian Languages cites this paper.

IberBench: LLM Evaluation on Iberian Languages LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:27.380569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:27.380569Z digest=sha256:d4952a0db082c9dd2631acffb106d5c8edd849b910d2936afaed211e519d6d10

Observation fc7a9f1d-e85b-46ff-aeaf-ba4d4a536bed · inbound

Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection cites this paper.

Evaluating Evaluation Metrics -- The Mirage of Hallucination Detection LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:29:31.183005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:29:31.183005Z digest=sha256:8c3adc6f0a8cb5e3556ba402110d8c12be157233d3ba604dcd1f3d79cb0f5a26

Observation 910af266-d3ef-4788-97c6-e44e8c4e72fc · inbound

Patterns and Mechanisms of Contrastive Activation Engineering cites this paper.

Patterns and Mechanisms of Contrastive Activation Engineering LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:00:08.628192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:00:08.628192Z digest=sha256:a22787cc13b7a34c2b327f356770ca6d66aa165cd1ebd95a2cec6a124e988c06

Observation b54337a7-2f13-4943-81bd-a9ddc7108038 · inbound

Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering cites this paper.

Integrating Neural and Symbolic Components in a Model of Pragmatic Question-Answering LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.419950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:46:59.419950Z digest=sha256:6493728077c0828fc81b780d756f318d06208cdf79d8227fa42f184d9bff3516

Observation 538c5280-8f31-4ead-84f0-46f36e653094 · inbound

Do Large Language Models Judge Error Severity Like Humans? cites this paper.

Do Large Language Models Judge Error Severity Like Humans? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:54.541102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:54.541102Z digest=sha256:ed37ec92c23f775f046753ae85e1d592f95b39dac6c8de219ebc2a1b4b975347

Observation 179b2d20-9325-4ff1-b447-3b8fc0cc03b0 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.849677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.849677Z digest=sha256:5dcbe44ffff4e4891c6e0906d017180a4c237e5d11701de37a75c3255febc73b

Observation 15a2d426-3228-482a-9c22-0ce3cae0da50 · inbound

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning cites this paper.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.572175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.572175Z digest=sha256:9d45653e8c32d3f7206f268112c2716dfe1a77ee56a53ed7079afe4839a10c87

Observation a6320319-de2b-45dd-8b78-e569c378aecc · inbound

Hatevolution: What Static Benchmarks Don't Tell Us cites this paper.

Hatevolution: What Static Benchmarks Don't Tell Us LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:08.843925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:08.843925Z digest=sha256:390742db262b4385a712fdb2187ea21540f2c982f1db822953875a159fe548ea

Observation f69b0178-882e-4c9b-beba-e99237dd2353 · inbound

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress cites this paper.

Has Machine Translation Evaluation Achieved Human Parity? The Human Reference and the Limits of Progress LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:12:43.431998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:12:43.431998Z digest=sha256:fa037cfc0e8d63d20a2fad30914a4725697a75fcdef6e1de8320c8f820432334

Observation 6ddfcbc4-d589-44a4-a627-a6805ccc57bc · inbound

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It cites this paper.

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:07.494636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:07.494636Z digest=sha256:5cfb7798fc29925debed02cb9745c87c971a866250371b1351f1f335c121cea5

Observation 4d057c26-ba66-41aa-8fc7-1033fdc3a767 · inbound

IMPACT: Inflectional Morphology Probes Across Complex Typologies cites this paper.

IMPACT: Inflectional Morphology Probes Across Complex Typologies LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:12.400522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:12.400522Z digest=sha256:74819c829b6d4b91a9c68679bdc0e42fbdd79a0b927472246d174ae8c19d5eae

Observation eb83d160-535e-4026-9385-50edf7de0464 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:20.815544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:20.815544Z digest=sha256:9c67c2bf92802f5d6da8d806c44a04830181641cd907f029aff7eb33439ad266

Observation 1472de57-2950-4def-b719-a5b691e572ed · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:11.055518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:11.055518Z digest=sha256:56c9d2abc03161f664bd8b7bf0397577cafd07bd6180741017b40cd3cf7fff5d

Observation d20c4c79-101e-4bc7-bbb9-d1b2020c5a7e · inbound

Real-World Summarization: When Evaluation Reaches Its Limits cites this paper.

Real-World Summarization: When Evaluation Reaches Its Limits LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:28.574651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:10:28.574651Z digest=sha256:4a4da4187c7ad4f5dfa428b5e0f2a443ad734e1b0fb524169a6531815511c0b3

Observation 7c29cf48-4c00-4b26-80f0-b9c83b4cd5ec · inbound

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? cites this paper.

Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge? LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:21.680804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:02:21.680804Z digest=sha256:6dc48011cf67ccccd5a3f9e7a6017ee15dd40ff81965e73d5fbedd85acc883f6

Observation 143bd1cf-07e1-414b-95d2-f560bfb658f2 · inbound

Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems cites this paper.

Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:56:12.081316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:56:12.081316Z digest=sha256:e0e3bed4d5b333a385deabe61ce18ac52571d7a4c52ce3cc9440754c6f6111c7

Observation 4b5e49d5-924a-4dee-9172-01983c6a4b12 · inbound

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts cites this paper.

FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:56:44.032983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:56:44.032983Z digest=sha256:d64910059f9faf36b645785eef3157f9697e551e6e6729e3cde9b55f3381df45

Observation 96e10f27-aa64-468c-b9c8-fd9a0f02af6a · inbound

LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data cites this paper.

LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T19:40:50.542299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:40:50.542299Z digest=sha256:11542290ed86327c2c020868ba6d57aa7f89cc45bbe5be291f1c6645df600026

Observation 0a5c050c-6db0-4d80-b492-286daf836952 · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:02:52.402406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T22:02:36.307598Z digest=sha256:276a1bb2c2148dc9404d74978689872174dda9b985dc5756b4b7649c9c1a93d0

Observation d20dbc68-7d9a-4ff9-9261-76849587813b · inbound

Guidelines for Empirical Studies in Software Engineering involving Large Language Models cites this paper.

Guidelines for Empirical Studies in Software Engineering involving Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:20:31.421064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-25T08:18:18.448122Z digest=sha256:a42d11961501c3b7c196e0774f768733015c2525efa6853d7e61d3010187c73c

Observation a8d66146-810a-4fb2-bea0-7ec6e93c4326 · inbound

AI Propaganda factories with language models cites this paper.

AI Propaganda factories with language models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:21:32.532441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:21:32.532441Z digest=sha256:954be8a1dab3ff6c81f06e46825c00bb4209698122088aa4f45deafbee8f96d5

Observation 4b2cd72e-8192-4832-a563-6d696062a4a8 · inbound

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing cites this paper.

E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T19:10:09.763083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:10:09.763083Z digest=sha256:c8f043dc1eea9f75ac56defebcba689a85c535488b89381a6595a59adf40c3d2

Observation b781d0d5-4c57-4c73-9a6c-66a0c0c80a2c · inbound

Personalized AI Practice Replicates Learning Rate Regularity at Scale cites this paper.

Personalized AI Practice Replicates Learning Rate Regularity at Scale LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:20:08.660634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T15:17:21.305341Z digest=sha256:badbcec4399ca391bf7a588a4f23c07ea03f2afdddbd118e64e95bafadca60c9

Observation 894f4d4a-d92d-4f67-8db3-f8d7c631899c · inbound

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents cites this paper.

Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:24.964601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:00:45.789649Z digest=sha256:39c624debf6ae13a7e3ace1d58225d578ade102864b61033d17d380daaaafd19

Observation 14dd8efc-1881-4428-b542-536b9c5c141b · inbound

Mixed response geometry and critical crossover in the Ising model cites this paper.

Mixed response geometry and critical crossover in the Ising model LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T19:18:49.672005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:18:49.672005Z digest=sha256:3eacc6e69197327c3272f260f5b14628df7cb389610b8a8596b50bdbf35f6920

Observation 67dceee5-e2ae-4fd2-835b-083e57a810d2 · inbound

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility cites this paper.

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 221

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:25:52.555656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:21:07.188482Z digest=sha256:955e17f4c096834acb80994c59b976d079e2fae7c6acc462adc43c85db7cb1fd

Observation b4a85437-cc68-4617-a7b3-c9a9735460ec · inbound

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility cites this paper.

Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 221

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:36:21.798788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T04:31:41.220297Z digest=sha256:e679b5a0539de004a59439bdfe48eff6a3054dc675f42244963f68b1825418b3

Observation b62dad26-4a6c-4138-b313-3daa2778c68a · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T03:12:09.221341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T03:09:02.902912Z digest=sha256:d29c700a921a3494470d6a1d0fa938d4246bc815bfa2b2b187581955cf666ee1

Observation 7dd30433-18a3-46c7-8688-f37703eb8a72 · inbound

Instructions Shape Production of Language, not Processing cites this paper.

Instructions Shape Production of Language, not Processing LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:02:58.149627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-14T21:02:02.135970Z digest=sha256:e36e62d53d445869e91af72cce273d5cbb3000f0590bd0aaa5b92ec9b5bd86e3

Observation 4d5897e8-6118-4e7c-ad90-2babf4c7b44b · inbound

Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication cites this paper.

Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:05:00.239582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-30T19:04:36.443335Z digest=sha256:ed9e07a96aba88f8ed7f0a5fadd4ce18856a0c181fab534831bebd2ed2b23eb7

Observation f7eb2240-5837-4610-9687-d8cf96f1b79f · inbound

Natural Language-Focused Software Engineering via Code-Documentation Equivalence cites this paper.

Natural Language-Focused Software Engineering via Code-Documentation Equivalence LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-26T11:29:24.171382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T11:24:30.932755Z digest=sha256:22aadcfc4fdd894e18fc53991708212d500df301fb57f94713ae42877f453128

Observation 7aeec3ce-04c7-45d7-b70f-9361993fc159 · inbound

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models cites this paper.

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:55:56.897891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:55:56.897891Z digest=sha256:9fa6855aad89bc133a5dfbf23da002d9b669b361e533967236669f2ab81305b7