Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

cs.CL · 2025-09-01 · conditional · novelty 6.0

Using LLM-as-a-Judge grading instead of heuristic answer matching dramatically reduces measured prompt sensitivity and stabilizes model rankings across prompt templates, with human annotations confirming the judge's view.

citing papers explorer

Showing 1 of 1 citing paper.

  • Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs cs.CL · 2025-09-01 · conditional · none · ref 4

    Using LLM-as-a-Judge grading instead of heuristic answer matching dramatically reduces measured prompt sensitivity and stabilizes model rankings across prompt templates, with human annotations confirming the judge's view.