In 50 LLM measurement tasks from 27 top-journal papers, LLM outputs are often central to claims yet validation is limited, mostly convergent, and frequently incomplete.
Messick, Validity and washback in language testing.Language testing13(3), 241–256 (1996)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2026 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Validating LLMs in social science: Epistemic threats and emerging norms
In 50 LLM measurement tasks from 27 top-journal papers, LLM outputs are often central to claims yet validation is limited, mostly convergent, and frequently incomplete.