EQMs, sixty LLM-scored reasoning patterns, predict forecast accuracy at both item and person levels and outperform prior text-analysis methods in a large pre-registered tournament dataset.
Rathje,et al., GPT is an effective tool for multilingual psychological text analysis
5 Pith papers cite this work, alongside 235 external citations. Polarity classification is still indexing.
years
2026 5representative citing papers
AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
In 50 LLM measurement tasks from 27 top-journal papers, LLM outputs are often central to claims yet validation is limited, mostly convergent, and frequently incomplete.
Grain calibration decomposes theoretical constructs into clause-level components, tests each with extractive evidence, and combines results through explicit theory-derived rules to validate LLM coding beyond agreement with human annotators.
Semantic mapping of 8,954 definitions and 2,700 scales from 14,000+ papers shows learner agency and autonomy span task regulation, personal motivation, and sociocultural dimensions, with existing scales and generative AI research underrepresenting the sociocultural dimension.
citing papers explorer
-
Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments
EQMs, sixty LLM-scored reasoning patterns, predict forecast accuracy at both item and person levels and outperform prior text-analysis methods in a large pre-registered tournament dataset.
-
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
AMALIA matches larger models on agreement with human coders for moral-authority annotation, yet its recovery gap shows most of that performance is not reproduced by the theory-defined route—revealing a validity shortfall invisible to agreement metrics.
-
Validating LLMs in social science: Epistemic threats and emerging norms
In 50 LLM measurement tasks from 27 top-journal papers, LLM outputs are often central to claims yet validation is limited, mostly convergent, and frequently incomplete.
-
Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
Grain calibration decomposes theoretical constructs into clause-level components, tests each with extractive evidence, and combines results through explicit theory-derived rules to validate LLM coding beyond agreement with human annotators.
-
Large-scale semantic mapping of learner agency and autonomy reveals what measurement and generative AI research overlook
Semantic mapping of 8,954 definitions and 2,700 scales from 14,000+ papers shows learner agency and autonomy span task regulation, personal motivation, and sociocultural dimensions, with existing scales and generative AI research underrepresenting the sociocultural dimension.