HEDS 3.0 is a revised human-evaluation datasheet with a longer verbatim form, two new intra-annotator agreement questions, and a latex export tool for reproducibility reporting.
The Human Evaluation Datasheet 1.0: A Template for Recording Details of Human Evaluation Experiments in NLP
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper introduces the Human Evaluation Datasheet, a template for recording the details of individual human evaluation experiments in Natural Language Processing (NLP). Originally taking inspiration from seminal papers by Bender and Friedman (2018), Mitchell et al. (2019), and Gebru et al. (2020), the Human Evaluation Datasheet is intended to facilitate the recording of properties of human evaluations in sufficient detail, and with sufficient standardisation, to support comparability, meta-evaluation, and reproducibility tests.
fields
cs.HC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
HEDS 3.0: The Human Evaluation Data Sheet Version 3.0
HEDS 3.0 is a revised human-evaluation datasheet with a longer verbatim form, two new intra-annotator agreement questions, and a latex export tool for reproducibility reporting.