REVIEW 4 major objections 6 minor 1 cited by
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Choosing test items by metric variance, output diversity, or estimated difficulty reproduces full-set human rankings with ~70% of the annotations, even when selection uses only source texts.
desk verdict Useful empirical study of subset selection for human evaluation, but the abstract overclaims: the 70% figure is against random sampling, not full-data evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a set of item-utility functions that turn subset selection into a knapsack problem. The load-bearing pieces are: metric-score moments (average and variance across models), metric consistency measured by Spearman rank correlation between an item's model scores and the corpus-wide model scores, output diversity measured by negative average embedding similarity across model outputs, and a continuous Item Response Theory model, $$p(r_{m,x}=1)=\frac{c_x}{1+\exp[-a_x(\theta_m-b_x)]},$$ whose fitted discriminability $a_x$ and difficulty $b_x$ define the DiffDisc utility $a_x \times b_x$. For source-based selection, the paper distills these utilities into a regression model trained on source text alone, using the architecture of a learned quality-estimation metric, so item usefulness can be predicted before model outputs exist. Soft pairwise accuracy, which compares subset rankings with full-set rankings while accounting for statistical confidence, is the meta-evaluation that carries all comparisons.
What would settle it
Take the trained source-based selectors and apply them to a deliberately out-of-distribution test set, such as a new low-resource language pair or a domain absent from the training years, then compare the soft pairwise accuracy of the selected subsets with random sampling across budgets; if source-based selection stops beating or even matches random sampling, the transfer assumption is falsified.
Extended reading notes
Core claim
The paper's central claim is that item usefulness for human evaluation can be approximated by a few utility functions, and that choosing the top-scoring items under those utilities outperforms random sampling. Output-based selectors score an item by the variance of automated metric scores across models (MetricVar), by how well the item's metric-based model ranking matches the corpus-wide ranking (MetricCons), by the diversity of the model outputs (Diversity), or by the product of item-response-theory difficulty and discriminability, $a_x \times b_x$ (DiffDisc). In the WMT23 machine-translation experiments the best output-based selector reaches random sampling's soft pairwise accuracy with about 71 percent of the data, and the distilled source-based selector reaches it with about 79 percent; on SummEval the metric-consistency selector needs about 71 percent of human-annotated items and about 78--89 percent for LLM-based evaluation. The paper frames the problem as a 0-1 knapsack over item utilities, which is why cost-aware and document-aware variants follow naturally.
Load-bearing premise
The load-bearing premise is that item usefulness learned from past human evaluations transfers to a new evaluation's items; if the new test distribution shifts in language pair, genre, or difficulty, the source-based utility predictions and the claimed budget savings can break down.
Editorial extensions
If this is right
- Human evaluation of machine translation and summarization can be run on roughly 70 percent of the test items without degrading the model ranking relative to random sampling.
- Shared-task organizers can pre-select evaluation items from source texts alone, before participant models or their outputs are known, using distilled source-based estimators.
- The quality of the automated metric that drives selection matters: string-matching metrics give little or no gain over random, while stronger learned metrics translate directly into smaller required budgets.
- Accounting for per-item annotation cost during selection avoids the trap of choosing the hardest items, which tend to be the most expensive to annotate.
- The same selectors apply to LLM-as-a-judge evaluation, where each judgment also has a cost, reducing the number of outputs that need to be judged.
Reading between the lines
- The headline savings are measured against random sampling, not against the full-set ranking's absolute accuracy; if random sampling is already far from the full-set ranking, the correct reading is that the selectors match a cheaper-but-weak baseline, not that they recover the full evaluation.
- The source-based selectors inherit the training distribution's biases, and the paper itself notes they only work for items similar to the training data; a shift in genre, language pair, or item difficulty is the most likely failure mode in real deployment.
- A natural testable extension is to apply the diversity and metric-consistency selectors to other generation tasks with variable-length outputs, such as dialogue or code generation, where no ground-truth answer exists and human ranking is the target.
- Because source-based distillation is trained on past human evaluations, a feedback loop could emerge: if future test sets are built by these selectors, the pool of available training items shrinks and shifts, which would need to be monitored for drift.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses the problem of selecting a budget-limited subset of test items for human evaluation of NLG models, with the goal of preserving the model ranking obtained on the full set. It formalizes output-based selection (model outputs and automated metrics available) and source-based selection (only source texts available). The proposed output-based selectors use metric mean, variance, consistency, output diversity, and IRT difficulty-discriminability; source-based selectors are obtained by distilling these utilities into a predictor or by using an artificial crowd of models. The methods are evaluated on WMT19-23 machine translation data and on SummEval, using soft pairwise accuracy as the meta-metric. The central empirical claims are that several selectors outperform random selection and that roughly 70% of the test data suffices to reach the same evaluation result as random sampling.
Significance. If the empirical results hold, the paper offers a practical and reproducible toolkit (subset2evaluate) for reducing human-evaluation cost, and the source-based distillation idea is a useful contribution for settings where model outputs are unavailable at selection time. The paper is careful to evaluate against held-out human scores and to train source-based estimators only on data up to WMT22 when evaluating on WMT23. The reported gains over random selection are small (about 1-2 points of soft pairwise accuracy on average), but they are consistent across most language pairs in the aggregate plots. However, the headline cost-saving claim is overstated, and the aggregate averages hide per-language cases where the selectors do not outperform random; these issues need to be addressed before the practical claims can be accepted.
major comments (4)
- [Abstract; Section 5.4, Eq. (12), Table 5] The abstract's statement 'only ~70% of the test data is needed to produce the same evaluation result as the entire data' is not supported by the reported analysis. Equation (12) defines the budget C as the smallest budget at which a selector's soft pairwise accuracy (SPA) reaches the SPA of a random subset Y^R_B, not the SPA of the full set X. Figure 4 shows that random selection at the tested budgets has SPA around 91.6%, well below 100%, so matching random does not mean matching the full-data evaluation. Table 5 therefore reports budget savings relative to random sampling, not relative to evaluating the entire test set. The abstract, Section 5.4, and Section 8 should be rewritten to state this target explicitly, and ideally the authors should add a comparison against the full-set SPA or at least report the SPA of the full-data evaluation as a reference point.
- [Figure 4; Table 10] The claim that the selectors 'outperform random selection' is based on averages across languages and budgets, but the per-language results in Table 10 show many counterexamples. For example, WMT23 Zh-En MetricCons achieves 90.3% vs. random 91.4%, and WMT22 De-En all proposed methods are below random (e.g., 76.7% for MetricAvg vs. 80.2% random). The paper should either report a significance test across the 33 campaigns (e.g., a paired test on per-campaign SPA) or explicitly state that the improvement holds on average but not in every language pair. Without this, the blanket claim in the abstract and Section 1 is too strong.
- [Section 4.1; Figure 5; Section 1] The source-based distillation selectors are evaluated only on WMT23, and Section 4.1 itself states the method 'only works on novel items that are similar to those in the training data.' This is a serious limitation for the practical deployment the paper advocates, since shared tasks change language pairs and item distributions every year. The abstract and Section 1 present source-based selection as a demonstrated capability without this caveat. I recommend adding at least one held-out year (e.g., WMT24) or a cross-language-pair evaluation, and in any case carrying the caveat into the abstract and conclusions.
- [Table 5; Section 5.4] The cost-saving ratios in Table 5 are point estimates without confidence intervals, computed by comparing to a random baseline whose 90% confidence interval is about ±0.3 points of SPA (Figure 4). Since the selectors' advantage over random is only about 1-2 points on average, the ratio estimates (e.g., 71.4% for MetricCons) are sensitive to this uncertainty. The authors should provide bootstrap or jackknife confidence intervals for the ratios, or at least a sensitivity analysis with respect to the random baseline.
minor comments (6)
- [Section 2.1] The definition of soft pairwise accuracy should be given in the main text (or at least Eq. (17) from Appendix B should be referenced at the first mention), since it is the central meta-evaluation measure.
- [Table 5] The star markers (⋆) in the table are not explained in the caption or text; please clarify what they indicate.
- [Section 5.2] The cost-aware selection relies on a linear approximation of annotation time whose correlation with true time is only ρ=0.24; the sensitivity of Table 3 to this weak approximation should be discussed.
- [Section 4.2, Table 4] The artificial crowd evaluation scores are computed on M \ M′, and the random baseline in that table also excludes the crowd models; this setting should be described more prominently so readers do not compare Table 4 to Figure 5 directly.
- [Figures 4 and 5] The confidence intervals shown are only for the random baseline; displaying the spread of the deterministic methods across languages (e.g., as error bars) would make the significance of the differences easier to assess.
- [Section 1] The paper would benefit from a link to the subset2evaluate package in the published version; currently the footnote only promises release.
Circularity Check
No circularity: the selectors are meta-evaluated on held-out human scores, and source-based estimators are trained on WMT22-and-earlier data before being applied to WMT23.
full rationale
The central derivation chain is self-contained. Output-based selectors (MetricAvg, MetricVar, MetricCons, Diversity, DiffDisc) are computed from automated metric scores or output embeddings and are meta-evaluated against held-out human scores on WMT23/SummEval, so the target quantity (human ranking on the full set) is not an input to the selectors. Source-based estimators are trained only on WMT22 and earlier: Section 5.3 states 'We train all item utility estimator models by distilling utilities (Equation 10) only up to WMT22 to avoid contamination of the evaluation on WMT23,' and are then applied to WMT23; this is a genuine transfer, not a fitted-input-called-prediction loop. The IRT model is fit to metric scores (Section 3.4), and its difficulty/discrimination parameters are used only to rank items; the subsequent evaluation uses human scores, so there is no definitional equivalence. The only notable issue is the abstract's wording that roughly 70% of the data 'produce[s] the same evaluation result as the entire data,' whereas Equation (12) and Table 5 define the saving relative to random subset selection (SPA(Y^dagger) >= SPA(Y^R)), not to the full set; this is an overstatement of the empirical comparison, not a circular derivation. Self-citations (Zouhar et al. 2023, 2024a,b, 2025a,b) appear only in related-work or implementation context and do not carry the argument.
Assumptions & free parameters
free parameters (4)
- IRT item parameters (a_x, b_x, c_x) =
per-item latent values, not reported in main text
- human annotation time coefficients =
0.15 per source word and 33.7 intercept
- utility positivity shift =
not reported
- PreCOMET distillation weights =
trained on WMT22 and earlier human scores
assumptions (5)
- domain assumption Item utilities are additive: the utility of a subset is the sum of item utilities, and item interactions can be ignored.
- domain assumption Automated metrics are sufficiently aligned with human judgments for selection purposes.
- domain assumption Soft pairwise accuracy against the full-set human ranking is the right objective for 'same evaluation result'.
- domain assumption The IRT three-parameter logistic model (Equation 8) with normal priors appropriately describes how metric scores vary across models and items.
- domain assumption Source-based distillation transfers across years and language pairs.
Cite this review
Pith. "Pith review of How to Select Datapoints for Efficient Human Evaluation of NLG Models?." pith.science (2026). https://pith.science/paper/444NKONJ
@misc{pith2026250118251,
author = {Pith},
title = {Pith review of: How to Select Datapoints for Efficient Human Evaluation of NLG Models?},
year = {2026},
howpublished = {\url{https://pith.science/paper/444NKONJ}},
note = {Machine review of arXiv:2501.18251}
}
abstract
Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is often chosen in practice for human evaluation. However, randomly selected data may not accurately represent test performance, making this approach economically inefficient for model comparison. Thus, in this work, we develop and analyze a suite of selectors to get the most informative datapoints for human evaluation, taking the evaluation costs into account. We show that selectors based on variance in automated metric scores, diversity in model outputs, or Item Response Theory outperform random selection. We further develop an approach to distill these selectors to the scenario where the model outputs are not yet available. In particular, we introduce source-based estimators, which predict item usefulness for human evaluation just based on the source texts. We demonstrate the efficacy of our selectors in two common NLG tasks, machine translation and summarization, and show that only $\sim$70\% of the test data is needed to produce the same evaluation result as the entire data.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation
A supervised pairwise metric that predicts graded quality differences between two translations outperforms matched single-candidate QE baselines and larger QE models on WMT24.
Reference graph
Works this paper leans on
-
[1]
Ibrahim Said Ahmad, Antonios Anastasopoulos, Ond r ej Bojar, Claudia Borg, Marine Carpuat, Roldano Cattoni, Mauro Cettolo, William Chen, Qianqian Dong, Marcello Federico, Barry Haddow, D \'a vid Javorsk \'y , Mateusz Krubi \'n ski, Tsz Kin Lam, Xutai Ma, Prashant Mathur, Evgeny Matusov, Chandresh Maurya, John McCrae, Kenton Murray, Satoshi Nakamura, Matte...
2024
-
[2]
Shir Ashury Tahan, Ariel Gera, Benjamin Sznajder, Leshem Choshen, Liat Ein-Dor, and Eyal Shnarch. 2024. https://doi.org/10.18653/v1/2024.acl-long.456 Label-efficient model selection for text generation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8384--8402. Association for Computati...
-
[3]
Luca Benedetto, Andrea Cappelli, Roberto Turrin, and Paolo Cremonesi. 2020. https://dl.acm.org/doi/abs/10.1145/3375462.3375517 R2DE : A NLP approach to estimating IRT parameters of newly generated questions . In Proceedings of the tenth international conference on learning analytics & knowledge, 412--421
arXiv 2020
-
[4]
Matthew Byrd and Shashank Srivastava. 2022. https://doi.org/10.18653/v1/2022.acl-short.15 Predicting difficulty and discrimination of natural language questions . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 119--130. Association for Computational Linguistics
-
[5]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning at scale . In Proceedings of the 58th Annual Meeting of the Association for Comp...
-
[6]
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021. https://doi.org/10.1162/tacl_a_00417 A statistical analysis of summarization evaluation metrics using resampling methods . Transactions of the Association for Computational Linguistics, 9:1132--1146
-
[7]
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022. https://doi.org/10.18653/v1/2022.naacl-main.442 Re-examining system-level correlations of automatic summarization evaluation metrics . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 6038--6052. Association for...
-
[8]
Daniel Deutsch, George Foster, and Markus Freitag. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.798 Ties matter: Meta-evaluating modern metrics with pairwise accuracy and tie calibration . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12914--12929. Association for Computational Linguistics
Show all 64 references
-
[9]
Shachar Don-Yehiya, Leshem Choshen, and Omri Abend. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.767 P re Q u EL : Quality estimation of machine translation outputs in advance . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 111...
2022 doi
-
[10]
Fabbri, Wojciech Kry \'s ci \'n ski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev
Alexander R. Fabbri, Wojciech Kry \'s ci \'n ski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. https://doi.org/10.1162/tacl\_a\_00373 S umm E val: Re-evaluating summarization evaluation . Transactions of the Association for Computational Linguistics, ...
2021 doi
-
[11]
J \'u lia Falc \ a o, Claudia Borg, Nora Aranberri, and Kurt Abela. 2024. https://aclanthology.org/2024.lrec-main.315 COMET for low-resource machine translation evaluation: A case study of E nglish- M altese and S panish- B asque . In Proceedings of the 2024 Joint Internationa...
2024
-
[12]
Kehua Feng, Keyan Ding, Kede Ma, Zhihua Wang, Qiang Zhang, and Huajun Chen. 2024. http://arxiv.org/abs/2404.08008 Sample-efficient human evaluation of large language models via maximum discrepancy competition
2024 arXiv
-
[13]
Ronald A. Fisher. 1935. The Design of Experiments. Oliver & Boyd, Edinburgh
1935
-
[14]
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023. https://doi.org/10.18653/v1/2023.wmt-1.51 Results of ...
2023 doi
-
[15]
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and Andr \'e F. T. Martins. 2022. https://aclanthology.org/2022.wmt-1.2/ Results of WMT 22 metrics shared task: Stop using BLEU -- neural metrics...
2022
-
[16]
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ond r ej Bojar. 2021. https://aclanthology.org/2021.wmt-1.73/ Results of the WMT 21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news...
2021
-
[17]
Phillip Good. 2000. https://doi.org/https://doi.org/10.1007/978-1-4757-3235-1 Permutation Tests A Practical Guide to Resampling Methods for Testing Hypotheses . Springer
2000 doi
-
[18]
Yvette Graham, Timothy Baldwin, and Nitika Mathur. 2015. https://doi.org/10.3115/v1/N15-1124 Accurate evaluation of segment-level machine translation metrics . In Proceedings of the 2015 Conference of the North A merican Chapter of the Association for Computational Linguistics...
2015 doi
-
[19]
Q Huangfu and JAJ Hall. 2018. https://link.springer.com/article/10.1007/s12532-017-0130-5 Parallelizing the dual revised simplex method . Mathematical Programming Computation, 10(1):119--142
2018 doi
-
[20]
Ibrahim Jubran, Alaa Maalouf, and Dan Feldman. 2019. http://arxiv.org/abs/1910.08707 Introduction to coresets: Accurate coresets
2019 arXiv
-
[21]
Juraj Juraska, Daniel Deutsch, Mara Finkelstein, and Markus Freitag. 2024. http://arxiv.org/abs/2410.03983 Metricx-24: The google submission to the WMT 2024 metrics shared task
2024 arXiv
-
[22]
Richard M Karp, Raymond E Miller, and James W Thatcher. 1975. https://link.springer.com/chapter/10.1007/978-3-540-68279-0_8 Reducibility among combinatorial problems . Journal of Symbolic Logic, 40(4)
1975 doi
-
[23]
M. G. Kendall. 1938. http://www.jstor.org/stable/2332226 A new measure of rank correlation . Biometrika, 30(1/2):81--93
1938
-
[25]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, ...
2024 doi
-
[26]
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondrej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Kenton Murray, Masaaki Nagata, Martin Popel, Maj...
2024 arXiv
-
[27]
Tom Kocmi, Rachel Bawden, Ond r ej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Nov \'a k, Ma...
2022
-
[28]
Tom Kocmi and Christian Federmann. 2023. https://doi.org/10.18653/v1/2023.wmt-1.64 GEMBA - MQM : Detecting translation quality error spans with GPT -4 . In Proceedings of the Eighth Conference on Machine Translation, 768--775. Association for Computational Linguistics
2023 doi
-
[29]
Tom Kocmi, Vil \'e m Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popovi \'c , Mrinmaya Sachan, and Mariya Shmatova. 2024 d . https://doi.org/10.18653/v1/2024.wmt-1.131 Error span annotation: A balanced approach for human evaluation of machine tra...
2024 doi
-
[30]
John Patrick Lalor and Pedro Rodriguez. 2023. https://doi.org/10.1287/ijoc.2022.1250 py-irt : A scalable item response theory library for python . INFORMS Journal on Computing, 35(1):5–13
2023
-
[31]
John Patrick Lalor, Hao Wu, and Hong Yu. 2019. https://pmc.ncbi.nlm.nih.gov/articles/PMC6892593/ Learning latent parameters without human response patterns: Item response theory with artificial crowds . In Proceedings of the Conference on Empirical Methods in Natural Language ...
2019
-
[32]
Xiang Lisa Li, Farzaan Kaiyom, Evan Zheran Liu, Yifan Mai, Percy Liang, and Tatsunori Hashimoto. 2025. http://arxiv.org/abs/2407.08351 AutoBencher : Towards declarative benchmark construction
2025 arXiv
-
[33]
Yang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba, and Graham Horwood. 2024. http://arxiv.org/abs/2410.05952 Active evaluation acquisition for efficient LLM benchmarking
2024 arXiv
-
[34]
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.153 G -eval: NLG evaluation using gpt-4 with better human alignment . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language ...
2023 doi
-
[35]
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014. https://ddd.uab.cat/record/130144 Multidimensional quality metrics (MQM) : A framework for declaring and describing translation quality metrics . Tradum \`a tica , 12:0455--463
2014
-
[36]
Frederic M Lord and Melvin R Novick. 2008. https://psycnet.apa.org/record/1968-35040-000 Statistical theories of mental test scores . IAP
2008
-
[37]
V \^a nia Mendon c a, Ricardo Rei, Lu \'i sa Coheur, and Alberto Sardinha. 2023. https://doi.org/10.1162/coli\_a\_00473 Onception: Active learning with expert advice for real world machine translation . Computational Linguistics, 49(2):325--372
2023 doi
-
[38]
V \^a nia Mendon c a, Ricardo Rei, Luisa Coheur, Alberto Sardinha, and Ana L \'u cia Santos. 2021. https://doi.org/10.18653/v1/2021.acl-long.242 O nline L earning meets M achine T ranslation evaluation: Finding the best systems with the least human effort . In Proceedings of t...
2021 doi
-
[39]
Toshiaki Nakazawa and Isao Goto, editors. 2024. https://doi.org/10.18653/v1/2024.wat-1.0 Proceedings of the Eleventh Workshop on Asian Translation (WAT 2024) . Association for Computational Linguistics
2024 doi
-
[40]
Jinjie Ni, Fuzhao Xue, Xiang Yue, Yuntian Deng, Mahir Shah, Kabir Jain, Graham Neubig, and Yang You. 2024. http://arxiv.org/abs/2406.06565 MixEval : Deriving wisdom of the crowd from LLM benchmark mixtures
2024 arXiv
-
[41]
Yvonnick Noel and Bruno Dauvier. 2007. https://journals.sagepub.com/doi/abs/10.1177/0146621605287691 A beta item response model for continuous bounded responses . Applied Psychological Measurement, 31(1):47--73
2007 doi
-
[42]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 B leu: a method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, 311--318. As...
2002
-
[43]
Dongmin Park, Dimitris Papailiopoulos, and Kangwook Lee. 2022. https://openreview.net/forum?id=PAgpyQ5rGS Active learning is a strong baseline for data subset selection . In Has it Trained Yet? NeurIPS 2022 Workshop
2022
-
[44]
Stefano Perrella, Lorenzo Proietti, Alessandro Scir \`e , Edoardo Barba, and Roberto Navigli. 2024. https://doi.org/10.18653/v1/2024.acl-long.856 Guardians of the machine translation meta-evaluation: Sentinel metrics fall in! In Proceedings of the 62nd Annual Meeting of the As...
2024 doi
-
[45]
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024. http://arxiv.org/abs/2402.14992 tinyBenchmarks : Evaluating LLMs with fewer examples
2024 arXiv
-
[46]
Maja Popovi \'c . 2015. https://doi.org/10.18653/v1/W15-3049 chr F : character n-gram F -score for automatic MT evaluation . In Proceedings of the Tenth Workshop on Statistical Machine Translation, 392--395. Association for Computational Linguistics
2015 doi
-
[47]
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213 COMET : A neural framework for MT evaluation . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2685--2702. Associ...
2020 doi
-
[48]
Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence- BERT : Sentence embeddings using S iamese BERT -networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference...
2019 doi
-
[49]
Lalor, Robin Jia, and Jordan Boyd-Graber
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021. https://doi.org/10.18653/v1/2021.acl-long.346 Evaluation examples are not equally informative: How should that change NLP leaderboards? In Proceedings of the 59th Ann...
2021 doi
-
[50]
Jie Ruan, Xiao Pu, Mingqi Gao, Xiaojun Wan, and Yuesheng Zhu. 2024. http://arxiv.org/abs/2406.07967 Better than random: Reliable NLG human evaluation with constrained active sampling
2024 arXiv
-
[51]
Bel \'e n Sald \'i as Fuentes, George Foster, Markus Freitag, and Qijun Tan. 2022. https://doi.org/10.18653/v1/2022.humeval-1.7 Toward more effective human evaluation for machine translation . In Proceedings of the 2nd Workshop on Human Evaluation of NLP Systems (HumEval), 76-...
2022 doi
-
[52]
Darcy A Santor and James O Ramsay. 1998. https://psycnet.apa.org/record/1998-11993-004 Progress in the technology of measurement: Applications of item response models . Psychological assessment, 10(4):345
1998
-
[53]
Spearman
C. Spearman. 1904. http://www.jstor.org/stable/1412159 The proof and measurement of association between two things . The American Journal of Psychology, 15(1):72--101
1904
-
[54]
Brian Thompson, Nitika Mathur, Daniel Deutsch, and Huda Khayrallah. 2024. https://doi.org/10.18653/v1/2024.wmt-1.118 Improving statistical significance in human evaluation of automatic metrics via soft pairwise accuracy . In Proceedings of the Ninth Conference on Machine Trans...
2024 doi
-
[55]
Saeed Vahidian, Baharan Mirzasoleiman, and Alexander Cloninger. 2020 . https://proceedings.mlr.press/v124/vahidian20a.html Coresets for estimating means and mean square error with limited greedy samples . In Proceedings of the 36th Conference on Uncertainty in Artificial Intel...
2020
-
[56]
Rajan Vivek, Kawin Ethayarajh, Diyi Yang, and Douwe Kiela. 2024. https://aclanthology.org/2024.eacl-long.95/ Anchor points: Benchmarking models with much fewer examples . In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguis...
2024
-
[57]
Kai Wei, Rishabh Iyer, and Jeff Bilmes. 2015 . https://proceedings.mlr.press/v37/wei15.html Submodularity in data subset selection and active learning . In Proceedings of the 32nd International Conference on Machine Learning , volume 37 of Proceedings of Machine Learning Resea...
2015
-
[58]
Davis, Benjamin W
Mike Wu, Richard L. Davis, Benjamin W. Domingue, Chris Piech, and Noah Goodman. 2020. http://arxiv.org/abs/2002.00276 Variational item response theory: Fast, accurate, and expressive
2020 arXiv
-
[59]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. http://arxiv.org/abs/1904.09675 BERTScore : Evaluating text generation with BERT
2020 arXiv
-
[60]
Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daum \'e III, Kaheer Suleman, and Alexandra Olteanu. 2022. https://doi.org/10.18653/v1/2022.naacl-main.24 Deconstructing NLG evaluation: Evaluation practices, assumptions, and their implications . In Proceedings of the 2022 Co...
2022 doi
-
[61]
Vil \'e m Zouhar, Pinzhen Chen, Tsz Kin Lam, Nikita Moghe, and Barry Haddow. 2024 a . https://doi.org/10.18653/v1/2024.wmt-1.121 Pitfalls and outlooks in using COMET . In Proceedings of the Ninth Conference on Machine Translation, 1272--1288. Association for Computational Linguistics
2024 doi
-
[62]
Vil \'e m Zouhar, Shehzaad Dhuliawala, Wangchunshu Zhou, Nico Daheim, Tom Kocmi, Yuchen Eleanor Jiang, and Mrinmaya Sachan. 2023. https://doi.org/10.18653/v1/2023.eacl-main.95 Poor man`s quality estimation: Predicting reference-based MT metrics without the reference . In Proce...
2023 doi
-
[63]
Vil \'e m Zouhar, Shuoyang Ding, Anna Currey, Tatyana Badeka, Jenyuan Wang, and Brian Thompson. 2024 b . https://doi.org/10.18653/v1/2024.acl-short.45 Fine-tuned machine translation metrics struggle in unseen domains . In Proceedings of the 62nd Annual Meeting of the Associati...
2024 doi
-
[64]
Vil \'e m Zouhar, Tom Kocmi, and Mrinmaya Sachan. 2025 a . https://aclanthology.org/2025.naacl-long.255/ AI -assisted human evaluation of machine translation . In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin...
2025
-
[65]
Vilém Zouhar, Maike Züfle, Beni Egressy, Julius Cheng, and Jan Niehues. 2025 b . http://arxiv.org/abs/2502.14429 Early-exit and instant confidence translation quality estimation
2025 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.