Pith. sign in

REVIEW 2 cited by

GENIE: Toward Reproducible and Standardized Human Evaluation for Text Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.06561 v4 pith:LIGLZA36 submitted 2021-01-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords geniegenerationhumandifferenttextevaluationevaluationsstandardized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While often assumed a gold standard, effective human evaluation of text generation remains an important, open area for research. We revisit this problem with a focus on producing consistent evaluations that are reproducible -- over time and across different populations. We study this goal in different stages of the human evaluation pipeline. In particular, we consider design choices for the annotation interface used to elicit human judgments and their impact on reproducibility. Furthermore, we develop an automated mechanism for maintaining annotator quality via a probabilistic model that detects and excludes noisy annotators. Putting these lessons together, we introduce GENIE: a system for running standardized human evaluations across different generation tasks. We instantiate GENIE with datasets representing four core challenges in text generation: machine translation, summarization, commonsense reasoning, and machine comprehension. For each task, GENIE offers a leaderboard that automatically crowdsources annotations for submissions, evaluating them along axes such as correctness, conciseness, and fluency. We have made the GENIE leaderboards publicly available, and have already ranked 50 submissions from 10 different research groups. We hope GENIE encourages further progress toward effective, standardized evaluations for text generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating the Evaluators: Are readability metrics good measures of readability?

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Classic readability formulas correlate weakly with human readability judgments for plain-language summaries (FKGL r=0.16), the best LLM judge reaches r=0.56, and the two evaluator families rank datasets nearly opposite.

  2. Min-p, Max Exaggeration: A Critical Analysis of Min-p Sampling in Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A comprehensive reanalysis finds that min-p sampling does not outperform top-p, top-k, or basic sampling once the original data are re-tested and hyperparameter budgets are equalized.

Pith tools