Pith. sign in

REVIEW 1 cited by

A Gold Standard Methodology for Evaluating Accuracy in Data-To-Text Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.03992 v1 pith:WAY6RWYN submitted 2020-11-08 cs.CL

classification cs.CL
keywords accuracymethodologysystemsdata-to-textevaluationgeneratedgoldstandard
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Most Natural Language Generation systems need to produce accurate texts. We propose a methodology for high-quality human evaluation of the accuracy of generated texts, which is intended to serve as a gold-standard for accuracy evaluations of data-to-text systems. We use our methodology to evaluate the accuracy of computer generated basketball summaries. We then show how our gold standard evaluation can be used to validate automated metrics

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Multimodal LLMs frequently accept hallucinated emotion claims on a new adversarial benchmark, with the worst failures on image, audio, and video perception rather than on textbook emotion knowledge.

Pith tools