REVIEW 2 cited by
Towards a Better Metric for Evaluating Question Generation Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
There has always been criticism for using $n$-gram based similarity metrics, such as BLEU, NIST, etc, for evaluating the performance of NLG systems. However, these metrics continue to remain popular and are recently being used for evaluating the performance of systems which automatically generate questions from documents, knowledge graphs, images, etc. Given the rising interest in such automatic question generation (AQG) systems, it is important to objectively examine whether these metrics are suitable for this task. In particular, it is important to verify whether such metrics used for evaluating AQG systems focus on answerability of the generated question by preferring questions which contain all relevant information such as question type (Wh-types), entities, relations, etc. In this work, we show that current automatic evaluation metrics based on $n$-gram similarity do not always correlate well with human judgments about answerability of a question. To alleviate this problem and as a first step towards better evaluation metrics for AQG, we introduce a scoring function to capture answerability and show that when this scoring function is integrated with existing metrics, they correlate significantly better with human judgments. The scripts and data developed as a part of this work are made publicly available at https://github.com/PrekshaNema25/Answerability-Metric
Forward citations
Cited by 2 Pith papers
-
Let's Ask Again: Refine Network for Automatic Question Generation
A two-pass refinement decoder with dual attention improves automatic question generation over single-pass models on SQuAD, HOTPOT-QA, and DROP.
-
Reinforcement Learning Based Graph-to-Sequence Model for Natural Question Generation
A reinforcement-learning graph-to-sequence model with answer-aware alignment reports new state-of-the-art question generation scores on SQuAD, with the gain partly explained by BERT embeddings and direct BLEU-4 optimization.
Discussion (0). Continue with ORCID to comment.