Pith. sign in

REVIEW 1 cited by

Evaluation of Question Answering Systems: Complexity of judging a natural language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.12617 v1 pith:7ERZDXYL submitted 2022-09-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords systemsevaluationsystemansweringbeenembeddingsimportantlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Question answering (QA) systems are among the most important and rapidly developing research topics in natural language processing (NLP). A reason, therefore, is that a QA system allows humans to interact more naturally with a machine, e.g., via a virtual assistant or search engine. In the last decades, many QA systems have been proposed to address the requirements of different question-answering tasks. Furthermore, many error scores have been introduced, e.g., based on n-gram matching, word embeddings, or contextual embeddings to measure the performance of a QA system. This survey attempts to provide a systematic overview of the general framework of QA, QA paradigms, benchmark datasets, and assessment techniques for a quantitative evaluation of QA systems. The latter is particularly important because not only is the construction of a QA system complex but also its evaluation. We hypothesize that a reason, therefore, is that the quantitative formalization of human judgment is an open problem.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators

    cs.CL 2025-07 conditional novelty 5.0 of 10

    In the ELOQUENT 2025 Sensemaking task, LLM-based evaluators rated clearly garbled or mismatched question-answer pairs as acceptable, showing that LLM-as-a-Judge scores cannot be trusted.

Pith tools