REVIEW 2 cited by
A Qualitative Comparison of CoQA, SQuAD 2.0 and QuAC
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We compare three new datasets for question answering: SQuAD 2.0, QuAC, and CoQA, along several of their new features: (1) unanswerable questions, (2) multi-turn interactions, and (3) abstractive answers. We show that the datasets provide complementary coverage of the first two aspects, but weak coverage of the third. Because of the datasets' structural similarity, a single extractive model can be easily adapted to any of the datasets and we show improved baseline results on both SQuAD 2.0 and CoQA. Despite the similarity, models trained on one dataset are ineffective on another dataset, but we find moderate performance improvement through pretraining. To encourage cross-evaluation, we release code for conversion between datasets at https://github.com/my89/co-squac .
Forward citations
Cited by 2 Pith papers
-
Attentive History Selection for Conversational Question Answering
A BERT-based model with position-aware history answer embeddings and a learned history attention mechanism improves QuAC F1 by about one point over strong baselines, but multi-task learning with dialog acts does not h...
-
FlowDelta: Modeling Flow Information Gain in Reasoning for Conversational Machine Comprehension
Modeling the difference between consecutive reasoning states, called FlowDelta, improves conversational machine comprehension accuracy across FlowQA and BERT on CoQA, QuAC, and SCONE.
Discussion (0). Continue with ORCID to comment.