Across SQuAD and QuAC, RoBERTa and DistilBERT are generally more stable under Monte Carlo dropout, while ALBERT and BERT-Base are less consistent, and paraphrase perturbation changes the ranking.
A deep network model for paraphrase detection in short text 29 messages.Information Processing & Management, 54(6):922–937, 2018
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Assessing Reliability of BERT-Based Models on Question Answering Tasks
Across SQuAD and QuAC, RoBERTa and DistilBERT are generally more stable under Monte Carlo dropout, while ALBERT and BERT-Base are less consistent, and paraphrase perturbation changes the ranking.