On SQuAD2, fine-tuned RoBERTa beats all tested LLMs, but LLaMA-3.1-70B outperforms the fine-tuned models on 3 of 5 out-of-distribution QA datasets.
Edg-based question decomposition for complex question answering over knowledge bases
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Question: How do Large Language Models perform on the Question Answering tasks? Answer:
On SQuAD2, fine-tuned RoBERTa beats all tested LLMs, but LLaMA-3.1-70B outperforms the fine-tuned models on 3 of 5 out-of-distribution QA datasets.