On 1,847 COVID-19 tweets, GPT-4 scored highest at detecting and classifying scientific claims (F1 0.65-0.76), but the evaluation lacks error bars, baselines, and open artifacts.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
REJECT 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Evaluating the Performance of Large Language Models in Scientific Claim Detection and Classification
On 1,847 COVID-19 tweets, GPT-4 scored highest at detecting and classifying scientific claims (F1 0.65-0.76), but the evaluation lacks error bars, baselines, and open artifacts.