SciFaultyQA is a 1,333-question benchmark of intentionally faulty science questions on which GPT-4o detects only 16% of faults, improved to 65% with web search integration.
AI now beats humans at basic tasks — new benchmarks are needed, says major report
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation
SciFaultyQA is a 1,333-question benchmark of intentionally faulty science questions on which GPT-4o detects only 16% of faults, improved to 65% with web search integration.