TRACES probes 30 LLMs with 42 unreliable papers and finds that models design follow-up studies for impossible premises in 93% of agentic attempts and 81% of interactive attempts.
Department of Energy labs embrace Genesis AI push.Science, 391(6791):1191–1192, 2026
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
TRACES probes 30 LLMs with 42 unreliable papers and finds that models design follow-up studies for impossible premises in 93% of agentic attempts and 81% of interactive attempts.