TRACES probes 30 LLMs with 42 unreliable papers and finds that models design follow-up studies for impossible premises in 93% of agentic attempts and 81% of interactive attempts.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.IR 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
TRACES probes 30 LLMs with 42 unreliable papers and finds that models design follow-up studies for impossible premises in 93% of agentic attempts and 81% of interactive attempts.