{"id":"421e8dd0-b032-4ed8-a43e-a7d90637d7b5","arxiv_id":"2608.11415","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"TRACES probes 30 LLMs with 42 unreliable papers and finds that models design follow-up studies for impossible premises in 93% of agentic attempts and 81% of interactive attempts.","lead":"This paper introduces TRACES, a benchmark that feeds language models near-verbatim text from 42 retracted or pseudoscientific papers and asks them to design follow-up studies. Most models do so happily, suggesting they lack the epistemic judgment needed to act as unsupervised scientific agents.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Construct-validity assumption that every probe is answerable only by accepting the flawed premise is not independently audited; if some requests admit legitimate hypothetical or critical responses, IFR-a overstates epistemic failure.","rationale":"The reader's weakest assumption and my independent read converge: the validity of IFR-a rests on the claim that the 42 operational requests have no scientifically legitimate non-refusal answer except premise acceptance. This is asserted in §2.1–§2.3 and illustrated for one probe, but not independently established for the corpus. The paper has real strengths that I do not want to discount: 12,600 responses across 30 models, a frozen deterministic scorer with human validation (96/96 on REFUSED/ENGAGED, 94/96 on RECOGNIZED), a stability sweep, the Fable boundary-case analysis, and an honest Limitations section. None of those strengths protects the construct assumption, because the scorer validates labels, not the premise-boundary design. I also noted the abstract's matched-structure-control claim, which is not backed by a reported control in the full text, and the Fable exclusion vs. the 'every evaluated model' phrasing; both are secondary to the construct concern and would be addressed by the same careful reporting. If the proposed expert audit shows that all or nearly all probes are premise-required, the central claim stands; if it shows a material fraction are not, the headline failure rates need to be recomputed and the interpretation softened. That is exactly the condition the reader already attached, so I recommend no change to the CONDITIONAL verdict.","tokens_in":32461,"tokens_out":9624,"duration_ms":85856,"concrete_test":"For each of the 42 probes, have 3–5 domain experts per probe independently answer, blinded to the paper's expected boundary: 'Can a scientifically valid response satisfy this request without adopting the paper's central premise?' For any 'yes', have experts write such a response (falsification protocol, hypothetical with explicit caveat, or critical analysis). Run the frozen deterministic scorer on those expert-written non-accepting responses. If any are labeled ENGAGED, or experts judge more than, say, 5 of 42 probes answerable without premise acceptance, recompute aggregate IFR-a excluding those probes. A material drop (e.g., below 0.85) or a change in model ordering would require the central claim to be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline IFR-a numbers treat every non-refusal as engagement with an untenable premise. That inference depends on the construct assumption in §2.3: each operational request is designed so that the model has an opportunity to refuse, and more strongly so that no scientifically valid response exists unless the model adopts the paper's flawed premise. The paper offers one worked example (Staker, §2.3) and author-run development-panel checks, but it does not report an independent, blinded audit of all 42 probe boundaries. The distinction is not academic: a request to 'design a setup to observe excess heat' can be answered by a falsification-oriented protocol that tests the claim rather than accepting it; a hypothetical framing can elicit a conditional design with an explicit caveat; a critical-analysis response can use the paper's vocabulary while rejecting its ontology. The frozen scorer's REFUSED/RECOGNIZED axes cannot separate these from premise acceptance: generic disclaimers are demoted, and any substantive study-design content is scored as ENGAGED. Because aggregate IFR-a is 0.93, even a modest number of probes with a legitimate non-accepting answer would shift the headline number and, more importantly, invalidate the interpretation that engagement equals epistemic endorsement. The paper's own §2.4 concedes that 'specific operational requests may deviate from these templates as the source material requires,' and §2.3 promises a construct-scope discussion in §2.6 that §2.6 does not actually provide. The human scorer validation confirms label consistency, not the premise-boundary annotation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TRACES, a benchmark and methodology for measuring whether LLMs engaged as scientific agents can recognize and refuse to build on unreliable scientific literature. A corpus of 42 probes derived from retracted, fraudulent, and pseudoscientific papers pairs near-verbatim preambles with operational study-design requests; responses are scored on two binary axes (REFUSED vs. ENGAGED, RECOGNIZED vs. UNRECOGNIZED) that yield two influence-failure rates, IFR-a (agentic: only refusal passes) and IFR-i (interactive: refusal or content-specific recognition passes), plus an Engagement Depth Index (EDI) for engaged responses. The evaluation covers 30 models over 10 repeated runs, reporting aggregate IFR-a of 0.93 ± 0.004 and IFR-i of 0.809 ± 0.009, with 22 of 30 models failing more than 90% of agentic probes. Refusals are found to be concentrated on a small number of high-notoriety or lexically distinctive probes, and the authors interpret this as topic-keyed safety behavior rather than robust epistemic competence. The paper concludes that no model in the panel refuses often enough for safe unsupervised deployment and calls for credibility-annotation infrastructure.","tokens_in":32730,"tokens_out":5881,"duration_ms":53555,"significance":"If the central construct holds, TRACES addresses a genuine and previously under-measured failure mode: single-shot agentic behavior on unreliable scientific premises, as opposed to factuality on questions with known answers. The released corpus, deterministic scoring harness, and complete run artifacts are valuable concrete contributions, and the distinction between agentic (IFR-a) and interactive (IFR-i) failure, with the 'disclaimer tax' gap, is a useful analytic frame. The qualitative finding that models overwhelmingly engage with untenable premises in a plausible research context is robust to most of the reviewer concerns and is itself an important result. The main risks are construct validity (whether every non-refusal really constitutes premise acceptance) and the support for several headline claims that depend on controls and exclusions not fully described in the main text.","major_comments":[{"comment":"The central inference that engagement equals premise acceptance rests on the claim that each operational request has no scientifically legitimate answer unless the model adopts the flawed premise, but the promised construct-scope discussion is missing. Section 2.3 states that \"Every TRACES probe contains at least three requests that require accepting the flawed premise\" and refers the reader to §2.6 for the construct-scope issue; §2.6, however, is devoted to EDI and does not audit or enumerate the answerability boundaries of the 42 probes. Section 2.4 adds that \"specific operational requests may deviate from these templates as the source material requires.\" Because a request such as designing a setup to observe \"excess heat\" could in principle be answered by a falsification-oriented protocol or an explicitly hypothetical design, an independent, blinded audit of each probe's boundary (or a clearly stated sampling plan) is needed to show that no legitimate non-accepting answer exists.","section":"§2.3, §2.6"},{"comment":"The headline claim that \"every evaluated model fails more than 71% of agentic probes\" and the aggregate IFR-a figures are computed after excluding Fable, the only model in the paper with an IFR-a below 0.71 (its IFR-a is 0.214 under the paper's counting convention). The exclusion is defensible because Fable's empty completions reflect a blanket content gate rather than reliability judgment, but it must be disclosed wherever the headline numbers appear. As written, the abstract implies universality over all models the paper evaluates, which is contradicted by the separately reported Fable results.","section":"Abstract, §3.3"},{"comment":"The abstract's claim that rejections \"disappear under matched-structure controls\" is a load-bearing piece of evidence for the topic-keyed-safety conclusion, yet no matched-structure control design, procedure, or quantitative result appears in §2, §3, §4, or the appendices. The closest material is the Wakefield/Epel, LK-99/Dias/holey-graphyne, and Bem/Persinger/Cohen comparisons in §3.1 and §3.2; if these are intended as the controls, the matching criteria and the comparison must be described and reported. Otherwise the sentence should be removed or explicitly labeled as a hypothesis.","section":"Abstract, §3.1, §3.2"},{"comment":"The validation of the frozen scorer is thin relative to the load it carries. Human validation covers 96 responses from 3 models with a single annotator, and the blind LLM panel audit covers 18 weakly scored rows with exact four-class agreement of 6/18. The reported disagreement asymmetry (the scorer under-credits recognition rather than over-credits it) is reassuring for IFR-a, but agreement of this size does not adequately support the boundary decisions that separate hypothetical, critical, or falsification-oriented responses from true premise acceptance. The manuscript should report inter-annotator agreement on a stratified sample spanning all 42 probes and describe how each probe's construct boundary was independently reviewed.","section":"§4, Human scorer validation and LLM panel audit"}],"minor_comments":[{"comment":"There is a typo in \"V oight-Kampff test\": the space inside the name should be removed.","section":"§2.3"},{"comment":"The paper alternates between \"Fable\" and \"Fable 5\" for the same model; a single name should be used, and the text should state explicitly that Fable is not one of the 30 models in Appendix B.","section":"§3.1, §3.3"},{"comment":"The audit section mentions \"ATLAS ontology annotations\" without defining ATLAS or providing a citation; this should be clarified.","section":"§4, LLM panel audit"},{"comment":"The 200-character gate for EDI is asserted without justification; a sentence explaining why reproduction is not meaningful below that length would help readers calibrate the metric.","section":"§2.5"},{"comment":"In the Schön discussion, the quotation \"legitimate'\" contains a stray apostrophe inside the closing quotation mark.","section":"§3.3"},{"comment":"The paper reports temperature 1.0 and seeds 1-10 \"where honored\" but does not state which providers or endpoints ignored the seed parameter; this is relevant to interpreting the stability statistics.","section":"§4, Model panel"}],"recommendation":"major_revision","confidential_remarks":"The benchmark contribution is real and the qualitative engagement result is likely robust, but the manuscript currently overstates the universality of the failure rates and under-supports the matched-structure-control narrative. The construct-validity audit is essential; without it, the precise meaning of the 93% IFR-a number remains uncertain. All of the major concerns are addressable within the manuscript's scope, so rejection is not warranted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth reading. TRACES measures a genuine gap in LLM evaluation: whether models act on unreliable scientific literature when asked to extend it, not just whether they answer known-answer questions correctly. The probe design is thoughtful—near-verbatim preambles, operational requests with an explicit exit, five claim types, two complementary failure metrics, a depth index, and a frozen deterministic scorer validated against human labels. The central qualitative finding is robust: across 30 models and 420 prompts each, models engage with obviously untenable premises at near-total rates, and refusals concentrate on notorious papers like Wakefield while structurally comparable but less famous papers get zero refusals. That pattern appears in raw refusal counts, not just in the aggregate metrics.\n\nThe main soft spot is the premise-boundary assumption behind the 0.93 IFR-a headline. The paper treats every non-refusal as engagement with the flawed premise. That's plausible for requests like 'suggest a number of Biofield Energy Healers and treatment time,' but for the full corpus, the paper does not provide an independent, blinded audit showing all 42 probes have only one legitimate answer. Some requests could be answered as a falsification protocol or explicit hypothetical without endorsing the paper. The authors say they checked this during development and promise a construct-scope discussion in §2.6 that §2.6 doesn't actually give. Even a few such probes would shift the aggregate, though not the qualitative conclusion. This is an addressable validity issue, not a fatal one.\n\nMinor issues: the abstract's 'every evaluated model fails more than 71%' silently excludes Fable, which is disclosed and justified in §3.3 but should be stated in the abstract. The matched-structure control claim isn't backed by a corpus-wide systematic control. The release artifacts are described but not linked, which matters for a benchmark paper.\n\nWho this is for: anyone evaluating LLMs for scientific or clinical deployment, or building credibility guardrails. It deserves a serious referee; the central finding is important even if the headline number gets qualified, and the construct-validity concern is fixable with an independent probe audit. I'd bring it to reading group mostly to discuss how much weight to put on IFR-a as a safety metric.","headline":"TRACES measures a real failure mode—models build study designs on retracted or incoherent premises—but the 93% engagement number needs an independent audit of probe boundaries.","tokens_in":33290,"tokens_out":3437,"would_cite":true,"duration_ms":32336,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 30 models and 42 flawed papers, language models design follow-up research on false premises in 95% of non-empty responses.","keywords":["epistemic reliability","scientific reasoning","large language models","influence failure rate","retracted papers","pseudoscience","agentic science","benchmark"],"falsifier":"Ask independent domain experts whether each of the 42 operational requests admits a legitimate non-accepting answer, such as an explicitly hypothetical study design or a critical analysis of the premise; if even a handful do, the engagement rate would no longer measure premise acceptance. A quicker check is to run the same probes against matched control papers with sound premises — the paper reports that refusals disappear under such controls, so if control refusals rose, the interpretation would collapse.","tokens_in":32238,"feed_emoji":"🧪","tokens_out":6974,"duration_ms":70554,"temperature":0.7,"pith_summary":"TRACES tries to establish that large language models, when handed the framing of a retracted, fraudulent, or pseudoscientific paper in a plausible research context, will mostly design follow-up work on top of the false premise instead of refusing. Across 30 models and 10 repeated runs on a 42-probe corpus, only 7% of responses refused outright, giving an aggregate agentic failure rate of 0.93; even after crediting content-specific warnings, the interactive failure rate is 0.809. The paper argues this is not a knowledge defect that more data fixes but a behavioral pattern: refusals concentrate on a handful of famous or flamboyantly pseudoscientific topics and disappear on structurally similar but obscure papers. If the result holds, no current model is safe to deploy as an unsupervised scientific research agent.","feed_headline":"LLMs design studies on false science 93% of the time","feed_subtitle":"A 42-probe test across 30 models finds refusals are rare, topic-specific, and absent on obscure but equally unsound papers.","key_machinery":"The instrument is the TRACES probe: a near-verbatim preamble from an unreliable paper, a plausible study-design request with no legitimate answer unless the flawed premise is accepted, and withheld paper-specific details scored by level. The central metrics are the two Influence Failure Rates — IFR-a (agentic: refusal is the only pass) and IFR-i (interactive: engagement with a content-specific warning also passes) — plus the Engagement Depth Index (EDI), which measures how much paper- or field-specific withheld vocabulary the model reproduces. The cross-tabulation of REFUSED/RECOGNIZED signals is what lets the paper separate 'safe because it reasoned' from 'safe because a classifier blitzed on a notorious topic,' and the gap between IFR-a and IFR-i quantifies the disclaimer tax.","core_discovery":"The paper's central claim is that scientific reasoning benchmarks that score correct answers cannot see the failure mode that matters for agentic science: a model that fluently adopts and extends an unsound premise. The authors build a probe corpus from 42 unreliable papers, each probe pairing near-verbatim text from the paper with a first-person request to design a follow-up study whose every reasonable path requires accepting the flawed premise. They measure two failure rates: IFR-a, where only outright refusal counts as a pass, and IFR-i, where engagement after a content-specific warning also passes. Across the panel, aggregate IFR-a is 0.93 ± 0.004 and IFR-i is 0.809 ± 0.009, with 95% of non-empty responses engaging with untenable premises; every model fails more than 71% of agentic probes and 22 of 30 fail more than 90%. The authors conclude that the observed refusals are consistent with topic-keyed safety filters rather than epistemic competence, and that credibility assessment must become shared infrastructure for scientific deployment of language models.","pith_inferences":["Because the refusals are topic-keyed, a benchmark that samples famous retractions will overestimate epistemic reliability; the same models would likely engage with unfamiliar but equally unsound papers, so TRACES' mix of notoriety levels is what makes its aggregate meaningful.","A natural follow-up the authors leave open is prompt-level mitigation: adding an explicit instruction to assess source reliability could change IFR-a, and measuring whether the improvement is uniform across claim types would separate guardrail effects from genuine reasoning.","The single-shot design suggests a multi-turn test: after a model refuses or blanks, re-asking with rephrased framing would reveal whether the refusal is a stable epistemic decision or a shallow filter that can be talked past.","A direct deployment test would give a retrieval agent the retraction status of each target paper at inference time; if IFR-a barely moves despite the retrieved evidence, then credibility signals must be enforced by guardrails rather than left to the model's judgment."],"forward_implications":["No model in the 30-model panel refused often enough to be deployed as an unsupervised research agent; even the best interactive models warned users in less than half of responses.","A model that produces a 'rigorous double-blind trial of homeopathy' has not produced a safe agentic output, so downstream agents consuming response bodies inherit the false premise regardless of disclaimers.","Refusals that do occur cluster on a few high-notoriety topics and largely disappear on structurally similar but obscure papers, implying current safety is topic-keyed rather than epistemic.","EDI results indicate contamination is mostly field-level: models across all sizes reproduce field vocabulary, and larger models do so more fluently, so the failure is not fixed by model scale alone.","Because 81% of responses contained no warning at all, the paper concludes that credibility annotations and retrieval-based credibility checking are needed infrastructure for scientific LLM deployment."],"supporting_citations":[{"why":"Supplies the open-ended task gap the paper targets: a leading model scored 77% on structured science questions but 25% on open-ended ones.","marker":"[15]"},{"why":"Supplies evidence that reworded logic puzzles collapse accuracy, motivating the claim that answer-centric benchmarks measure recall rather than reasoning.","marker":"[17]"},{"why":"The MMR-autism retraction is a corpus probe whose refusal pattern is analyzed as evidence of a paper-specific classifier rather than epistemic reasoning.","marker":"[20]"},{"why":"The cold-fusion palladium paper is the worked example of probe construction, showing how operational requests become premise-locked.","marker":"[21]"},{"why":"The original cold-fusion electrolysis report supplies the 'excess heat' vocabulary that anchors the first refusal point in the worked example.","marker":"[22]"},{"why":"The telomere-stress study is the methodological control that draws no refusals, used to argue that Wakefield refusals are notoriety-based rather than reasoning-based.","marker":"[28]"},{"why":"The retracted tracheal-transplant study is a corpus probe demonstrating sanewashing, where a model cites the fraudulent early cases as successful implants.","marker":"[29]"},{"why":"The LK-99 preprint is a corpus probe in the notoriety gradient, drawing far more refusals than equally flawed but less famous retractions.","marker":"[32]"},{"why":"The clinical study on alternative-medicine survival supplies the harm evidence for why engaging with unsound protocols can cost lives.","marker":"[50]"}],"fun_headline_variants":["LLMs engage with false science in 95% of responses","Scientific integrity: 22 of 30 LLMs fail 90% of probes","LLMs rarely refuse pseudoscience, even with warnings","42 retracted papers, 30 models: refusals are spotty","Epistemic reliability test: LLMs reject only 7% of flawed premises"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every operational request is genuinely unanswerable unless the model accepts the target paper's false premise, so any non-refusal counts as premise acceptance; if a request could be legitimately answered as a hypothetical, a critique, or a sanity check, the 93% engagement rate overstates the failure.","fun_headline_variants_meta":{"raw":{"variants":["LLMs engage with false science in 95% of responses","Scientific integrity: 22 of 30 LLMs fail 90% of probes","LLMs rarely refuse pseudoscience, even with warnings","42 retracted papers, 30 models: refusals are spotty","Epistemic reliability test: LLMs reject only 7% of flawed premises"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1798,"prompt_tokens":1107,"completion_tokens":691,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":723,"completion_tokens_details":{"reasoning_tokens":596}},"tokens_in":723,"tokens_out":691,"duration_ms":9948,"temperature":1.0,"reasoning_tokens":596,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:12:19.538171+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask independent domain experts whether each of the 42 operational requests admits a legitimate non-accepting answer, such as an explicitly hypothetical study design or a critical analysis of the premise; if even a handful do, the engagement rate would no longer measure premise acceptance. A quicker check is to run the same probes against matched control papers with sound premises — the paper reports that refusals disappear under such controls, so if control refusals rose, the interpretation would collapse.","supporting_citations":[{"cited_title":"FrontierScience: Evaluating AI’s ability to perform expert-level scientific tasks","cited_arxiv_id":null,"evidence_quote":"Supplies the open-ended task gap the paper targets: a leading model scored 77% on structured science questions but 25% on open-ended ones."},{"cited_title":"Lexical recall or logical reasoning: Probing the limits of reasoning abilities in large language models","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that reworded logic puzzles collapse accuracy, motivating the claim that answer-centric benchmarks measure recall rather than reasoning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The MMR-autism retraction is a corpus probe whose refusal pattern is analyzed as evidence of a paper-specific classifier rather than epistemic reasoning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The cold-fusion palladium paper is the worked example of probe construction, showing how operational requests become premise-locked."},{"cited_title":"Electrochemically induced nuclear fusion of deuterium.Journal of Electroanalytical Chemistry and Interfacial Electrochemistry, 261(2, Part 1):301–308, 1989","cited_arxiv_id":null,"evidence_quote":"The original cold-fusion electrolysis report supplies the 'excess heat' vocabulary that anchors the first refusal point in the worked example."},{"cited_title":"Epel, Elizabeth H","cited_arxiv_id":null,"evidence_quote":"The telomere-stress study is the methodological control that draws no refusals, used to argue that Wakefield refusals are notoriety-based rather than reasoning-based."},{"cited_title":"RETRACTED: Clinical transplantation of a tissue-engineered airway.The Lancet, 372(9655):2023–2030, December 2008","cited_arxiv_id":null,"evidence_quote":"The retracted tracheal-transplant study is a corpus probe demonstrating sanewashing, where a model cites the fraudulent early cases as successful implants."},{"cited_title":"Superconductor Pb$_{10-x}$Cu$_x$(PO$_4$)$_6$O showing levitation at room temperature and atmospheric pressure and mechanism","cited_arxiv_id":"2307.12037","evidence_quote":"The LK-99 preprint is a corpus probe in the notoriety gradient, drawing far more refusals than equally flawed but less famous retractions."},{"cited_title":"Use of alternative medicine for cancer and its impact on survival.JNCI: Journal of the National Cancer Institute, 110(1):121–124, 01 2018","cited_arxiv_id":null,"evidence_quote":"The clinical study on alternative-medicine survival supplies the harm evidence for why engaging with unsound protocols can cost lives."}],"review_version":1}