{"id":"ae74f01b-a012-46bd-be16-26099563fcf7","arxiv_id":"2412.10635","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Commercial LLMs agree only modestly with expert confounder lists for the Coronary Drug Project and flip answers under trivial prompt changes, so they cannot yet act as reliable causal knowledge repositories.","lead":"This paper tests whether three commercial chatbots can reliably identify confounders for the Coronary Drug Project, using expert lists as the answer key. It finds mediocre accuracy and large sensitivity to wording, multiple-choice order, and model choice, concluding LLMs are not yet ready to automate causal-link reporting.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'non-confounder' gold standard is contaminated with clinical covariates that are plausible confounders (e.g., dyspnea, ankle swelling in post-MI patients), so the headline 'only slightly more likely' comparison is not supported.","rationale":"The paper has two distinct contributions: (1) a benchmark comparison of expert confounders versus non-confounders, and (2) evidence of response instability across models, prompts, and option ordering. Contribution (2) is well supported and largely independent of the ground-truth lists. Contribution (1), which drives the abstract's headline, depends on the validity of the negative control. The non-confounder set is the authors' own construction, and inspection of Appendix C shows it contains clinical findings that are plausible confounders in a post-MI cohort. This is a concrete correctness risk rather than a mere consensus disagreement: if 'shortness of breath at night' is a genuine confounder, then an LLM labeling it as a confounder is not a false positive. The reader identified the expert ground-truth lists generally as the weakest assumption; I partially agree, but the sharper problem is the non-confounder list, because it is not expert-derived and its categories are not clean negative controls. I do not recommend rejection because the instability results and the qualitative 'cannot automate reporting of causal links' conclusion retain support, but the quantitative 'only slightly more likely' phrasing should be conditional on validating the negative set and reporting uncertainty. Hence the verdict remains CONDITIONAL.","tokens_in":15362,"tokens_out":4242,"duration_ms":40124,"concrete_test":"Use the authors' OSF code to recompute the direct-method separation after restricting the negative set to clearly administrative or irrelevant variables (e.g., blood type, date of study entry, cause of death). Independently have two clinical epidemiologists, blind to the study design, classify all 60 non-confounders as definitely-not, possible, or likely confounders for the placebo-arm adherence-mortality question. If the confounder-designation gap between expert confounders and raters' 'definitely-not' variables is substantially larger than the reported gap, or if the reported gap falls within sampling noise when clustering by variable, the headline comparison must be revised.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract's central quantitative claim is the separation between expert confounders and non-confounders. That separation is only meaningful if the negative set is genuinely free of confounders. The 60 'Non-confounders' in Appendix C were assembled by the authors, not by expert elicitation, and the list includes 'general medical examination results' such as 'Shortness of breath at night', 'Swelling ankles', 'Rales', 'Palpable liver', 'Gouty arthritis', and 'Nervous system abnormalities'. In a post-myocardial-infarction population, these are prognostic markers that plausibly influence both adherence and mortality, i.e., they are confounders for the adherence-mortality effect. The category 'anticipated side-effect or known metabolite of active study medication' is also problematic for the generic prompt, since side effects can affect both the decision to continue taking X and downstream mortality. If many negative-set variables are true confounders, LLM 'false positives' are partly correct, and the reported 'only slightly more likely' separation is artificially compressed. This is not merely the paper's own caveat that experts might be wrong: the non-confounder list is not expert-derived at all, so the benchmark's negative control is unvalidated. The inconsistency and option-order results are less affected, but the precise headline comparison is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper tests whether three commercial LLMs (GPT-4o, GPT-o1-preview, Claude 3.5 Sonnet) can identify confounders for the adherence-mortality relationship in the Coronary Drug Project, using expert-selected confounder lists from CDPRG (1980), Murray and Hernán (2016), and Debertin et al. (2024), plus an author-constructed list of 60 supposed non-confounders. Across direct and indirect prompting, with and without encouraged reasoning, and with shuffled multiple-choice options, the authors report moderate true-positive rates, high false-positive rates on the non-confounder set, and substantial sensitivity to iteration, model, prompt, and option order. They conclude that LLMs cannot yet be used to automate the reporting of causal links for this kind of task.","tokens_in":15618,"tokens_out":4773,"duration_ms":44369,"significance":"If the empirical claims hold, the paper offers a useful, reproducible benchmark for evaluating whether LLMs can recall rather than reason about causal structure. The strengths are considerable: the positive ground truth comes from published expert lists rather than from the LLMs themselves; the study uses multiple models and prompt variants; repeated sampling reveals response distributions rather than single outputs; and the code is promised on OSF, providing a concrete falsifiable test for future models. The inconsistency and option-order results are robust and practically important regardless of any disagreement with the expert lists. However, the headline quantitative claim about the small separation between expert confounders and non-confounders depends on a negative control whose validity is questionable, and the effect sizes are presented without uncertainty quantification.","major_comments":[{"comment":"The 'Non-confounder' negative control is not expert-derived and appears contaminated with plausible confounders. In a post-myocardial-infarction cohort, variables such as 'Shortness of breath at night', 'Swelling ankles', 'Rales', 'Palpable liver', 'Gouty arthritis', and 'Nervous system abnormalities' are prognostic markers that can plausibly influence both adherence and mortality; several items in the 'anticipated side-effect' category can likewise affect continued adherence and downstream mortality. If a substantial share of these 60 variables are true confounders, the reported LLM false-positive rates are overstated and the abstract's 'only slightly more likely' separation is artificially compressed. Because this list was constructed by the authors rather than elicited from experts, it cannot serve as an unvalidated benchmark for the headline claim. Please recompute the separation using only administrative and sub-study variables that are unlikely to be confounders, or obtain independent expert validation of the negative set.","section":"Methods, Appendix C"},{"comment":"The central claim that expert confounders are 'only slightly more likely' to be labeled as confounders than non-confounders is presented without any inferential quantification. No confidence intervals, standard errors, or tests are given for the mean designation rates or their differences, even though the repeated prompting design provides the raw material for such estimates. As written, the reader cannot assess whether the observed separation is statistically distinguishable from zero for any model-prompt combination. Please add variable-level means with confidence intervals, or a compact table of differences with uncertainty, for at least the direct and indirect with-reasoning conditions.","section":"Abstract, Results, Table 1"},{"comment":"The GPT-o1-preview non-confounder data are incomplete and were collected under different conditions: a content-flagging workaround was added to direct prompts, follow-up variables were re-prompted as baseline measurements, and indirect-method prompts for non-confounders were apparently unavailable. Nevertheless, Table 1 reports a single row for GPT-o1 across all variable sets, and the figures plot non-confounder rates without indicating which cells are missing or based on modified prompts. This can distort cross-model comparisons and should be explicitly flagged, with affected cells either excluded from aggregate tables or analyzed separately.","section":"Data, Table 1, Figures 1-2"},{"comment":"The claim that 'text about the ground truth is in their training data' is supported only by asking the LLMs themselves whether they know about the Coronary Drug Project. Such self-reports are weak evidence of training-data inclusion, since models can produce plausible text without having seen a specific document. Please either verify inclusion by external means (for example, retrieval of distinctive phrases or documentation of publication dates in known corpora) or soften the claim to 'likely in the training data'.","section":"Introduction, Coronary Drug Project"}],"minor_comments":[{"comment":"There are several typos in the variable-set headings: 'Counfounders' should be 'Confounders', and 'abnormalitites' should be 'abnormalities'.","section":"Appendix A"},{"comment":"In the paragraph on option-order sensitivity, 'GPA-4o' should be 'GPT-4o'.","section":"Results, Consistency of Confounder Designations"},{"comment":"In the sub-study category, 'Alpha-lipoprotein cholecystitis' appears to be a copy-paste error and should read 'Alpha-lipoprotein cholesterol'.","section":"Appendix C"},{"comment":"The footnote explaining that 'Mixed' pools all intermediate percentages is important and should appear in the main text, because the claim that reasoning made 'no difference at all' for indirect prompting is an artifact of this coarse binning.","section":"Results, Figure 7"},{"comment":"The phrase 'only slightly more likely' is informal and underspecified; please replace it with a numeric difference or range once the uncertainty analysis is added.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"One of the ground-truth papers, Debertin et al. (2024), is co-authored by the manuscript's second author, and this is not disclosed in the text. The benchmark lists themselves predate the current submission, so the concern is not about data fabrication, but an explicit disclosure would be prudent. The paper fits the journal's scope and the OSF code availability is a clear strength; I would be willing to review a revised version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a useful empirical check on the claim that LLMs can automate causal modeling, and the consistency evidence is the strongest part. The headline comparison between expert confounders and non-confounders is the soft spot; the non-confounder list is self-constructed and probably contaminated, so the \"slightly more likely\" sentence should not be quoted without error bars or a cleaner negative control.\n\nWhat is new: they take a real observational problem (confounder selection in the Coronary Drug Project) with published expert lists, query three current LLMs with direct and indirect prompts, repeat draws, and test option-order sensitivity. The option-order effect (Cohen's kappa 0.13-0.41; 65.7% of variables switching for GPT-4o) and the iteration-level inconsistency are real, reproducible, and practically important. The paper also provides a reusable testbed and code. The \"causal parrot\" conclusion is not novel, but this is concrete negative evidence with current models.\n\nWhere it wobbles: the abstract's precise claim that expert confounders are \"only slightly more likely\" to be labeled than non-confounders depends on a negative control that is not expert-derived. The authors built the non-confounder list themselves, and several entries—shortness of breath at night, ankle swelling, rales, palpable liver, gouty arthritis in post-MI patients—look like plausible confounders of the adherence-mortality relationship. Side effects are especially problematic because they can plausibly affect both adherence and mortality. So the false-positive rate is partly disagreement with the authors' list, not necessarily a failure of the LLM. That said, this does not sink the paper: the order sensitivity and across-prompt inconsistency would remain worrying even with a cleaner negative set. I also note the missing inferential statistics on the headline comparison, and the o1-preview protocol deviation for non-confounders; the latter is minor but should be disclosed more prominently.\n\nThe ground truth is mostly fine: the confounder lists come from published work, though one list is the second author's own (Murray and Hernán 2016; Debertin et al. includes her). That is not disqualifying, but the authors should not treat their own list as independent expert opinion.\n\nWho this is for: applied researchers considering using LLMs to propose or screen confounders, and people building LLM-based causal inference tools. It deserves a serious referee. I would ask for a reanalysis with a cleaner non-confounder set and explicit confidence intervals before the abstract wording is taken as quantitative.\n\nRecommendation: send it out; require the o1 deviation to be reported clearly and either fix the negative control or soften the headline.","headline":"A useful, reproducible empirical check on LLMs as causal-knowledge repositories; the inconsistency results are solid, but the headline non-confounder comparison needs a cleaner negative control before it is quoted quantitatively.","tokens_in":16155,"tokens_out":1893,"would_cite":true,"duration_ms":17787,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Asking commercial LLMs to label confounders in a well-studied clinical trial produces only slightly better-than-chance discrimination against non-confounders, and answers flip with prompt wording and option order.","keywords":["large language models","causal inference","confounder selection","Coronary Drug Project","causal knowledge recall","prompt sensitivity","observational studies"],"falsifier":"Run the same confounder-labeling prompts with the full text of CDPRG (1980), Murray and Hernán (2016), and Debertin et al. (2024) inserted into the context window, then measure discrimination between expert confounders and the non-confounder list. If the models are repositories of causal knowledge, giving them the answer key should produce near-perfect and stable labels; if discrimination stays weak and answers still flip when the multiple-choice order is changed, the repository theory fails.","tokens_in":15152,"feed_emoji":"🤖","tokens_out":5607,"duration_ms":47086,"temperature":0.7,"pith_summary":"This paper asks whether commercial large language models can serve as a database of causal knowledge: given a variable and a study setting, will they correctly report whether experts treat that variable as a confounder? The authors test this using the Coronary Drug Project, a clinical trial where expert-selected confounder lists from 1980, 2016, and 2024 provide a ground truth, and where the relevant texts are believed to be in the models' training data. Across three models and several prompting styles, expert-confirmed confounders were only slightly more likely to be labeled confounders than variables experts explicitly consider non-confounders, and answers were unstable across models, prompts, and even the order of multiple-choice options. The paper concludes that LLMs do not yet have the ability to automate the reporting of causal links, so a non-expert cannot currently use them to build the covariate adjustment sets needed for valid observational causal inference.","feed_headline":"LLMs flunk confounder test even with answers in training data","feed_subtitle":"Three commercial models confused expert-selected confounders with non-confounders and flip-flopped when options were reordered.","key_machinery":"The test apparatus is a ground-truth benchmark built from the Coronary Drug Project placebo arm, where adherence affects mortality and confounding is well documented. Three expert sources supply confounder lists — the original 1980 trial report, a 2016 reanalysis, and a 2024 expert-curated causal diagram — and the authors add a list of 60 variables considered non-confounders. The models are queried directly ('is this variable a confounder?') and indirectly (first asking whether the variable affects adherence, then whether it affects mortality), with variants that add or remove step-by-step reasoning and that shuffle the answer options. A confounder is a pre-treatment variable that, if left unadjusted, distorts the estimated effect of treatment on outcome; the benchmark measures whether LLMs can reproduce expert judgments about which variables play that role.","core_discovery":"The central claim is that current mass-market LLMs, when used as recall devices rather than reasoners, cannot reliably retrieve expert causal judgments. Even though the models can produce text showing that the Coronary Drug Project and its confounder discussions are in their training data, their confounder labels discriminate only weakly between expert-chosen confounders and expert-designated non-confounders. The paper finds that the apparent success of some configurations is driven by a tendency to label almost everything a confounder rather than by genuine agreement with experts, and that designations change substantially when the prompt is direct versus indirect, when reasoning is encouraged versus suppressed, and when the multiple-choice options are reordered. In the authors' words, LLMs do not yet have the ability to automate the reporting of causal links.","pith_inferences":["If this pattern extends beyond the CDP, then any observational study that auto-generates adjustment sets from current LLMs risks systematically biased effect estimates, because false-positive confounders and false negatives enter the model in ways that are hard to detect without expert review.","The option-order flips suggest the models are choosing plausible continuations of the prompt rather than consulting a stable internal representation of causal structure; a testable extension would be to see whether supplying the relevant study text in the prompt removes the instability.","The authors note that experts might not be more correct than the LLMs in cases of disagreement; an editorial inference is that a multi-panel expert elicitation on the same variable set would separate 'LLMs lack causal knowledge' from 'experts disagree among themselves.'","A natural next experiment is to apply the same benchmark to other well-studied questions with published confounder sets, such as postmenopausal hormone therapy and cardiovascular disease, to see whether the mediocre discrimination is specific to the CDP or general."],"forward_implications":["A researcher cannot yet delegate confounder selection to an off-the-shelf LLM: in the best-performing configuration, 65–74% of expert-rejected variables were still labeled confounders.","A prompt that looks good on one model or dataset is not portable, because answers shift with option order and with direct versus indirect questioning.","Reported success rates overstate true causal knowledge, since much of the agreement comes from a general tendency to say 'confounder' rather than from discriminating between confounders and non-confounders.","The CDP setting gives future models a ready-made benchmark: the same code and ground-truth lists can be rerun as LLMs improve to test whether the gap closes.","If the inconsistency persists after the models have clearly absorbed more causal literature, that is evidence the failure is structural, not just a matter of training data coverage."],"supporting_citations":[{"why":"Supplies the original expert confounder list for the Coronary Drug Project and establishes the adherence–mortality confounding problem that is the test case.","marker":"CDPRG 1980"},{"why":"Supplies the modern reanalysis confounder set ('Added in 2016') that serves as part of the ground truth.","marker":"Murray and Hernán (2016)"},{"why":"Supplies the expert-curated causal diagram and trimmed covariate sets used as additional ground-truth labels.","marker":"Debertin et al. (2024)"},{"why":"Earlier recall-based test of GPT-3 on ground-truth causal links, motivating the authors' recall framing and prompt-sensitivity findings.","marker":"Long, Schuster, and Piché (2023)"},{"why":"Introduces the 'causal parrots' view of LLMs as unreliable causal-link databases and provides a prior benchmark showing poor matching and prompt sensitivity.","marker":"Zečević et al. (2023)"},{"why":"Source of the option-ordering test used to show that LLM confounder judgments shift with irrelevant answer-order changes.","marker":"Nori et al. (2023)"}],"fun_headline_variants":["Causal recall? LLMs flunk expert confounder test","LLMs fail to retrieve causal links from memory","Expert confounders elusive for LLMs in drug trial","LLMs inconsistent on confounder labels, study finds","Confounder identification: LLMs lag behind experts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The expert-selected confounder lists from the three studies, and the authors' list of non-confounders, are treated as the correct ground truth for the adherence–mortality relationship; if those lists are wrong or context-dependent, the measured performance is disagreement with particular experts rather than absence of causal knowledge.","fun_headline_variants_meta":{"raw":{"variants":["Causal recall? LLMs flunk expert confounder test","LLMs fail to retrieve causal links from memory","Expert confounders elusive for LLMs in drug trial","LLMs inconsistent on confounder labels, study finds","Confounder identification: LLMs lag behind experts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1291,"prompt_tokens":873,"completion_tokens":418,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":341}},"tokens_in":489,"tokens_out":418,"duration_ms":4365,"temperature":1.0,"reasoning_tokens":341,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:45:36.336457+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same confounder-labeling prompts with the full text of CDPRG (1980), Murray and Hernán (2016), and Debertin et al. (2024) inserted into the context window, then measure discrimination between expert confounders and the non-confounder list. If the models are repositories of causal knowledge, giving them the answer key should produce near-perfect and stable labels; if discrimination stays weak and answers still flip when the multiple-choice order is changed, the repository theory fails.","supporting_citations":[],"review_version":1}