{"id":"9c4b554b-77ed-4d80-8548-272d327baa7e","arxiv_id":"2412.10509","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLMs show a causal illusion bias: they often report causal relationships when evidence is purely correlational, null-contingency, or temporally impossible, particularly in 0-100 scaled judgments.","lead":"Researchers tested three large language models on tasks where correlations were spurious, drug-outcome contingency was absent, and temporal order ruled out causation. The models frequently still produced causal claims, especially when answer scales ran from 0 to 100, suggesting they have not reliably internalized normative causal principles. A generalist might care because the same failure could amplify superstition, misinformation, and misleading health headlines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0-100 effectiveness ratings are read as unbiased evidence of causal illusion, but the prompt's '50 = quite effective' anchor and real-drug priors could produce the same numbers without any causal inference, undermining the numeric-task conclusion.","rationale":"The strongest claim is specifically about the contingency-judgment task; the open-ended headline task is less affected by this concern because it is annotated categorically and has inter-annotator agreement (kappa 0.84/0.80), so its evidence for some level of bias is credible. The load-bearing step is the inference from 'rated effectiveness 75.21' to 'has a causal illusion.' Once the prompt's midpoint label and the use of real drugs are taken into account, high ratings can be generated by response conventions and prior knowledge rather than by an assessment of the presented contingency. The authors themselves flag the 0-100 scale as potentially unsuitable for LLMs, which supports evaluating this concern before accepting the headline claim. A two-format replication, one de-anchored numeric scale and one forced-choice causal question, would directly test the interpretation; if the effect survives both, the concern is retired. I agree with the reader that this is the weakest assumption; the null-contingency generation algorithm is also underspecified, but it is secondary because the code and data are released for direct inspection, whereas the scale semantics are baked into every numeric response. The reader's CONDITIONAL verdict already conditions on this kind of validation, so no verdict adjustment is needed.","tokens_in":10545,"tokens_out":9094,"duration_ms":88217,"concrete_test":"Re-run the 1,000 null-contingency scenarios from Section 5.2 with the same trial data but (a) a neutral prompt that removes the '50 signifies quite effective' label, using only '0 = no effect on recovery, 100 = complete effect on recovery'; and (b) a forced-choice causal version for the same scenarios: 'Does medicine A have any causal effect on recovery? Yes/No.' If GPT-4o-Mini's mean rating in (a) drops below 50 and/or its Yes rate in (b) falls to chance or below while the original mean stays near 75.21, the high ratings are an artifact of the loaded response scale rather than evidence of causal illusion. Recompute all pairwise significance tests under both formats.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline numeric claim depends on treating a 0-100 rating as a neutral readout of perceived causal contingency. That link is not secure. Appendix A's prompts instruct the model that '0 indicates non-effective, 50 signifies quite effective, and 100 represents totally effective,' and Section 2 defines any score above 0 in a null-contingency scenario as evidence of the illusion. For an LLM, 'effective' is task-pragmatically loaded and semantically anchored: a labeled midpoint at 50 invites moderate-to-high numbers, and real drug names such as paracetamol (Appendix A) bring world-knowledge priors that the drug treats fever, independent of the zero-contingency trial list. The doctor/researcher role framing further pushes toward helpful numeric answers rather than a contingency calculation. GPT-4o-Mini's mean of 75.21 with zero 0-responses is therefore consistent with response-format artifacts and does not by itself establish that the model inferred a causal relation. The authors' Limitations section concedes that the 0-100 scale 'may not be ideal for evaluating LLMs, and could partly explain the results'; that admission is in-scope evidence and weighs directly against the abstract's 'significantly higher bias' conclusion for numeric tasks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether three LLMs (GPT-4o-Mini, Claude-3.5-Sonnet, Gemini-1.5-Pro) exhibit the illusion of causality across three tasks: headline generation from spurious-correlation abstracts, a 0-100 contingency judgment task in medical scenarios, and a superstitious-thinking inference task with temporal and alternative-cause cues. The authors report strong causal-illusion bias in open-ended headline generation at levels comparable to or lower than human press-release exaggeration, and significantly higher bias in the numeric 0-100 tasks, concluding that the models have not reliably internalized normative causal-learning principles. The dataset and generation code are made available.","tokens_in":10794,"tokens_out":3754,"duration_ms":34165,"significance":"If the numeric-task results are valid, this is an important contribution to the study of cognitive biases in LLMs and to AI-safety evaluation. The headline-generation task is carefully annotated with high inter-annotator agreement (kappa 0.80-0.84), the three-domain dataset is a useful public resource, and the paper is transparent about several limitations. However, the central numeric-task conclusion is not yet secure because the 0-100 response scale is pragmatically anchored and confounded with world-knowledge priors; the paper's own Limitations section concedes that the scale 'may not be ideal for evaluating LLMs, and could partly explain the results.' The finding therefore needs additional control conditions or a more cautious interpretation before the 'significantly higher bias' claim can be accepted.","major_comments":[{"comment":"The 0-100 effectiveness rating is not a neutral readout of perceived contingency. The prompt explicitly tells models that 0 means 'non-effective', 50 means 'quite effective', and 100 means 'totally effective', and several prompts use real drug names such as paracetamol for fever. This invites scale anchoring and activation of prior knowledge about drug efficacy, independent of the zero-contingency trial list. Since Section 2 defines any score above 0 in a null-contingency scenario as evidence of the illusion, the GPT-4o-Mini result of mean 75.21 with zero 0-responses (Section 6.2) is consistent with response-format artifacts. The Limitations section (Section 7) concedes that the scale 'may not be ideal for evaluating LLMs, and could partly explain the results'; this concession directly weakens the abstract's 'significantly higher bias' claim for numeric tasks.","section":"Section 5.2 and Appendix A"},{"comment":"The claim that GPT-4o-Mini 'failed to recognize, in any of the 1,000 zero-contingency scenarios, that there was no causal relationship' equates 'no causal relationship' with a response of exactly 0. However, the prompt describes 50 as 'quite effective' and never tells the model that 0 is the only correct answer under null contingency. A score around 50 could be a neutral default or an anchoring artifact rather than an inference of causation. The normative definition of bias should be adapted to the LLM setting, or control prompts with a neutral midpoint and invented or generic drug names should be used.","section":"Section 6.2"},{"comment":"The paper states that the contingency-task results 'bear a resemblance' to human results but provides no human baseline data or statistical comparison. The abstract's 'significantly higher bias' claim for numeric tasks needs an explicit comparison group or a pre-registered threshold; qualitative resemblance to the literature is not sufficient to support the strength of the conclusion. Running a small human sample or reusing published human distributions would make the claim testable.","section":"Section 5.2 and Section 7"},{"comment":"The description of null-contingency trial generation is under-specified. The text says that '80% of each half assigned to combinations where one variable remained constant while the other varied (e.g., potential cause present and potential outcome absent), and the remaining 20% assigned to configurations where both variables either remained fixed or varied together.' This does not uniquely determine the joint probabilities needed to verify that delta P = 0 for every generated scenario. Please provide the exact generation rule or include a code-level check that all scenarios satisfy the null-contingency condition.","section":"Section 4"}],"minor_comments":[{"comment":"In the sentence 'They recommend to extent the evaluation to more real-world false beliefs,' 'extent' should be 'extend'.","section":"Section 3"},{"comment":"The text uses the decimal comma in '17,5%' while other percentages use a decimal point; please standardize the notation.","section":"Section 6.1"},{"comment":"The captions state 'All model pairs show statistically significant differences (p <0.0001)' without reporting the test used or whether multiple-comparison corrections were applied; please add this information.","section":"Figures 4 and 6"},{"comment":"The superstitious-thinking prompts are described but not shown in Appendix A, which only contains contingency-task prompts; including them would improve reproducibility.","section":"Section 5.3"},{"comment":"The sentence about Claude-3.5-Sonnet says its performance 'aligns closely' with a 22% human exaggeration rate while 'shows a lower bias (17,5%)'; please clarify whether 17.5% is the direct/conditional causal rate or a combined measure, and whether the comparison is statistically meaningful.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an AI/cognitive-science venue and does not have a circularity problem: the claims are empirical measurements against normative definitions, not derivations that assume their conclusions. The main risk is construct validity of the numeric tasks. If the authors add control conditions or substantially soften the numeric-task conclusions, the paper could be publishable; otherwise the headline claim overreaches. I would not recommend rejection, because the open-ended headline task and the dataset remain valuable even if the numeric-task conclusion is qualified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take after reading it through. The headline-generation task is the genuine contribution. The authors built a 1,000-abstract corpus of spurious correlations, prompted three LLMs from journalist and researcher perspectives, and manually annotated the resulting headlines for causal claims. The kappa values (0.80 and 0.84) are solid, and the comparison to the 22% human exaggeration rate in Yu et al. is a nice external anchor. That part of the paper is credible and worth building on.\n\nThe numeric contingency task is where the main claim gets shaky. The prompt explicitly says 50 means 'quite effective.' The paper then counts any score above 0 in a null-contingency scenario as evidence of the illusion. That logic assumes the 0–100 rating is a neutral readout of perceived contingency, but for an LLM it isn't. 'Effective' is semantically loaded, a labeled midpoint invites moderate-to-high responses, and real drug names like paracetamol bring world-knowledge priors that don't depend on the trial list at all. GPT-4o-Mini's mean of 75.21 with zero zeros is exactly what an anchoring artifact would look like. The authors' own Limitations section concedes the scale 'may not be ideal for evaluating LLMs, and could partly explain the results.' That concession is in the paper, and it directly contradicts the abstract's 'significantly higher bias' framing for the numeric tasks.\n\nWhat else? The dataset and code are released, and the variable taxonomy (invented, indeterminate, pseudo-medicine, validated drugs) is thoughtful. The null-contingency generation algorithm is described too loosely—'controlled distribution of 80% and 20%' doesn't pin down the joint distribution—so reproduction would be guesswork. The superstitious task shares the same scale issue, though it's less central. And with only three models, 'not uniformly, consistently, or reliably' is overreach.\n\nWho gets value from this? People designing LLM bias evaluations, and psychologists who want an automated probe of the illusion of causality. The headline-task half deserves a serious referee. The numeric half needs a neutral-scale replication or a same-stimulus human baseline before the strong conclusion is justified. I'd send it to review, but with a clear revision request: fix or drop the numeric task as evidence of causal bias, or at least reframe the conclusion to match the admitted limitation.","headline":"Good headline-generation study, but the numeric contingency task that carries the main claim is compromised by the prompt's own '50 = quite effective' anchor.","tokens_in":11331,"tokens_out":3092,"would_cite":true,"duration_ms":27803,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models infer causation where the evidence shows none, with GPT-4o-Mini averaging 75.21/100 on null-contingency medical cases and never answering 0 in 1,000 trials.","keywords":["Causal Learning","Illusion of causality","Large Language Models","contingency judgment task","null contingency","spurious correlation","causal bias","GPT-4o-Mini"],"falsifier":"Run the same null-contingency medical trials but change the response instruction to 0 = 'no causal relationship' and 100 = 'definite causal relationship', or ask a binary causal question (cause or no cause) before the numeric rating. If the near-75 average and the absence of zero responses persist, the illusion is robust; if ratings collapse toward 0 or binary causal choices become rare, the reported bias is largely an artifact of the 0–100 effectiveness scale with its '50 = quite effective' anchor.","tokens_in":10382,"feed_emoji":"🧠","tokens_out":7163,"duration_ms":59117,"temperature":0.7,"pith_summary":"This paper asks whether large language models fall prey to the same illusion of causality that humans do—believing one thing causes another when the evidence shows no such link. To find out, the authors built over 2,000 scenarios: abstracts with bare correlations, medical null-contingency trial lists, and superstitious testimonies where the effect precedes the cause. Three models (GPT-4o-Mini, Claude-3.5-Sonnet, Gemini-1.5-Pro) were prompted to state causal claims or rate effectiveness on 0–100 scales. The paper reports a strong causal-illusion bias, most striking in null-contingency numeric judgments: GPT-4o-Mini averaged 75.21/100 and gave zero in none of 1,000 cases. The authors conclude that the models have not uniformly, consistently, or reliably internalized the normative principles that should guide causal learning.","feed_headline":"GPT-4o-Mini rates null-effect drugs at 75/100 on average","feed_subtitle":"In 1,000 null-contingency cases, GPT-4o-Mini never rated a drug at 0; all three models showed causal illusion.","key_machinery":"The central mechanism is the adapted contingency judgment task, a standard human experimental procedure in which trial-by-trial observations of a potential cause (drug given) and outcome (recovery) are summarized, and the model rates the cause's effectiveness on a 0–100 scale. The paper's null-contingency datasets ensure that the outcome probability is identical whether the cause is present or absent, which normatively should yield a rating of 0 (no causal effectiveness). Alongside this, the headline-generation task classifies outputs into correlational, conditional-causal, direct-causal, and no-claim categories, and the superstitious-thinking task inserts temporal and alternative-cause cues; both are designed to detect causal claims where evidence is absent.","core_discovery":"The central claim is that LLMs exhibit a strong, measurable illusion of causality in causal learning tasks, particularly when quantitative judgment scales are used. Across three tasks—generating headlines from spurious-correlation abstracts, rating drug effectiveness from null-contingency trial lists, and judging superstitious outcomes given temporal counterevidence—the models made causal claims unsupported by the evidence. In the numeric tasks the effect was pronounced: GPT-4o-Mini's mean null-contingency effectiveness rating was 75.21 (SD = 12.52) with no zero responses out of 1,000; Claude-3.5-Sonnet averaged 43.46 with 12.1% zeros; Gemini-1.5-Pro averaged 33.75 with 28.5% zeros but also high variability. The paper interprets this as evidence that the models have not uniformly, consistently, or reliably internalized normative principles such as contingency and temporal ordering, and it reads the failure to use temporal cues as contrary to some prior LLM results.","pith_inferences":["Editorial inference: the 0–100 response scale itself may inflate the illusion measure, since the prompt defines 50 as 'quite effective'; a null-contingency judgment near 50 may express scale anchoring or prior beliefs about drugs rather than an inferred causal link.","Testable extension: repeat the contingency task with a causal-only scale (0 = no causal relationship, 100 = definite causal relationship) and with abstract variables like 'X' and 'Y'; if mean ratings fall near zero, the reported bias is partly a scale artifact.","Editorial inference: if the illusion is genuinely driven by text priors, fine-tuning on explicit null-contingency trials or chain-of-thought instructions may reduce it—a mitigation the paper lists as future work, not something it demonstrates.","Testable extension: a direct human-vs-model comparison on the exact same prompts would place the reported effect sizes in clearer perspective, since the paper's human comparisons come mostly from earlier psychology studies rather than a matched sample."],"forward_implications":["If the paper is right, LLM outputs in health, science communication, and belief-related domains cannot be assumed to reflect evidence-based causal reasoning.","Numeric causal ratings from LLMs are not reliable for null-contingency evidence; a widely deployed assistant could tell a patient that an ineffective treatment is quite effective.","Explicit textual cues such as 'correlation does not imply causation' do not reliably reduce causal illusion in headline generation.","Models differ substantially in bias strength, so safety assessments for causal reasoning should be model-specific.","Temporal information that places the effect before the cause fails to negate causal inference in most cases, despite earlier evidence that some LLMs use temporal cues."],"supporting_citations":[{"why":"Defines the illusion of causality and the contingency judgment task that the paper adapts, supplying the human-baseline framing.","marker":"Matute et al., 2015"},{"why":"Provided the human contingency-judgment instruction format and prompt inspiration for the medical null-contingency task.","marker":"Moreno-Fernandez et al., 2021"},{"why":"Supplies the headline exaggeration categories, the language cues, and the 22% human press-release exaggeration baseline used in the headline task.","marker":"Yu et al., 2020"},{"why":"Supports the criterion that any effectiveness rating above 0 in null contingency indicates the presence of causal illusion.","marker":"Vinas et al., 2023"},{"why":"Provides the normative principles of temporal ordering and contingency that the paper uses to define correct causal inference.","marker":"Blanco, 2017"},{"why":"Prior evidence that LLMs mirror human biased causal judgments, motivating the expectation that biases transmitted through language rather than experience.","marker":"Keshmirian et al., 2024"},{"why":"Contrast case: an earlier finding that LLMs infer absence of causal relations from temporal cues, which this paper's superstitious-thinking results do not reproduce.","marker":"Joshi et al., 2024"}],"fun_headline_variants":["LLMs show causal illusion bias, rate null-effect drugs at 75/100","GPT-4o-Mini rates inert drugs 75/100 in null-contingency tests","LLMs fail temporal cues in causal learning, show strong illusion bias","Causal illusion in LLMs: no zeros from GPT-4o-Mini on 1,000 null cases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measurement of causal bias assumes that a model's 0–100 effectiveness rating is a neutral readout of the contingency evidence, yet the prompts in Section 5.2 and Appendix A explicitly tell the model that 50 means 'quite effective', so higher scores could reflect anchoring or prior beliefs about drugs rather than an inferred causal relationship.","fun_headline_variants_meta":{"raw":{"variants":["LLMs show causal illusion bias, rate null-effect drugs at 75/100","GPT-4o-Mini rates inert drugs 75/100 in null-contingency tests","LLMs fail temporal cues in causal learning, show strong illusion bias","Causal illusion in LLMs: no zeros from GPT-4o-Mini on 1,000 null cases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000746,"raw_usage":{"total_tokens":3370,"prompt_tokens":1036,"completion_tokens":2334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":2240}},"tokens_in":652,"tokens_out":2334,"duration_ms":15727,"temperature":1.0,"reasoning_tokens":2240,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:53:13.543250+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same null-contingency medical trials but change the response instruction to 0 = 'no causal relationship' and 100 = 'definite causal relationship', or ask a binary causal question (cause or no cause) before the numeric rating. If the near-75 average and the absence of zero responses persist, the illusion is robust; if ratings collapse toward 0 or binary causal choices become rare, the reported bias is largely an artifact of the 0–100 effectiveness scale with its '50 = quite effective' anchor.","supporting_citations":[{"cited_title":"Vadillo, and Itxaso Barberia","cited_arxiv_id":null,"evidence_quote":"Defines the illusion of causality and the contingency judgment task that the paper adapts, supplying the human-baseline framing."},{"cited_title":"Scarcity affects cognitive biases: The case of the illusion of causality","cited_arxiv_id":null,"evidence_quote":"Supports the criterion that any effectiveness rating above 0 in null contingency indicates the presence of causal illusion."},{"cited_title":"Positive and negative implications of the causal illusion","cited_arxiv_id":null,"evidence_quote":"Provides the normative principles of temporal ordering and contingency that the paper uses to define correct causal inference."},{"cited_title":"Chain versus common cause: Biased causal strength judgments in humans and large language models","cited_arxiv_id":null,"evidence_quote":"Prior evidence that LLMs mirror human biased causal judgments, motivating the expectation that biases transmitted through language rather than experience."},{"cited_title":"LLMs Are Prone to Fallacies in Causal Inference","cited_arxiv_id":"2406.12158","evidence_quote":"Contrast case: an earlier finding that LLMs infer absence of causal relations from temporal cues, which this paper's superstitious-thinking results do not reproduce."}],"review_version":1}