{"id":"4f750e45-60d0-4020-8c5c-d02ed790c91f","arxiv_id":"2501.05844","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The authors argue that causality is domain-specific mechanistic narrative, so causal machine learning methods cannot, on their own, certify true causes outside narrow physical or engineered systems.","lead":"This paper argues that the word \"cause\" does not have one fixed technical meaning, but instead takes on different forms inside different scientific fields while serving one shared function: describing the mechanism behind an influence. It uses the philosophy of ordinary language to conclude that causal machine learning, which treats cause as a graph of statistical dependencies, overstates what its results certify.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The argument's normative premise—that ordinary language in each scientific tribe fixes the legitimate meaning of 'cause'—is asserted rather than defended, and the paper's own unified functional definition undercuts it.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the paper relies on ordinary-language use within scientific communities as the normative standard for causal claims, rather than defending that standard. My attack sharpens this by pointing out that the paper itself supplies a unified functional characterization of cause—mechanism-based influence—which can be directly formalized in the structural causal model framework it criticizes. This does not refute the paper's more modest claims about the practical brittleness of causal discovery in open systems (causal sufficiency failures, model representation bias, type-1 error concerns), which are well-taken and appropriately conditional. But the strong conclusion that causal ML 'cannot' certify causes outside physics rests on an unargued semantic premise. This is a significant but addressable weakness, not a falsification of the central argument: the paper could defend its ordinary-language methodology or weaken the impossibility claim to a pragmatic caution. Since the reader already gave CONDITIONAL, my analysis does not change the verdict; it reinforces the condition under which the paper's strongest conclusion would be accepted.","tokens_in":26521,"tokens_out":6487,"duration_ms":73031,"concrete_test":"Analytical test: formalize the paper's functional definition of cause (Section 1) as a structural causal model with mechanisms f_i and independent noise, and show that the DAG abstraction follows by the Markov factorization. Then re-read Section 6.2's claim 'nor can it' with this formalization in hand: if the DAG is an abstraction of the paper's own mechanism definition, ask what additional domain-specific content is missing that could not in principle be added to the model (e.g., scale-specific equations). If no such content is identified, the premise that only within-paradigm language can assert causation is doing no work, and the paper's rejection of causal ML is a terminological preference rather than an epistemic result.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion—that a DAG plus conditional-independence tests 'cannot' certify causes outside physics and 'amounts to redefining the word \"cause\"' (Sections 3.5, 6.2)—depends on a normative semantic premise: that the customary use of 'cause' inside each scientific language game is the correct standard for what a causal claim must mean. The paper documents diversity of usage but never argues for this authority. It even offers its own cross-domain functional definition—'the mechanism underlying fundamental forces of influence' (Section 1)—which is precisely what a structural causal model formalizes (each assignment X_i := f_i(Pa_i, U_i) is a mechanism). If that is so, the graphical model is not a redefinition but an abstraction of the paper's own notion of cause; the remaining objection is about the appropriate level of abstraction, a pragmatic matter. Without an independent defense of the ordinary-language premise, the strong impossibility claim is unsupported and the thesis reduces to a recommendation to treat domain-specific usage as authoritative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This is an argumentative, interdisciplinary paper that critiques the foundations of causal machine learning (causal inference/discovery as implemented via Bayesian networks, DAGs, and conditional independence tests). The authors argue, using ordinary language philosophy, that the word \"cause\" has no single formal definition; its meaning is fixed by the language games of distinct scientific communities. They then demarcate physics/engineering (where mathematical models can fully capture causality), biology (where emergence requires multi-scale evidence), and the social sciences (where hermeneutic interpretation is valuable but precise causal claims require multi-domain convergence). Their central critical claim is that DAG-based causal learning should not be read as certifying that 'A causes B', and they argue that definitive causal claims about open systems require an agglomeration of evidence across multiple domains and levels of abstraction.","tokens_in":26693,"tokens_out":6523,"duration_ms":59441,"significance":"If its central argument were fully supported, the paper would be a significant contribution to the ongoing debate about the overinterpretation of causal machine learning results. Its main strengths are its synthetic use of philosophy of science (Wittgenstein, Kuhn, Lakatos, Quine) to frame the issue, its clear typology of causal semantics across physics, biology, and social science, and its explicit concession that causal formalism is correct for physics/engineering when the underlying equation model is known and exact. The paper also makes a useful practical recommendation: that causal discovery results in open, complex systems should be treated as hypothesis-generating rather than as definitive certification. The paper does not claim to provide mathematical proofs; its value is conceptual, and its credibility hinges on the normative premise about the authority of ordinary language in fixing the meaning of 'cause'.","major_comments":[{"comment":"The paper moves from the descriptive observation that 'cause' is used differently in different scientific domains to the normative conclusion that the ordinary-language usage within each scientific tribe fixes what a causal claim must mean. This premise is never defended. The strongest statement in §6.2 that DAG-based causal discovery 'nor can it' certify causes, and in §3.5 that it 'amounts to redefining the word \"cause\"', depends entirely on that premise. Furthermore, the paper's own functional definition in §1—'the mechanism underlying fundamental forces of influence'—is precisely what a structural causal model's assignment X_i := f_i(Pa_i, U_i) formalizes. The graphical model is therefore not a redefinition but an abstraction of the paper's own concept; the remaining dispute is about whether the level of abstraction preserves the domain-relevant mechanisms, which is a pragmatic and gradable epistemic question. The categorical 'cannot' is not supported. I recommend reframing the conclusion as a counsel of caution about overclaiming from DAG+CI in open systems, rather than an in-principle impossibility.","section":"§3.4–§3.5, §6.2"},{"comment":"The positive thesis that definitive causal claims require an 'agglomeration of consistent evidence across multiple domains' is asserted but not operationalized. If, as the paper argues with Kuhn in §3.3, scientific paradigms are incommensurable language games, then it is nontrivial to say what counts as consistency of evidence across those games. The examples (smoking, Weber's Protestant Ethic, cognitive science) are suggestive, but no criterion is given for when cross-domain results harmonize rather than merely coexist. Without such a criterion, the proposed mixed-methods framework is underspecified as a research program.","section":"§6.3, §3.3"},{"comment":"Several empirical generalizations are made without supporting evidence. The claim in §2.1 that the authors could not identify a single real-data causal-learning paper with all conditional independence tests passing is anecdotal, not a systematic survey. Similarly, §6.2 states that sparse causal models in social sciences are 'uncommon in the literature' and that the few that exist 'present type-1 error concerns', but no citation is given. These empirical claims are used to bolster the critique of causal learning's practical value and should either be substantiated with a proper literature review or removed.","section":"§2.1 and §6.2"}],"minor_comments":[{"comment":"The final sentence of the abstract is a fragment: 'Given the role of epistemic hubris ... optimizing integration of different findings.' It needs a main clause to be a complete sentence.","section":"Abstract"},{"comment":"The name 'Halpern' is misspelled as 'Harpen' twice in the discussion of actual causality.","section":"§3.4"},{"comment":"'paropagation' should be 'propagation'.","section":"§5.3"},{"comment":"'feedforwark' should be 'feedforward'.","section":"§5.4"},{"comment":"The quotation 'bewitchment of intelligence by language' is a paraphrase of Wittgenstein's Philosophical Investigations §109; an exact citation would be helpful.","section":"§6.2"},{"comment":"The phrase 'temporal difference in differences at interventions' is unclear and should be clarified.","section":"§2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is best categorized as a philosophy-of-science position paper rather than a technical CS contribution. Its fit with a cs.LG venue depends on whether the journal values such critical, cross-disciplinary essays. The main technical weakness is the unsupported move from descriptive diversity of usage to a normative conclusion about the meaning of 'cause'; the authors should be encouraged to either defend that premise explicitly or soften the impossibility claim. The paper also cites its own 'extended version' (arXiv v2) for several important details; the editor may wish to request that the version under review be self-contained on those points."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is best read as an essay, not a technical result. What is genuinely new is the ordinary-language framing paired with a three-way demarcation: physics and engineering have closed equation systems where causal claims can be complete; biology requires consistency across scales because of emergence; social science offers hermeneutic and instrumental value even without precise predictive models. The smoking example is a good illustration of how causal consensus actually forms in medicine, and the paper is honest enough to concede that causal ML tools are not worthless when the equation form is right. That is real credit.\n\nThe soft spots are proportionate but not trivial. The main one is the load-bearing premise: the paper asserts that the customary use of “cause” inside each scientific language game is the correct standard for what causal claims must mean, and never really argues for it. Absent that defense, the strong claim in Section 6.2 that conditional-independence causal discovery “nor can it” certify causes does not follow. The paper itself defines cause as “the mechanism underlying fundamental forces of influence,” which is close to what structural causal models formalize; so the graphical model is better seen as an abstraction of that notion, not a redefinition. Once you frame it that way, the remaining dispute is about the right level of abstraction, which is a pragmatic matter, not a refutation. The stress-test note lands.\n\nThere are also a few empirical assertions without support: the claim that the authors could not find one real-data causal learning paper with all conditional-independence tests passing, the claim that few social-science DAG papers exist without type-1 error concerns, and the “99% of the exact function is known” figure. Minor in an essay, but should be fixed. The abstract also ends with a grammatical fragment, which needs rewriting.\n\nThe argument is otherwise coherent and the literature engagement is broad and mostly fair. It is not a new technical result, and the central critical content partly overlaps with Eberhardt and Peters et al., but the domain taxonomy and the hermeneutics framing are a real contribution to how causal ML results get communicated.\n\nFor whom? Philosophers of science, causal ML researchers worried about overclaiming, and anyone teaching responsible use of causal methods. It deserves a serious referee, not a desk reject. If I were the editor, I would send it out with a referee instruction to focus on the ordinary-language authority premise and the unsupported empirical claims; with those addressed, it is publishable in a venue that welcomes philosophical critique of ML.","headline":"A well-read philosophical essay whose descriptive core about domain-specific causal language largely works, but whose strong conclusion that causal ML cannot certify causes rests on an under-defended ordinary-language premise.","tokens_in":27217,"tokens_out":2099,"would_cite":true,"duration_ms":25881,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Causal machine learning redefines 'cause' and overstates what statistical graphs can certify, a philosophical critique argues.","keywords":["causality","causal machine learning","causal discovery","ordinary language philosophy","language games","epistemology","scientific domains","hermeneutics"],"falsifier":"A single well-documented case where a causal DAG discovered purely from conditional independence in an open biological or social system led to a verified intervention that the field's mechanistic theories had rejected, and where the mechanism was later confirmed, would directly contradict the paper's claim that independence-based discovery cannot certify real causes outside physics and engineering.","tokens_in":1518,"feed_emoji":"🧠","tokens_out":2883,"duration_ms":56244,"temperature":0.7,"pith_summary":"The paper argues that the word 'cause' has no single formal definition that a general-purpose computational method can capture. Instead, what counts as a cause depends on the scientific domain: physics expresses causes as terms in equations, biology requires consistent mechanisms across scales, and the social sciences rely on interpretive narratives. Because of this, the paper concludes that a discovered directed acyclic graph from conditional independence tests should not be read as demonstrating that 'A causes B' in any domain. This matters because causal machine learning is being used in medicine and policy, and overstating what its graphs certify could mislead real decisions. The paper proposes that definitive causal claims about open systems require convergent evidence across multiple scientific language games.","feed_headline":"Statistical graphs alone can't certify causes, paper argues","feed_subtitle":"A philosophical critique says 'cause' is a domain-specific mechanistic narrative, so causal ML overstates what its DAGs prove.","key_machinery":"The central machinery is the Ordinary Language method, drawn from Wittgenstein's language-game analysis, which treats the meaning of 'cause' as its actual use within each scientific community's practices. This method lets the paper show that the grammar of causal claims differs across physics, biology, and social science, and it uses Lakatos's 'hard core' idea to locate causal mechanisms within each domain's foundational assumptions. The argument then proceeds by contrasting this domain-relative meaning with the graphical independence model of causal learning, which the paper says amounts to redefining the word 'cause'.","core_discovery":"The authors discover that causality is a functional concept whose form is fixed by the grammatical conventions of each scientific community, not a single relation that statistical independence can certify. They argue that causal learning takes the exact equation systems of physics as an analogy and applies that analogy to open, emergent, and interpretive systems where it does not carry the same authority. Consequently, a fitted DAG with passing conditional-independence tests is at best a plausibility suggestion, not a certified cause. Definitive causality requires what the authors call an agglomeration of consistent evidence across domains, scales, and narrative forms.","pith_inferences":["A concrete test of the paper's position would be to run causal discovery on well-understood physics benchmarks versus open biological or social datasets and compare whether the discovered graphs align with the domains' known mechanisms.","The argument implies a communication standard for high-stakes 'causal AI': outputs should be phrased as 'consistent with' rather than 'causes' unless the domain mechanism is known and identified.","The authors' sketch of analogical reasoning via CP-logic suggests a research program for formalizing cross-domain evidence integration, which could be tested by building systems that transfer causal rules between modeled domains.","The critique can be read as a call for pluralistic causal models, where physics, biology, and social science each get a distinct causal calculus rather than one universal statistical framework."],"forward_implications":["Causal discovery results on real datasets should be reported as generating hypotheses, not as certified causes.","Research evaluating causal machine learning should demand domain-specific mechanistic validation before accepting a graph as evidence of causation.","The replication crisis in social science will not be fixed by more sophisticated statistics alone, but by lowering confidence and seeking convergence across multiple disciplines.","Causal statements in physics and engineering remain legitimate where the equation models are known and validated.","The burden of proof in medicine and policy shifts from a single DAG or p-value to multi-scale, multi-domain evidence."],"supporting_citations":[{"why":"Supplies the Ordinary Language method and the notion of language games that grounds the entire argument about domain-relative meaning.","marker":"[Wittgenstein, 2009]"},{"why":"Defines the graphical causal framework, Bayesian networks, and do-calculus that the paper critiques as an inappropriate abstraction level.","marker":"[Pearl, 2009]"},{"why":"Codifies the formal foundations of causal inference, including structural causal models and additive noise models, which are the target of the critique.","marker":"[Peters et al., 2017]"},{"why":"Provides the web-of-knowledge and inductive-bias argument that the paper uses to show causal claims depend on theoretical commitments.","marker":"[van Orman Quine, 1976]"},{"why":"Supplies the 'hard core' of scientific paradigms, which the paper uses to locate what counts as a causal mechanism in each domain.","marker":"[Lakatos and Feyerabend, 2019]"},{"why":"Provides the hermeneutic account of truth that the paper relies on for its treatment of social science and interpretive causal narratives.","marker":"[Gadamer, 2013]"},{"why":"Represents the prominent causal representation learning claims that the paper argues overstate what independence-based methods can certify.","marker":"[Schölkopf et al., 2021]"},{"why":"Explicitly associates graphical independence models with causal conclusions, providing the specific overstatement the paper challenges.","marker":"[Eberhardt, 2009]"}],"fun_headline_variants":["Cause is a narrative, not a graph","Causal ML overstates what its DAGs prove","True cause needs evidence across domains","Domain grammar defines cause, not statistics","Causal claims demand cross-domain consensus"],"cache_read_input_tokens":29440,"weakest_assumption_plain":"The argument depends on the premise that the everyday meaning of 'cause' used inside each scientific community is the correct standard for what a true causal claim must mean, which is a normative interpretive choice rather than a mathematical fact.","fun_headline_variants_meta":{"raw":{"variants":["Cause is a narrative, not a graph","Causal ML overstates what its DAGs prove","True cause needs evidence across domains","Domain grammar defines cause, not statistics","Causal claims demand cross-domain consensus"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1282,"prompt_tokens":955,"completion_tokens":327,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":262}},"tokens_in":571,"tokens_out":327,"duration_ms":3348,"temperature":1.0,"reasoning_tokens":262,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:05:23.001627+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single well-documented case where a causal DAG discovered purely from conditional independence in an open biological or social system led to a verified intervention that the field's mechanistic theories had rejected, and where the mechanism was later confirmed, would directly contradict the paper's claim that independence-based discovery cannot certify real causes outside physics and engineering.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Ordinary Language method and the notion of language games that grounds the entire argument about domain-relative meaning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Codifies the formal foundations of causal inference, including structural causal models and additive noise models, which are the target of the critique."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the web-of-knowledge and inductive-bias argument that the paper uses to show causal claims depend on theoretical commitments."},{"cited_title":"and Feyerabend, P","cited_arxiv_id":null,"evidence_quote":"Supplies the 'hard core' of scientific paradigms, which the paper uses to locate what counts as a causal mechanism in each domain."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the hermeneutic account of truth that the paper relies on for its treatment of social science and interpretive causal narratives."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explicitly associates graphical independence models with causal conclusions, providing the specific overstatement the paper challenges."}],"review_version":1}