{"id":"597653a4-b87b-407d-a117-9117c4998eb2","arxiv_id":"2411.18656","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Machine learning and deep learning applications can repeat the errors of physiognomy and Lombrosianism because they treat correlations as causal, and bias reduction alone does not fix this.","lead":"This paper argues that many machine learning systems, especially deep learning, mistake correlation for causation and can revive discredited pseudosciences like physiognomy and Lombrosianism. It recommends rethinking AI models and metrics rather than only cleaning training data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The harm estimate in §4.2 treats every false-positive AI flag as a wrongful conviction, skipping the arrest/prosecution/adjudication chain; the magnitude claim is unsupported, though the qualitative thesis survives.","rationale":"The paper is a position essay, not an empirical study, and its central argument is coherent: predictive models built on correlations can be deployed in ways that reproduce pseudoscientific reasoning, and debiasing training data alone does not guarantee that a model is safe or valid. The reader's weakest-assumption analysis correctly identifies the load-bearing gap in the quantitative harm argument. The simulation in Section 4.2 is transparent arithmetic, but the text overstates it by converting false-positive flags into 'wrongly convicted people' without evidence of the intervening legal process. This matters because the paper's conclusion leans on the scale of harm to advocate for a fundamental rethinking of ML models; if the scale is materially smaller, the policy urgency is weakened, though not eliminated. The qualitative thesis—that high accuracy metrics do not justify causal or ethical claims, and that bias curation alone is insufficient—remains intact. I considered whether the paper's 'correlation does not imply causation' framing conflates prediction with causal inference, but the paper explicitly attributes the problem to 'designers and final users' who attribute causality, not to the models' internal epistemology, so that concern is less decisive. The reader's verdict of CONDITIONAL is appropriate: the paper should correct or clearly caveat the wrongful-conviction framing before the quantitative claims are used in policy discussions.","tokens_in":13211,"tokens_out":5922,"duration_ms":60091,"concrete_test":"Recompute Table 3's London row at 95% accuracy/precision/recall using an empirically measured conversion rate from a false-positive risk flag to a wrongful conviction. For a concrete check: obtain a matched cohort from the UK Ministry of Justice's evaluation of OASys, link high-risk risk scores to subsequent custodial sentences while controlling for human officer recommendations, and estimate the probability that a false-positive flag directly caused a conviction. If that probability is, say, 10% or less, the '4800–9600 wrongly convicted' claim drops below a few hundred and should be relabeled as an upper bound on flags, not convictions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.2 computes numbers of false-positive classifications for a one-time screening of each city's population, then immediately interprets the London result as '4800 to 9600 wrongly convicted people.' This leap equates an automated flag with a wrongful conviction without any evidence that a false flag leads to arrest, charge, prosecution, or conviction. In the cited deployed systems (e.g., OASys in the UK), the AI output is a risk score that feeds a human probation officer's decision, not a direct conviction. The paper's own assumptions state that each person is tested once and that algorithm predictions are consistent, so Table 3 at best counts distinct individuals who would be flagged, not individuals who would be convicted. The unexamined causal chain from model output to legal outcome is load-bearing because the paper uses these figures to argue that high-accuracy correlational models 'would be a serious danger if implemented at large scale' and to justify 'a complete rethinking' of model design. Without the conviction step, the quantitative urgency is unsupported, although the qualitative claim that biased or invalid target constructs can cause harm does not depend on the exact numbers.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that modern machine learning and deep learning systems, by relying on correlational patterns without causal understanding, risk reviving pseudoscientific practices such as physiognomy, Lombrosianism, and social astrology. It surveys controversial applications in justice, surveillance, and profiling; critiques the 'theory-free' ideal and the sufficiency of bias-reduced training data; and presents a quantitative simulation in Section 4.2 that estimates false-positive counts for hypothetical crime-prediction systems in four cities. The paper concludes that high-accuracy models can cause large-scale harm and calls for rethinking model design, using harm-relevant metrics, and maintaining human oversight.","tokens_in":13397,"tokens_out":3064,"duration_ms":29013,"significance":"The paper addresses an important and timely topic: the historical continuity between discredited pseudosciences and some contemporary AI applications. Its central conceptual claim—that high accuracy on correlational tasks does not justify causal or high-stakes decisions—is sound and well-aligned with existing critiques in the ML fairness and causality literature. The paper usefully catalogs problematic applications and clearly explains why accuracy and recall alone are inadequate for evaluating harm. The quantitative section, however, contains a load-bearing overstatement that currently undermines the paper's credibility, as it equates false-positive flags with wrongful convictions without any evidence of the intervening legal process. If that overstatement is corrected, the paper could serve as an effective perspective piece for a broad ML audience.","major_comments":[{"comment":"The statement that London 'would potentially feature 4800 to 9600 wrongly convicted people' equates a false-positive classification with a wrongful conviction. The computation in Table 3 counts distinct individuals flagged by a one-time hypothetical screening under assumed accuracy, precision, and recall; it does not count arrests, prosecutions, or convictions. No evidence is provided that a false flag from a CCTV or risk-assessment system would lead to conviction, especially since the cited deployed systems (e.g., OASys) are advisory risk scores feeding human decisions. This leap is load-bearing because the paper uses these numbers to argue that high-accuracy models 'would be a serious danger if implemented at large scale' and to justify a 'complete rethinking' of model design. The text should be revised to say 'wrongly flagged individuals' or 'false positive identifications,' and the downstream decision chain should be discussed explicitly.","section":"§4.2, Table 3"},{"comment":"The simulation's headline numbers are highly sensitive to assumptions that are not defended. First, the crime rate estimates are taken from unspecified 'estimates available online' and vary by two orders of magnitude across cities (0.56 to 93 per 1000), which directly drives the false-positive counts. Second, the assumption that accuracy, precision, and recall are equal is acknowledged but not varied independently, even though real classifiers often have trade-offs between precision and recall; lower precision at fixed accuracy would sharply increase false positives. Third, the 'Number of misclassified people' rows are trivially N*(1-accuracy) and do not inform harm. The paper should either provide a sensitivity analysis over realistic ranges or clearly state that the table illustrates a hypothetical scenario, not a prediction.","section":"§4.2, Eqs. (6)–(7), Table 3"},{"comment":"The claim that 'input variables are de facto causal variables for Machine Learning Models to make their decisions' is imprecise and overstates the matter. A discriminative classifier learns a function f(x) that maps features to labels; it does not itself assert a causal relation. The causal misattribution typically occurs when designers or users interpret the model's predictions as causal or when they choose features based on causal assumptions. The paper's argument would be more rigorous if it distinguished between the model's statistical dependence and the human interpretation or deployment that turns correlation into an implicit causal claim. This distinction is central to the 'correlation does not imply causation' thesis and should be clarified.","section":"§3.1"}],"minor_comments":[{"comment":"There are numerous typographical errors and missing words: 'real-wold' (abstract), 'and orientation' should be 'and sexual orientation' (Section 1), 'ssystems' (Section 2), 'countries countries' (Section 2), 'a,d' should be 'and' (Section 3.2), 'bow' should be 'now' (Section 3.2), 'mitiate' should be 'mitigate' (Section 4.1), 'Mister Lombroso' should be 'Lombroso' (Section 3.2), 'ethic courses' should be 'ethics courses' (Section 6), and 'humans oversight' should be 'human oversight' (Section 6).","section":"Throughout"},{"comment":"The citation placeholder '[61 ? ]' appears in the sentence 'biases have shown to be a problem [61 ? ]'; this should be resolved to a proper reference or removed.","section":"Section 5"},{"comment":"Some references have malformed URLs or missing publisher locations (e.g., [6], [44], [45]); these should be cleaned up for publication.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is an opinion/perspective paper rather than a traditional technical contribution. The quantitative simulation in Section 4.2 is presented as a central piece of evidence, but its conclusions overreach the arithmetic. The paper would be more credible if the simulation were substantially revised to focus on 'false positive flags' and accompanied by a careful discussion of the legal and decision-making context. The journal's editorial team may also want to consider whether the paper's scope fits a stat.ML readership or whether it would be better placed in an ethics or interdisciplinary venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it is a position essay, not an empirical study, and it does not claim to be one. Second, its most memorable numbers—thousands of \"wrongly convicted people\" in London—are not supported by the paper's own arithmetic, because they skip the entire chain from an automated flag to a conviction. The stress-test note is right about that, and it is the paper's real soft spot.\n\nWhat the paper does well: it gives a clear, historically informed mapping of current ML applications to physiognomy, Lombrosianism, and related pseudosciences, and it makes a sensible argument that bias-curation alone cannot fix problems rooted in the models' implicit causal claims. Section 5 on the \"theory-free\" ideal is the strongest part and engages seriously with philosophy of science. Table 3 is transparent: the reader can see exactly how the false-positive counts are computed from the stated assumptions. There is no fitted-parameter circularity; it is simple arithmetic, and the paper says so. Credit is also due for citing Andrews et al. (2024) and other prior critiques rather than pretending the thesis is new. The novelty is limited—the core idea is already in the literature—but the synthesis and the cross-city illustration are a legitimate extension.\n\nThe soft spots are proportionate. The main one is Section 4.2. The paper assumes a one-time screening of each city's population and then interprets the false-positive count as people \"wrongly convicted.\" An automated flag is not a conviction; deployed systems like OASys feed a human decision. The paper itself acknowledges some limitations but not this one, and the magnitude claim is load-bearing for the policy urgency. The crime-rate estimates are taken from online sources without citations, and the assumption that accuracy, precision, and recall are equal is stated but not justified. These are fixable with more careful language and better sourcing. There are also a few typos and awkward sentences, but they do not obscure the argument.\n\nFor whom is this paper? People working on AI fairness, ethics, or ML evaluation who want a compact, historically grounded critique to assign in a seminar or cite in a broader discussion. It is not a technical contribution, but it does not pretend to be.\n\nMy recommendation: send it to peer review. It deserves a serious referee, not a desk reject. A good reviewer will ask the authors to revise Section 4.2 so the harm estimates are framed as \"flags requiring investigation\" rather than convictions, and to source or caveat the crime-rate figures. The central thesis is coherent and worth engaging with, and the paper is honest about its own scope.","headline":"A coherent and readable position essay arguing that ML can revive pseudoscience by treating correlations as causal, but its central harm estimate overreaches by equating false-positive flags with wrongful convictions; the qualitative thesis survives the flaw.","tokens_in":13926,"tokens_out":1313,"would_cite":true,"duration_ms":14579,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Deep learning revives pseudosciences by mistaking correlation for causation","keywords":["correlation vs causation","deep learning ethics","pseudoscience revival","physiognomy","Lombrosianism","algorithmic harm","quality metrics","theory-free AI"],"falsifier":"A concrete test would be to examine the actual false positive and conviction rates of deployed AI systems like OASys or similar risk-assessment tools in the UK or US, comparing the number of false positives to the number of wrongful convictions actually caused by these systems. If the conviction rate from automated flags is negligible, the paper's quantitative harm claims would be overstated.","tokens_in":12982,"feed_emoji":"🧠","tokens_out":1654,"duration_ms":17089,"temperature":0.7,"pith_summary":"This paper argues that machine learning and deep learning systems, despite their high accuracy, routinely confuse correlation with causation, and that this confusion effectively resurrects pseudosciences such as physiognomy, Lombrosianism, and social astrology. The author contends that the primary ethical danger of AI is not just biased training data but the models themselves, which implicitly treat input features as causal variables. Using examples from justice, surveillance, and facial analysis, the paper claims that even well-intentioned bias-reduction efforts are insufficient. The paper also shows that standard quality metrics like accuracy and recall hide the harm caused by false positives, and it quantifies this harm in four major cities. The conclusion is that safe AI requires rethinking model design, adopting harm-sensitive metrics, and keeping human oversight.","feed_headline":"AI's accuracy hides a pseudoscience revival, paper argues","feed_subtitle":"High-scoring deep learning models still mistake correlation for causation, and that, not bias, is the real danger.","key_machinery":"The central mechanism is the implicit reversal of causal direction in deep learning inference: models are trained to map features to labels, so they treat symptoms, facial traits, or background data as causes of the predicted outcome, whereas in reality the direction of causation may be opposite or absent. The paper also uses a simple algebraic argument with classification metrics (accuracy, recall, precision) to show how false positive counts scale with population size, crime rate, and model precision, giving concrete estimates of harm for four cities. The 'theory-free' myth—the idea that data-driven models are value-free and unbiased—is identified as the ideological support that legitimizes these pseudoscientific applications.","core_discovery":"The central claim is that current machine learning and deep learning methods, because they are fundamentally inductive statistical tools, implicitly infer causation from correlation: input variables become de facto causal variables in their decision processes. This implicit causation, combined with the neglect of domain knowledge and historical context, revives discredited pseudosciences such as physiognomy, Lombrosianism, and social astrology. The paper further argues that the 'theory-free' ideal of data-driven AI is a fallacy, that bias removal from training data cannot achieve fairness because all data are biased and models can reconstruct sensitive features, and that the prevailing emphasis on accuracy and recall metrics obscures the real social harm of false positives, particularly in criminal justice and security applications. The paper demonstrates this harm with a quantitative simulation for London, Beijing, Hyderabad, and New York, showing that even at 95% accuracy, precision, and recall, thousands of false positive identifications would occur, potentially leading to thousands of wrongful convictions.","pith_inferences":["The paper's harm simulation could be extended to a formal analysis of how precision and recall trade-offs affect false positive counts under different base rates; the author's assumption of equal accuracy, precision, and recall is conservative, and real systems with lower precision would show even larger harms.","The argument suggests a concrete testable hypothesis: if deep learning models are indeed treating features as causes, then interventions that change the causal structure of the data (e.g., do-calculus style manipulations) should change model predictions in ways that are inconsistent with purely correlational learning.","The paper's critique of 'theory-free' AI could be connected to the broader debate on whether AI explainability methods can recover causal structure; one could test whether post-hoc explanation tools actually reveal causal mechanisms or merely rationalize correlations.","The author's conclusion implies that regulatory frameworks should mandate not only fairness audits but also harm-based impact assessments that quantify false positive rates in the deployment context, not just in the training distribution."],"forward_implications":["If the paper is right, high-accuracy AI systems deployed at scale in policing and justice will produce large absolute numbers of false positives, potentially leading to wrongful arrests or convictions, even in the absence of deliberate bias.","Fairness interventions that only curate training data or remove sensitive features will fail to prevent harm, because models can reconstruct protected attributes from correlated features and because bias is inherent to the application itself.","Evaluation practices need to shift from accuracy and recall toward precision and specificity, and toward metrics that explicitly capture the social cost of false positives, especially in high-stakes domains.","AI systems should not be treated as oracles that replace domain experts; human oversight and integration of domain knowledge are necessary to avoid repeating historical errors of pseudoscience.","The revival of physiognomy and Lombrosianism through facial-analysis and criminal-prediction AI is not a marginal curiosity but a systemic risk that demands a rethinking of core model design, not just data curation."],"supporting_citations":[{"why":"Bengio et al. on representation learning are cited to support the claim that deep learning treats input variables as de facto causal variables in its decision process.","marker":"[42]"},{"why":"Bengio et al. on disentangling causal mechanisms is used to argue that current deep learning lacks explicit causal reasoning, underpinning the paper's central critique.","marker":"[43]"},{"why":"Andrews, Smart, and Birhane on the reanimation of pseudoscience in machine learning is the direct predecessor that the paper builds on to argue that ML is not only vulnerable to pseudoscience but also legitimizes it.","marker":"[49]"},{"why":"Anderson's 'end of theory' essay is cited as the 'theory-free' ideology that the paper criticizes, providing the straw man it argues against.","marker":"[54]"},{"why":"Wang and Kosinski's study on detecting sexual orientation from facial images is a primary example of physiognomy-like AI that the paper uses to illustrate the revival of pseudoscience.","marker":"[32]"},{"why":"Kabir et al. on 'human abnormality' classification is cited as an example of Lombrosian-style criminal tendency detection from faces.","marker":"[25]"},{"why":"Hamilton and Ugwudike on the OASys black-box AI system in UK criminal justice is used to show that such systems are already in use, grounding the paper's concern in real deployments.","marker":"[9]"},{"why":"Howard and Dixon's validation of the OASys violence predictor is used to illustrate the use of AI in justice, showing the type of system the paper argues harms people via false positives.","marker":"[10]"},{"why":"Mehrabi et al.'s survey on bias and fairness in machine learning is used to represent the mainstream fairness approach that the paper argues is insufficient.","marker":"[57]"},{"why":"Wang et al.'s work on balanced datasets being insufficient is cited to support the paper's claim that removing sensitive features or rebalancing data does not eliminate bias.","marker":"[63]"}],"fun_headline_variants":["AI's accuracy hides a causal fallacy, paper argues","Correlation to pseudoscience: AI's dangerous oversight","Deep learning revives pseudoscience by ignoring statistics","Why AI's success ignores a core statistical lesson","Bias removal can't fix AI's causal mistakes, study warns"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative harm estimates assume that a false positive from an AI system, such as a CCTV flag or risk-assessment score, is equivalent to a wrongful conviction, without evidence that an automated flag leads to a legal conviction.","fun_headline_variants_meta":{"raw":{"variants":["AI's accuracy hides a causal fallacy, paper argues","Correlation to pseudoscience: AI's dangerous oversight","Deep learning revives pseudoscience by ignoring statistics","Why AI's success ignores a core statistical lesson","Bias removal can't fix AI's causal mistakes, study warns"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2826,"prompt_tokens":938,"completion_tokens":1888,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1811}},"tokens_in":554,"tokens_out":1888,"duration_ms":14327,"temperature":1.0,"reasoning_tokens":1811,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:28:24.075775+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to examine the actual false positive and conviction rates of deployed AI systems like OASys or similar risk-assessment tools in the UK or US, comparing the number of false positives to the number of wrongful convictions actually caused by these systems. If the conviction rate from automated flags is negligible, the paper's quantitative harm claims would be overstated.","supporting_citations":[{"cited_title":"IEEE transactions on pattern analysis and machine intelligence 35(8), 1798–1828 (2013)","cited_arxiv_id":null,"evidence_quote":"Bengio et al. on representation learning are cited to support the claim that deep learning treats input variables as de facto causal variables in its decision process."},{"cited_title":"Patterns (2024)","cited_arxiv_id":null,"evidence_quote":"Andrews, Smart, and Birhane on the reanimation of pseudoscience in machine learning is the direct predecessor that the paper builds on to argue that ML is not only vulnerable to pseudoscience but also legitimizes it."},{"cited_title":"Wired magazine 16(7), 16–07 (2008)","cited_arxiv_id":null,"evidence_quote":"Anderson's 'end of theory' essay is cited as the 'theory-free' ideology that the paper criticizes, providing the straw man it argues against."},{"cited_title":"Journal of personality and social psychology 114(2), 246 (2018)","cited_arxiv_id":null,"evidence_quote":"Wang and Kosinski's study on detecting sexual orientation from facial images is a primary example of physiognomy-like AI that the paper uses to illustrate the revival of pseudoscience."},{"cited_title":"In: 2020 IEEE 17th International Conference on Smart Communities: Improving Quality of Life Using ICT, IoT and AI (HONET), pp","cited_arxiv_id":null,"evidence_quote":"Kabir et al. on 'human abnormality' classification is cited as an example of Lombrosian-style criminal tendency detection from faces."},{"cited_title":"The Conversation (2023)","cited_arxiv_id":null,"evidence_quote":"Hamilton and Ugwudike on the OASys black-box AI system in UK criminal justice is used to show that such systems are already in use, grounding the paper's concern in real deployments."},{"cited_title":"Criminal Justice and Behavior 39(3), 287–307 (2012) https://doi","cited_arxiv_id":null,"evidence_quote":"Howard and Dixon's validation of the OASys violence predictor is used to illustrate the use of AI in justice, showing the type of system the paper argues harms people via false positives."},{"cited_title":"ACM computing surveys (CSUR) 54(6), 1–35 (2021)","cited_arxiv_id":null,"evidence_quote":"Mehrabi et al.'s survey on bias and fairness in machine learning is used to represent the mainstream fairness approach that the paper argues is insufficient."},{"cited_title":"In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp","cited_arxiv_id":null,"evidence_quote":"Wang et al.'s work on balanced datasets being insufficient is cited to support the paper's claim that removing sensitive features or rebalancing data does not eliminate bias."}],"review_version":1}