{"id":"48a5a31b-a9d8-4f81-aacb-cfe1a1ed2310","arxiv_id":"2511.14441","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"SkewD models skew-normal noise in location-scale noise models to infer cause-effect direction, keeping accuracy high under skewed noise while matching prior methods on standard benchmarks.","lead":"Standard causal-discovery tools assume the randomness in the data is symmetric; when noise is skewed, they often infer the wrong cause-effect direction. This paper introduces SkewD, which models skew-normal noise in location-scale noise models and stays accurate under high skewness.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SKEWD-LL's likelihood score assumes normal marginals; on Tübingen it scores 45.97%, so the 'on par in other settings' half of the central claim is not established for the likelihood variant.","rationale":"Read in good faith: the paper's contribution is modeling skew-normal noise in LSNMs, and the synthetic evidence is strong: SKEWD-LL is perfect across 6 datasets, SKEWD-IT improves with ECM, and out-of-family GNO tests show robustness. I do not object to the identifiability citation or lack of CIs as primary. The marginal-normality issue is the most load-bearing because it directly targets the likelihood variant's validity beyond the synthetic design. The paper itself flags it in the Limitations section. The reader identified the same weakness, so I agree. Since the reader already made the verdict CONDITIONAL, my recommendation is UNCHANGED: the paper should either restrict the likelihood claim to normal-marginal settings or replace the normal-marginal assumption (e.g., with KDE marginals) before claiming on-par performance on real benchmarks.","tokens_in":22321,"tokens_out":13977,"duration_ms":148222,"concrete_test":"Create a synthetic LSNM benchmark identical to the paper's LSs setups except draw the cause from a skewed distribution (e.g., standardized Exp(1) or chi-square_3) instead of N(0,σ²), keeping skew-normal/GNO noise. Run SKEWD-LL, SKEWD-IT, LOCI-LL, and ROCHE on 100 pairs per setting. If SKEWD-LL's accuracy drops substantially while SKEWD-IT and ROCHE do not, the normal-marginal assumption is the load-bearing condition; if SKEWD-LL remains near-perfect, the concern is largely theoretical.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: high-skewness synthetic superiority and on-par performance elsewhere. The first part is credible. The second part rests on SKEWD-LL, whose likelihood score neglects marginal log-likelihoods. Appendix A justifies this only when the true marginals are normal (Proposition 1). In the synthetic skew-noise benchmarks the cause is normal by construction, so the favorable result is exactly inside the assumption; on the Tübingen benchmark, where marginals are non-normal, SKEWD-LL drops to 45.97% accuracy, below every LSNM baseline (ROCHE 77.27%, GRCI 81.59%, Table 6). The authors explicitly list 'lifting the Gaussianity assumption of the marginals for the likelihood-based approach' as a limitation. This does not falsify the claim that skew-normal residuals help under skewed noise, but it means the headline claim 'performing on par with SOTA in other settings' is overstated for SKEWD-LL as published; only SKEWD-IT currently merits that part. The most load-bearing condition is therefore the normal-marginal assumption in Sec. 4/Likelihood Scoring and Appendix A.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SKEWD, a bivariate causal discovery method for location-scale noise models (LSNMs) with skew-normal noise. It models the location and scale functions via penalized B-splines, estimates parameters by a CMA-ES heuristic refined with an ECM algorithm, and performs cause-effect inference either by HSIC-based independence testing (SKEWD-IT) or by likelihood scoring that standardizes variables and neglects marginal log-likelihoods under a normality assumption (SKEWD-LL). The authors introduce synthetic ANs/LSs datasets with skew-normal and generalized-normal noise at skewness levels -0.455, 0.985, and 1.750, and evaluate on 13 established benchmarks. SKEWD-LL achieves 100% accuracy on all six synthetic settings; SKEWD-IT is competitive on benchmarks; SKEWD-LL is much weaker on the real-world Tübingen benchmark (45.97%). The paper concludes that SKEWD is the most reliable LSNM approach under high skewness and on par with state of the art elsewhere.","tokens_in":22596,"tokens_out":7311,"duration_ms":69059,"significance":"The paper addresses a real failure mode: skewed noise degrades Gaussian LSNM estimators, and the proposed skew-normal extension is natural. The empirical evidence on synthetic GNO noise, which is outside the skew-normal family, supports the claim of robustness at high skewness. The ECM ablation (Table 2) demonstrates that the refinement step is necessary, and the authors provide code for reproducibility. However, the headline claim is split: the likelihood variant's marginal-normality assumption is violated on real benchmarks, and the identifiability of the skew-normal LSNM is asserted rather than proven. If these gaps are fixed or the claims are appropriately qualified, the contribution would be a useful step for skew-robust causal discovery.","major_comments":[{"comment":"The likelihood scoring variant assumes that after standardization the marginals are standard normal, so log p(x) and log p(y) cancel. Appendix A's Proposition 1 only establishes this convergence for samples drawn from a normal distribution. On the Tübingen benchmark, where marginals are non-normal, SKEWD-LL achieves 45.97% accuracy (Table 6), the lowest among LSNM baselines (ROCHE 77.27%, GRCI 81.59%). This contradicts the conclusion that SKEWD 'performs on par with SOTA' on established benchmarks for the likelihood variant. The authors list 'lifting the Gaussianity assumption of the marginals' as a limitation, but the conclusion still presents SKEWD as a single reliable method. Please either qualify the on-par claim to SKEWD-IT, or modify likelihood scoring to estimate marginal densities without the normality assumption.","section":"Sec. 4 'Likelihood Scoring'; Appendix A 'Neglection of Marginal Log-Likelihoods'"},{"comment":"The paper states that general LSNM identifiability (Strobl and Lasko 2023; Immer et al. 2023) 'includes the skew-normal LSNM, so that identifiability should hold apart from pathological cases.' However, the cited theorem only gives a PDE condition; the authors do not verify that no skew-normal conditional densities solve this PDE. Since both likelihood scoring and independence testing require identifiability, this is a load-bearing assumption. Please provide a proof, or at least a systematic numerical verification over the skew-normal parameter range, that the non-identifiable ODE/PDE is not satisfied by the skew-normal LSNM.","section":"Sec. 4 'Notes on Identifiability'; Theorem 1 in Appendix A"},{"comment":"The mean accuracy in Table 1 (77.43% for SKEWD-LL) is driven by high synthetic benchmarks; on the real-world Tübingen subset SKEWD-LL is at chance. Reporting a single mean across 13 benchmarks hides this failure mode. The paper's claim that SKEWD is 'performing on par with SOTA in other settings' (Sec. 1 and Sec. 7) is therefore not supported for SKEWD-LL. We recommend reporting per-benchmark comparisons with the likelihood and independence variants separated.","section":"Sec. 6 'Results for Established Benchmarks'; Table 6"}],"minor_comments":[{"comment":"The text says 'We denote the corresponding cause-effect methods as CMA-ES-LL and CMA-ES-LL'; the second should be CMA-ES-IT.","section":"Appendix B, 'Ablation of the ECM Algorithm'"},{"comment":"The method is inconsistently capitalized as 'SkewD' in the abstract and Fig. 2 caption, while the rest of the paper uses 'SKEWD'.","section":"Abstract and Fig. 2 caption"},{"comment":"The text says 'natural cubic B-splines' but the penalty matrices are based on second-order differences of B-spline coefficients (P-splines). Please clarify the spline basis and penalty relationship.","section":"Sec. 2, Eq. (3)"},{"comment":"The skewness formula for the GNO distribution uses both k and κ notation; please unify and ensure parentheses are correct in the displayed equation.","section":"Appendix A, 'Generalized Normal'"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope of a statistical ML journal. The synthetic high-skewness result is credible and the ECM ablation is a strength. The main concerns are the unverified identifiability claim and the normal-marginal assumption behind SKEWD-LL, which undermines part of the headline claim. These are fixable with a proof/verification and with more careful claims or a modified likelihood score."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things. First, the paper delivers a genuinely useful tool: a skew-normal LSNM estimator for bivariate causal discovery, with an independence-test version and a likelihood-scoring version, plus new skewed synthetic benchmarks. The synthetic results are strong and go beyond in-family confirmation: the GNO datasets show SKEWD handling skewness outside the skew-normal family, and the ECM ablation shows the refinement step matters, not just the initial heuristic. Second, the headline claim — most reliable under high skewness, on par elsewhere — is only true in full for the independence-test variant. The likelihood variant (SKEWD-LL) assumes normal marginals in its scoring step, and that assumption is violated on real data: it gets 45.97% on Tübingen, below every LSNM baseline. The authors themselves list lifting the Gaussian-marginal assumption as a limitation, so this is not a hidden flaw, but it does mean the 'on par' half of the central claim is overstated for SKEWD-LL as published.\n\nFor what it is worth, the paper is honest about its scope. The identifiability argument is cited from Immer et al. and Strobl and Lasko rather than verified for this specific family, which is fine because the general LSNM result covers it, but it is worth remembering. The per-dataset results have no confidence intervals, and the repo has no commit hash, so exact reproducibility is a bit loose. None of these are fatal; they are addressable in revision.\n\nWho gets value? Anyone working on cause-effect inference under heteroscedastic or skewed noise, especially in applied settings where symmetric-noise estimators are the default. The novel datasets alone are worth having, and the ECM-vs-CMA-ES ablation is informative for anyone building similar estimators. I would cite this paper, and I would bring it to a reading group, though not this week.\n\nThe paper deserves peer review. The skeptical reader's main worry — the likelihood scorer's marginal-normality assumption — is real but contained, and the independence-test variant gives the core contribution a solid foundation. Send it to a serious referee with a request to check the likelihood-scoring claim and to ask for confidence intervals and a pinned code version.","headline":"Useful, credible extension of LSNM causal discovery to skewed noise, but the headline claim is only fully supported by the independence-test variant; the likelihood scorer leans on a normal-marginal assumption that breaks on real benchmarks.","tokens_in":23118,"tokens_out":1073,"would_cite":true,"duration_ms":14400,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that explicitly modeling skew-normal noise in location-scale noise models removes the failure mode of Gaussian LSNM estimators under skewed noise, and introduces SKEWD to do it.","keywords":["causal discovery","location-scale noise models","skew-normal distribution","bivariate cause-effect inference","skewness","independence testing","likelihood scoring","ECM algorithm"],"falsifier":"Construct LSNM pairs with truly skew-normal noise and causes drawn from a clearly non-normal distribution, such as exponential or bimodal. If SKEWD-LL's direction accuracy collapses while SKEWD-IT stays high, the marginal-normality assumption is implicated; if both collapse, the skew-normal conditional fit is the problem. A direct check is whether replacing the assumed standard-normal marginal with a kernel density estimate restores accuracy on the Tübingen pairs.","tokens_in":22139,"feed_emoji":"📊","tokens_out":5372,"duration_ms":50690,"temperature":0.7,"pith_summary":"The paper tries to establish that the poor performance of existing location-scale noise model (LSNM) estimators under skewed noise is a model misspecification problem, not an identifiability problem, and that it can be removed by modeling the noise as skew-normal instead of normal. To this end it proposes SKEWD, which fits LSNMs with skew-normal noise and chooses direction by residual independence or likelihood scoring. On new synthetic ANM/LSNM datasets with skewness levels -0.455, 0.985, and 1.750, the likelihood variant achieves perfect accuracy in all six settings, while Gaussian-based LOCI drops sharply under high skewness. The paper also reports SKEWD staying competitive on established benchmarks, with the independence-test variant more robust than the likelihood variant on real-world data.","feed_headline":"Skew-normal noise model hits 100% causal accuracy on skewed pairs","feed_subtitle":"Existing location-scale estimators fail when noise is asymmetric; adding one skewness parameter keeps direction recovery reliable.","key_machinery":"The key object is the skew-normal location-scale noise model, Y=f(X)+g(X)N with N~SN(0,1,λ), which makes the conditional Y|X=x skew-normal with location f(x), scale g(x), and shape λ. Because the skew-normal reduces to the normal when λ=0, this is a strict generalization of the Gaussian LSNM used by prior estimators, and identifiability follows from the known LSNM identifiability result except for a pathological differential-equation case. The estimation machinery is a penalized log-likelihood with B-spline bases for f and log g, optimized by Bayesian-tuned CMA-ES and refined by an ECM algorithm that treats the skew-normal as a normal with a latent variable; direction is decided by HSIC inde","core_discovery":"The central claim is that in the model Y=f(X)+g(X)N with independent noise N, replacing the usual normal assumption N~N(0,1) with skew-normal N~SN(0,1,λ) makes cause-effect inference reliable when the noise is asymmetric. SKEWD estimates f and g with penalized B-spline regressions and estimates λ jointly, using a CMA-ES heuristic refined by an ECM algorithm. Inference is then either HSIC-based independence testing on residual-cause pairs or joint likelihood scoring. The paper's key evidence is that SKEWD-LL reaches 100% accuracy on all six synthetic ANs/LSs datasets at skewness -0.455, 0.985, and 1.750, while LOCI-LL drops to 83%, 66%, and other lower values, and ROCHE degrades to 81% on the","pith_inferences":["Editorial inference: because the skew-normal distribution has a maximum skewness of about 0.995, SKEWD cannot exactly match the GNO(1.750) noise used in the highest-skewness datasets; its perfect accuracy there suggests the method exploits the asymmetric conditional shape rather than needing the exact noise family.","Editorial inference: the stark gap between SKEWD-LL's 45.97% accuracy on the Tübingen benchmark and its perfect synthetic results points to the marginal-normality assumption as the load-bearing part; replacing the assumed standard-normal marginal with a nonparametric density estimate is a direct, testable extension.","Editorial inference: the same penalized-B-spline-plus-ECM architecture could be extended to skew-t noise to cover heavy tails and skewness simultaneously, a direction the authors name as future work but do not pursue."],"forward_implications":["If the central claim is right, Gaussian LSNM estimators such as LOCI should not be trusted on data with skewed noise; their residual-independence tests are systematically misspecified.","Skewed noise in LSNMs is not a barrier to identifiability; explicitly modeling the skewness restores the asymmetry that makes the causal direction recoverable.","The likelihood-scoring variant of SKEWD is, on synthetic skewed data, the strongest single configuration, while the independence-test variant is the safer choice on varied real-world benchmarks.","The paper's new synthetic skewed-noise datasets provide a controlled benchmark on which future LSNM methods can be compared at specific skewness levels."],"fun_headline_variants":["SkewD keeps causal discovery accurate under skewed noise","Skew-normal noise model repairs causal direction when data is skewed","SkewD beats existing methods on skewed noise data","Causal direction stays reliable with skewed noise","Skewness-robust causal discovery handles skewed noise"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"For the likelihood-scoring version, the assumption that standardized marginals of cause and effect are approximately standard normal is load-bearing; Proposition 1 proves this only when the true marginal is exactly normal, and the Tübingen accuracy of 45.97% for SKEWD-LL indicates real marginals can violate it.","fun_headline_variants_meta":{"raw":{"variants":["SkewD keeps causal discovery accurate under skewed noise","Skew-normal noise model repairs causal direction when data is skewed","SkewD beats existing methods on skewed noise data","Causal direction stays reliable with skewed noise","Skewness-robust causal discovery handles skewed noise"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000344,"raw_usage":{"total_tokens":1773,"prompt_tokens":841,"completion_tokens":932,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":585,"completion_tokens_details":{"reasoning_tokens":855}},"tokens_in":585,"tokens_out":932,"duration_ms":9079,"temperature":1.0,"reasoning_tokens":855,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:36:37.823453+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct LSNM pairs with truly skew-normal noise and causes drawn from a clearly non-normal distribution, such as exponential or bimodal. If SKEWD-LL's direction accuracy collapses while SKEWD-IT stays high, the marginal-normality assumption is implicated; if both collapse, the skew-normal conditional fit is the problem. A direct check is whether replacing the assumed standard-normal marginal with a kernel density estimate restores accuracy on the Tübingen pairs.","supporting_citations":[],"review_version":1}