{"id":"7272d6a7-5784-4096-ae7e-326fa64d3d5b","arxiv_id":"1908.03932","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Under a faithfulness assumption, causal order among observed variables in linear non-Gaussian systems with latent confounders is identifiable, and all observationally equivalent causal effect matrices can be enumerated in polynomial time.","lead":"This paper shows that the order in which visible variables cause each other can still be recovered when hidden common causes exist in a linear non-Gaussian system, as long as no causal effect cancels out. It also lists all possible cause-effect strengths that fit the data, and gives conditions under which they are unique.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The no-cancellation Assumption 1 is the linchpin: Lemma 5 and Theorem 15 read descendant sets off support patterns of B', so an exact cancellation of total effects breaks the causal-order test and the effect-enumeration count.","rationale":"The reader's verdict was CONDITIONAL and identified Assumption 1. I agree that Assumption 1 is the most load-bearing condition, but I would not unambiguously call it stronger than standard d-separation faithfulness; the important point is that the proof's support-to-descendant correspondence is coextensive with it. The synthetic cancellation example shows the failure mode concretely. Since the assumption is explicitly stated and the theory is otherwise coherent, the appropriate verdict remains CONDITIONAL rather than REJECT: the paper should either adopt the no-cancellation assumption as a clearly separated condition and discuss its relation to d-separation faithfulness, or provide a way to test or relax it. The numerical experiments and lack of code or comparison remain secondary concerns already noted by the reader.","tokens_in":16275,"tokens_out":49003,"duration_ms":580134,"concrete_test":"Symbolically compute B' for the three-variable model V1→V3→V2 plus direct V1→V2 with coefficients a, c, b and b=-ac, V3 latent. Form B'', apply Lemma 5's n0*/n*0 rule. If it returns 'no causal path' from V1 to V2, the support test is false whenever Assumption 1 is violated. Then check whether this distribution is faithful in the standard d-separation sense (V1 and V2 are independent here, so it is not); if so, this confirms the paper's 'faithfulness' is a nonstandard no-cancellation assumption and should be stated and defended as such. If instead one can find a d-separation-faithful example with a zero total effect along a path, the concern becomes stronger.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3's path test and Section 4's enumeration both reduce to the statement that the support of a column of B' equals the set of observed descendants of the corresponding variable. This is exactly Assumption 1: [B]_{j,i} is nonzero whenever Vi↝Vj. Lemma 5's inference 'n0*>0 and n*0=0 iff Vi↝Vj' fails if a causal path exists but its total contribution cancels against another path. Concretely, take the DAG V1→V3→V2 and V1→V2 with coefficients a, c, b, where b=-ac, and V3 latent. Then B_{2,1}=0 even though V1↝V2. The columns of B'' are [1,0]^T and [0,1]^T after absorbing the latent column, so the n0*/n*0 test reports no path from V1 to V2. Lemma 1(i) also uses Assumption 1 to propagate nonzero entries along paths; without it, a column can omit a true descendant. Consequently, Theorem 15's count Π r_i, which assumes the only candidates for an observed column are columns with the same support as deso(V_i), inherits the same sensitivity. The paper labels this 'faithfulness,' but the proof requires the support condition, not merely d-separation faithfulness. This is the place where the central claim could break if the assumption is not met.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies linear non-Gaussian acyclic structural equation models with latent variables. Under a support-type faithfulness condition (Assumption 1), it proposes to solve an overcomplete ICA problem on the observed variables, recover the reduced mixing matrix B'', and read the descendant set of each observed variable from the support of the corresponding column. From these descendant sets it obtains a causal order, gives graphical conditions for identifiability of the number of variables, derives a product formula for the number of observationally equivalent total-effect matrices, and provides structural conditions for unique total causal effects. Synthetic experiments and a stock-index application are presented as illustrations.","tokens_in":1435,"tokens_out":1494,"duration_ms":270971,"significance":"If the main claims hold, the paper is a useful contribution: it offers a polynomial-time alternative to the combinatorial search in Hoyer et al. (2008), gives a clean graphical characterization of absorbable latent variables, and carefully demonstrates that total causal effects are not uniquely identifiable in general. The paper is also honest about the non-identifiability example and credits it properly. However, the central identifiability results rest on an assumption that is stronger than the usual faithfulness notion, and the enumeration in Theorem 15 as stated omits observationally equivalent models with absorbable latent variables. These issues are load-bearing and need to be resolved before the results can be accepted at face value.","major_comments":[{"comment":"Assumption 1 is a no-cancellation support condition, not the standard d-separation faithfulness used in causal discovery. The proofs of Lemma 1, Lemma 5, and Theorem 15 all identify the support of a column of B' with the observed descendant set of the corresponding variable. This equivalence fails under exact cancellation of total effects. For example, in the DAG V1 -> V3 -> V2 and V1 -> V2 with structural equations V1 = N1, V3 = a V1 + N3, and V2 = b V1 + c V3 + N2, choosing b = -ac gives [B]_{2,1} = 0 even though V1 reaches V2. The n0*/n*0 test in Lemma 5 then reports no causal path from V1 to V2. If the paper intends to claim identifiability under ordinary faithfulness, this claim is unsupported; if the stronger condition is intended, it should be stated as a separate assumption (e.g., 'no-cancellation faithfulness') and its restrictiveness should be discussed explicitly.","section":"Section 3, Assumption 1 and Lemma 5"},{"comment":"The enumeration formula in Theorem 15 is not correct as stated because the product is computed from B'', the reduced matrix obtained after deleting columns that are proportional to other columns. Models with absorbable latent variables generate exactly the same observed distribution but are not counted. A concrete case is the model of Example 4: V3 latent, V3 -> V1 -> V2. The column of B' corresponding to N3 is proportional to the column corresponding to N1, so after absorption B'' has only the two observed columns and the product in Theorem 15 equals 1. Yet both the original latent-variable model and the equivalent two-variable model are compatible with the same Vo. More generally, latent variables with no observed descendants can be added arbitrarily without changing the observed distribution, so the set of 'all possible D's' is infinite unless the statement is restricted to minimal representations or to a fixed number of variables. The theorem needs an explicit minimality or fixed-dimension restriction, or a revised statement of what is being counted.","section":"Section 4.2, Theorem 15"},{"comment":"The proof of Theorem 15 shows that each of the Pi r_i column selections can be realized by some assignment of the matrices A_oo, A_ol, A_lo, A_ll, but it does not prove completeness, i.e., it does not show that every D generating the same observed distribution corresponds to one of the enumerated selections. In particular, representations for which the associated B' is reducible are not covered by the identifiability argument based on Proposition 3. The completeness step needs to be made explicit, and in light of the previous comment it will require additional assumptions to be true.","section":"Section 4.2, proof of Theorem 15"}],"minor_comments":[{"comment":"In the proof of Lemma 5, the phrase 'there is no causal path between Vi and Vi' should read 'there is no causal path between Vi and Vj'.","section":"Lemma 5 proof"},{"comment":"In the non-identifiability example, the sentence 'The direct causal effects from Vk to Vi, from Vk to Vj, and from Vi to Vi are α, γ, and β, respectively' should say 'from Vi to Vj' rather than 'from Vi to Vi'.","section":"Section 4.1"},{"comment":"The statement that B' is not reducible if and only if the columns of [I | Aol(I-All)^{-1}] are 'not linearly independent' is confusing and, when there are more columns than rows, trivially true. The intended condition is that no two columns are linearly dependent (pairwise linear independence), which is the definition of reducibility used in the paper.","section":"Section 3, paragraph before Example 4"},{"comment":"The experiments use RICA, a heuristic overcomplete ICA method, and the paper does not discuss the conditions under which RICA recovers the true mixing matrix. The empirical results should therefore be presented as illustrative rather than as a validation of the identifiability theorems.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid core idea and several clean results, but the two load-bearing issues above need to be addressed before publication: the strength of Assumption 1 relative to standard faithfulness, and the scope of the enumeration in Theorem 15. I would like the authors either to restrict the claims to minimal or no-cancellation models explicitly and rename the assumption accordingly, or to prove the stronger statements. The stock-index experiment is interesting but not decisive evidence for the theory."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the causal-order identifiability result is real, and the paper is worth serious referee time. The main thing to know is that Assumption 1 is doing more work than the name suggests.\n\nWhat's actually new: under a linear non-Gaussian acyclic model with latent confounders, the paper shows the causal order among observed variables is identifiable, and it gives a polynomial-time enumeration of all observationally equivalent total causal effect matrices. The minimal-graph condition for identifying the number of variables (Theorem 11) is also new. These are genuine advances over Hoyer et al. 2008, which required combinatorial search, and over ParceLiNGAM, which fails on simple graphs like their Figure 1. The proofs in Sections 3 and 4 are mostly coherent; the non-identifiability example is credited correctly, and the ICA reduction is standard.\n\nThe soft spot, as your stress-test notes, is Assumption 1. It is not d-separation faithfulness; it is a no-cancellation condition on total effects. Lemma 5 and Theorem 15 read descendant sets directly off support patterns of B'', so an exact cancellation along multiple paths breaks the path test and the enumeration count. For generic coefficients this is a measure-zero event, so the theory still works in the usual generic sense, but the paper should say so explicitly. Calling it 'faithfulness' is misleading and will confuse readers who expect the standard graphical definition.\n\nTwo more issues, in proportion. Theorem 15's proof is terse: completeness of the enumeration is more asserted than shown, and the step assigning Aol and All needs more detail. The experiments are also weak: no code, no comparison with Hoyer et al. or ParceLiNGAM, and the real-data analysis is qualitative. None of this undermines the central theoretical claim, but it limits how far the practical claims can be trusted.\n\nThis paper is for researchers working on causal discovery with latent confounders, especially those who care about identifiability theory rather than applied benchmarking. It deserves a serious referee. My recommendation: send it out, and ask the authors to clarify the faithfulness assumption, tighten Theorem 15's proof, and add at least one comparison to existing methods on synthetic data.","headline":"Causal-order identifiability under latent variables is a real result, but the 'faithfulness' assumption is actually a no-cancellation condition that the paper should flag clearly.","tokens_in":17104,"tokens_out":2613,"would_cite":true,"duration_ms":30078,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Causal order among observed variables is identifiable even when latent confounders are present, via support patterns of an overcomplete ICA mixing matrix.","keywords":["causal discovery","latent variables","linear non-Gaussian acyclic models","overcomplete ICA","causal order","total causal effects","faithfulness","identifiability"],"falsifier":"Build the linear model $V_1 \\to V_2$ with direct coefficient $1$ and $V_1 \\to V_3 \\to V_2$ with coefficients $2$ and $-0.5$, so the total effect of $V_1$ on $V_2$ is $1 + 2(-0.5) = 0$, with non-Gaussian noises and only $V_1,V_2$ observed. Run the proposed support-detection procedure: if it fails to declare a causal path from $V_1$ to $V_2$, as Lemma 5 would predict when the entry is zero, the method breaks under cancellation, showing that the guarantee depends exactly on Assumption 1.","tokens_in":16042,"feed_emoji":"🔀","tokens_out":7922,"duration_ms":80159,"temperature":0.7,"pith_summary":"This paper claims that even when some causes are unobserved, the causal order among the observed variables in a linear non-Gaussian acyclic model can still be recovered, provided no total causal effect cancels to zero along multiple paths. The route is to run overcomplete independent component analysis on the observed variables and read directed-path information from which entries of the recovered mixing matrix are zero and which are nonzero. Once the order is known, the total causal effects among observed variables are shown to be non-unique in general; the paper gives an efficient way to enumerate exactly the set of all causal-effect matrices compatible with the observed distribution. It also gives a structural condition, that no latent variable may share its set of observed descendants with an observed variable, under which the effects are uniquely identified, and graphical conditions under which the number of variables in the system is identifiable.","feed_headline":"Causal order is recoverable despite latent variables","feed_subtitle":"Non-Gaussian noise exposes directed paths through mixing-matrix support patterns; all compatible effects can be enumerated.","key_machinery":"The load-bearing object is the set of support columns of the overcomplete-ICA mixing matrix $B''$, equivalently the observed-descendant sets $\\mathrm{des}_o(V_i)$. Lemma 5 shows that for any two observed variables, the pattern of zeros in their two rows across all recovered columns, counted as $n_{0*}$ and $n_{*0}$, tells whether one is an ancestor of the other. Absorbing latent variables, defined by merging a latent noise into another variable when all paths from the latent variable to observed variables pass through one node, is the device that decides when the number of variables is identifiable: the graph is minimal when no absorption is possible, and minimality is equivalent to non-reducibility of $B'$ almost surely. The descendant-set matching in Theorem 15 then enumerates all data-compatible total-effect matrices.","core_discovery":"On the paper's own terms, the central discovery is that the support pattern of the mixing matrix obtained from overcomplete ICA is a faithful mirror of ancestor–descendant relations among observed variables. Under Assumption 1, if $V_i\\leadsto V_j$, then exactly one directional pattern occurs: the pair of rows $(i,j)$ in the recovered matrix has some column with a zero in row $i$ and a nonzero in row $j$, and no column with the reverse pattern. This converts causal order discovery into support-pattern inspection. Total effects are not uniquely recoverable in general, but Theorem 15 counts the candidates: there are $\\prod_{i=1}^{p_o} r_i$ matrices $D$ consistent with the data, where $r_i$ is the number of variables, observed or latent, whose observed-descendant set equals that of observed variable $V_i$. Theorem 16 shows that when no latent variable's descendant set coincides with an observed variable's, the total causal effects are unique and are read directly from the normalized column of $\\tilde{B}''$.","pith_inferences":["The asymmetry between $n_{0*}$ and $n_{*0}$ suggests a robustness check: bootstrap the support matrix and require the asymmetry to be stable before declaring a directed path, which would give a direct test of the method's sensitivity to finite-sample ICA errors.","The enumeration of all compatible $D$ matrices gives a natural way to incorporate domain knowledge: any externally motivated constraint on effect signs or magnitudes can prune the product set, which could make the non-uniqueness example practically resolvable.","The method's use of i.i.d. noise and fixed causal structure might transfer to time-series causal discovery if each return series is pre-whitened; the authors' own stock-index experiment suggests this direction, but formal stationarity and lag treatment are not developed in the paper."],"forward_implications":["A practitioner with non-Gaussian observations and hidden common causes can obtain a correct causal order among observed variables without modeling the latents explicitly, using only support recovery from overcomplete ICA.","When the structural condition of Theorem 16 holds, total causal effects are read off from a single normalized ICA column, giving point estimates rather than a set.","The number of latent variables is identifiable almost surely exactly for minimal graphs; latent variables that can be absorbed into another node's noise are genuinely undetectable.","The enumeration in Theorem 15 runs in $O(p_o^2 p_r)$ time, avoiding the combinatorial search over $\\binom{p_r}{p_o}$ latent structures in earlier overcomplete-ICA causal discovery.","Because causal order is obtained before effect estimation, the method can be used as a preprocessing step for downstream effect-size analysis."],"supporting_citations":[{"why":"Introduces the LiNGAM model and the ICA-based identification of the mixing matrix that this paper extends to latent variables.","marker":"Shimizu et al. (2006)"},{"why":"Provides the canonical form for latent confounders and formulates causal effect identification as overcomplete ICA, the starting point for the support-based analysis.","marker":"Hoyer et al. (2008)"},{"why":"Supplies the identifiability theorem (Proposition 3) that the columns of a non-reducible mixing matrix are recovered up to scaling and permutation.","marker":"Eriksson and Koivunen (2004)"},{"why":"Baseline method that recovers unconfounded sets but returns an empty set in the graph of Figure 1, motivating the new order-identification approach.","marker":"Entner and Hoyer (2010)"},{"why":"Extends DirectLiNGAM to latent confounders and fails on Figure 1, giving the comparison case for the proposed method.","marker":"Tashiro et al. (2014)"},{"why":"Provides the reconstruction-ICA algorithm used in the synthetic experiments to solve the overcomplete ICA problem.","marker":"Le et al. (2011)"}],"fun_headline_variants":["Support patterns in mixing matrix reveal causal order despite latents","Non-Gaussian noise exposes causal paths even with hidden causes","When effects are ambiguous, enumerate all compatible causal effects","Even if effects aren't unique, all possible ones are computable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a cause never acts with exactly zero total effect along all its paths to an effect: if two paths from $V_i$ to $V_j$ have coefficients that cancel, the support pattern the method reads will be wrong.","fun_headline_variants_meta":{"raw":{"variants":["Support patterns in mixing matrix reveal causal order despite latents","Non-Gaussian noise exposes causal paths even with hidden causes","When effects are ambiguous, enumerate all compatible causal effects","Even if effects aren't unique, all possible ones are computable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000831,"raw_usage":{"total_tokens":3630,"prompt_tokens":947,"completion_tokens":2683,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2615}},"tokens_in":563,"tokens_out":2683,"duration_ms":22471,"temperature":1.0,"reasoning_tokens":2615,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:59:19.045630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build the linear model $V_1 \\to V_2$ with direct coefficient $1$ and $V_1 \\to V_3 \\to V_2$ with coefficients $2$ and $-0.5$, so the total effect of $V_1$ on $V_2$ is $1 + 2(-0.5) = 0$, with non-Gaussian noises and only $V_1,V_2$ observed. Run the proposed support-detection procedure: if it fails to declare a causal path from $V_1$ to $V_2$, as Lemma 5 would predict when the entry is zero, the method breaks under cancellation, showing that the guarantee depends exactly on Assumption 1.","supporting_citations":[{"cited_title":"Estimation of causal effects using linear non-gaussian causal models with hidden variables","cited_arxiv_id":null,"evidence_quote":"Provides the canonical form for latent confounders and formulates causal effect identification as overcomplete ICA, the starting point for the support-based analysis."},{"cited_title":"Identifiability, separability, and uniqueness of linear ica models","cited_arxiv_id":null,"evidence_quote":"Supplies the identifiability theorem (Proposition 3) that the columns of a non-reducible mixing matrix are recovered up to scaling and permutation."},{"cited_title":"Discovering unconfounded causal relationships using linear non-gaussian models","cited_arxiv_id":null,"evidence_quote":"Baseline method that recovers unconfounded sets but returns an empty set in the graph of Figure 1, motivating the new order-identification approach."},{"cited_title":"Parcelingam: a causal ordering method robust against latent confounders","cited_arxiv_id":null,"evidence_quote":"Extends DirectLiNGAM to latent confounders and fails on Figure 1, giving the comparison case for the proposed method."}],"review_version":1}