{"id":"de211119-93e6-4672-baa6-792d9449270e","arxiv_id":"2608.13229","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper rigorously proves that Gaussian-free independent sources in a linear ICA model are identifiable up to permutation, scale and translation even with arbitrary additive Gaussian noise.","lead":"This paper presents a self-contained mathematical treatment of the identifiability theory behind linear independent component analysis (ICA), proving when the independent sources can be recovered from mixtures. It also analyzes the stability of the standard equivariant gradient descent algorithm, showing that it converges to separating solutions exactly when the model is identifiable.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the Gaussian-free scope restriction is explicit and does not threaten the theorem as stated.","rationale":"The paper's central claim is a conditional identifiability theorem, and I checked the proof chain from Theorem 4.2 through Theorem 5.5 to Theorem 6.12. The column-dichotomy proof in Section B is detailed and internally sound; the Gaussian splitting in Section 6 is correct, with uniqueness up to translation established by a characteristic-function argument. The most delicate step, the cancellation in Theorem 6.12 where φ_{A(1)Z(1)} may vanish, is handled locally on a ball before extending the Gaussian polynomial identity, so the zeros do not break the argument. The Gaussian-free assumption is the only real restriction: it is strictly stronger than non-Gaussianity and excludes natural Gaussian-mixture sources, but this is openly stated and the theorem is formulated as a conditional statement. No red flags, no circular reasoning, and no overclaim are present. Hence the reader's ACCEPT verdict needs no adjustment; confidence remains moderate because the proofs are not machine-checked.","tokens_in":63998,"tokens_out":14914,"duration_ms":156367,"concrete_test":"Re-derive Theorem 6.12 for p=k=1 with Z(1) uniform on [-1,1], which is Gaussian-free but has characteristic function sin(t)/t vanishing at t=π, and with E(1)~N(0,1). Verify that the proof cancels φ_{A(1)Z(1)} only on a neighbourhood of the origin, not globally, and that the conclusion still holds; a global cancellation would be invalid for this valid example.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim is Theorem 6.12, and its proof is internally coherent: the Gaussian-free hypothesis is used exactly to rule out the residual Gaussian factors that Theorem 5.5 leaves, and the local cancellation around zeros of φ_{A(1)Z(1)} is handled by working on a neighbourhood of the origin before extending the Gaussian polynomial identity globally. The Gaussian splitting theorem, Theorem 6.7, and the finite-difference proof of Theorem 4.2 are self-contained and do not contain a hidden assumption. The one genuine limitation is that Gaussian-freeness is strictly stronger than non-Gaussianity and fails for common models such as mixtures of Gaussians, but this is explicitly acknowledged in Example 6.10(d) and Remark 6.13, and the theorem is formulated as a conditional statement. No red flags, no circularity, and no overclaim are present.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a self-contained mathematical treatment of linear independent component analysis (ICA). It develops the theory of characteristic functions of probability measures on R^d, including analyticity, cumulants, and the Gaussian characterisation theorems, and then proves a sequence of identifiability results for the linear ICA model under successively stronger source assumptions: non-constant, non-Gaussian, and Gaussian-free. The central result is Theorem 6.12, which identifies Gaussian-free independent sources up to permutation, scale, and translation even in the presence of additive Gaussian noise with an arbitrary covariance matrix. The second half of the paper studies the complete noiseless square ICA model, derives the relative-gradient (equivariant) online algorithm, proves local stability of separating solutions under explicit conditions on model scores, and derives LiNGAM identifiability as a corollary. Full proofs are provided in two appendices, with the Kagan–Linnik–Rao theorem proved by a self-contained finite-difference argument.","tokens_in":64098,"tokens_out":30618,"duration_ms":295629,"significance":"If the results are correct, this is a valuable rigorous reference for the mathematical foundations of ICA. The paper's main strengths are its completeness: the identifiability statements are proved from first principles, the source assumptions and remaining ambiguities are stated precisely, and the proof of the central Theorem 6.12 is internally coherent, with the Gaussian-free hypothesis used exactly where it is needed. The treatment of splittable Gaussian scales and the maximal Gaussian decomposition (Theorem 6.7) is a useful clarification of a notion that is often left informal. The paper also gives a careful account of the relationship between the relative gradient and the natural metric, and it states explicitly which results are quoted rather than proved. The contribution is more expository and foundational than revolutionary, but it meets a genuine need for a precise, self-contained presentation of the classical identifiability theory and its modern refinements.","major_comments":[],"minor_comments":[{"comment":"The proof asserts that for a continuous map T one has supp L(T(Z)) = T(supp L(Z)) and refers to T(supp L(Z)) as a closed set. This is not true in general: a linear map need not be a closed map, since a continuous image of a closed set need not be closed (for instance, a projection of the closed hyperbola xy=1 has non-closed image). The conclusion of the proposition is nevertheless correct; the proof should argue directly that aff(supp L(X)) = T(aff(supp L(Z))) using the fact that affine hulls commute with affine maps and are unchanged by taking closures.","section":"Section 4, Proposition 4.4"},{"comment":"The sentence 'the G_j were constructed componentwise, so we may and do take G to have independent components' deserves an explicit justification. The componentwise identities in Eq. (136) fix only the marginal laws of the components of Z(2); to pass to the vector identity Eq. (143), one should state that the G_j can be recoupled independently on an enlarged probability space, with Z(2) then defined by Eq. (143), and that the resulting vector has independent components with the required marginals. As written, the step is correct but requires the reader to reconstruct the coupling argument.","section":"Section 6.2, Theorem 6.12 proof"},{"comment":"The claimed stable non-separating equilibrium for k=2 with the specific value R* = 0.80993 is stated without derivation. Likewise, the entries in Table 2 are said to be obtained by numerical quadrature but no code or computational details are supplied. A short explanation of how these quantities were computed, or a reference to reproducible code, would strengthen the presentation.","section":"Section 7.4, after Corollary 7.23"},{"comment":"The statement that the online stochastic approximation algorithm converges almost surely to locally stable equilibria of the ODE under 'the usual regularity and boundedness conditions' is informal and is not proved. Since Theorem 7.20 establishes local stability only for the mean dynamics Eq. (202), the remark should either state precise hypotheses sufficient for the stochastic approximation result or explicitly label that convergence statement as heuristic.","section":"Section 7.4, Remark 7.27"}],"recommendation":"minor_revision","confidential_remarks":"The central mathematical claims appear sound, and the manuscript is a careful foundational contribution that fits the journal's scope. The minor issues listed above are local and do not affect the main theorems. The editor may wish to consider whether the essentially expository nature of parts of the paper is appropriate for the journal, but the technical level and completeness are high."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a review in the best sense — it re-proves the known ICA identifiability theorems from first principles, and the proofs are actually good. The reader's report is right: the central results are Comon 1994 and Eriksson-Koivunen 2004 restated, and Theorem 4.2 goes back to Kagan, Linnik and Rao. If you need new theorems, this isn't where to find them.\n\nBut what the paper does well is substantive. The self-contained finite-difference proof of the Kagan–Linnik–Rao column dichotomy (Appendix B) is the real contribution — it's a genuinely hard classical argument, and the version here is readable and complete. The Gaussian-free framework in Section 6, with the maximal Gaussian scale and the splitting theorem, is a clean way to organize the noise-identifiability results. The reinstated proportionality constant in Theorem 4.2 is a minor correction, but a correct one, and the authors say exactly where the earlier statement lost it. The stability analysis in Section 7 is a careful recasting of Amari et al. 1997, properly credited; the table of stability values is useful.\n\nSoft spots, in proportion. The novelty really is low. The paper presents itself as a note for readers with measure-theoretic probability background, and it is honest about being a survey. The Gaussian-free assumption is strictly stronger than non-Gaussianity — mixture-of-Gaussians sources are excluded — but the paper states this openly in Example 6.10(d) and Remark 6.13, and the theorem is conditional on it. That's a scope limitation, not a flaw. The regularity conditions in Theorem 7.20 (bounded second derivative of the score, third moments) exclude the cubic nonlinearity, which is handled by a separate polynomial argument; minor. A couple of standard external theorems (Bochner, unstable manifold) are quoted without proof, which is fine and disclosed.\n\nCitation pattern is clean. Self-citations appear only in the survey section on generalizations, where they are appropriate. No circularity — the identifiability theorems are proved from characteristic-function theory, not assumed.\n\nWho it's for: graduate students and researchers who want one rigorous place to send people for the ICA identifiability basics; teachers of courses on blind source separation or causal discovery. If the venue is a research journal expecting novelty, the paper will need framing as an expository/foundations contribution — but it deserves a referee, not a desk reject. I'd accept it with an expectation of minor revision, mostly tightening the scope claims.","headline":"A careful, self-contained re-proof of the classical ICA identifiability theory with a clean Gaussian-free formulation; not much new, but the rigor and clarity earn it a serious referee.","tokens_in":64685,"tokens_out":3620,"would_cite":true,"duration_ms":33483,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H25","60E10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that linear ICA is identifiable up to permutation, scale, translation and sign whenever the sources are mutually independent and Gaussian-free, even under additive Gaussian noise with arbitrary covariance.","keywords":["independent component analysis","identifiability","Gaussian-free sources","characteristic functions","cumulant generating functions","Kagan–Linnik–Rao theorem","equivariant gradient descent","LiNGAM"],"falsifier":"Try to find two full-rank representations of the same observed law, A¹Z¹+E¹ = A²Z²+E², where every source in both representations is Gaussian-free in the sense that no non-degenerate Gaussian convolution factor can be split off, and check whether the sources are related only by permutation, scale, and translation. Exhibiting such a pair without this relation—or computing for a concrete law that the set of splittable Gaussian scales is not a closed interval, contradicting Lemma 6.6(iii)—would refute the paper's central claim.","tokens_in":63765,"feed_emoji":"🧮","tokens_out":8101,"duration_ms":75664,"temperature":0.7,"pith_summary":"The paper's central claim is that linear independent component analysis is identifiable in the strongest sense—sources recoverable up to translation, permutation, scale, and sign—provided the sources are mutually independent and Gaussian-free, meaning that no non-degenerate Gaussian distribution can be split off from any source as a convolution factor, and then even when the observed mixture carries additive Gaussian noise with arbitrary, possibly degenerate covariance. The authors prove this by building the full theory of characteristic functions from scratch, then proving a two-sided identifiability theorem (Theorem 6.12) whose engine is a decomposition result: every real-valued random variable splits essentially uniquely into a Gaussian-free part plus independent Gaussian noise, and the Gaussian-free parts are what the mixture pins down. Along the way they prove the classical Kagan–Linnik–Rao dichotomy with a self-contained finite-difference argument, and show that merely non-Gaussian sources leave a residual additive Gaussian ambiguity, while Gaussian sources are not identifiable at all. The paper also derives the online equivariant gradient descent algorithm and shows that its separating fixed point is locally stable exactly when the identifiability theory declares the model identifiable. A sympathetic reader should care because these results delineate precisely when blind source separation is solvable in principle and what ambiguity must remain.","feed_headline":"ICA recovers sources uniquely even with arbitrary Gaussian noise","feed_subtitle":"New proof shows one condition makes blind source separation uniquely solvable.","key_machinery":"The argument is carried by characteristic functions and their distinguished logarithms (cumulant generating functions), combined with three classical rigidity results: Marcinkiewicz' theorem, which says the exponential of a polynomial is a characteristic function only in the Gaussian case; Cramér's decomposition theorem, which says a Gaussian sum can only have Gaussian independent summands; and the Kagan–Linnik–Rao theorem (Theorem 4.2), which this paper proves by a finite-difference argument over ridge functions, such that any column of one mixing matrix that is not proportional to a column of the other forces its source to be Gaussian. Around that core the paper wraps Lemma 5.2, which trades Gaussian noise vectors for extra columns of the mixing matrix and back, and Theorem 6.7, which splits every source into a Gaussian-free part plus independent Gaussian noise. The estimation half then shows that the natural-gradient (relative gradient) update for the mixing matrix induces a dynamics on the global system matrix whose stability condition—in terms of the quantity ζj = −βjσj²—matches the identifiability condition exactly when the model score equals the true score.","core_discovery":"On the paper's own terms, the central discovery is Theorem 6.12: under mutual independence, Gaussian-free sources, and full column rank of the mixing matrix, two representations of the same observed law must be related by a permutation, a diagonal scaling, and a translation, with both the source laws and the noise covariance then forced to agree. The key structural insight is Theorem 6.7, which shows that every real-valued random variable decomposes, essentially uniquely, as a Gaussian-free random variable plus independent Gaussian noise, and that the maximal Gaussian scale is always attained. This makes 'Gaussian-free' the exact hypothesis that removes the additive Gaussian ambiguity that remains when sources are merely non-Gaussian (Theorem 5.5), and it explains why Gaussian sources are hopeless: they can be rotated by any orthogonal matrix without changing the observed law. In the complete noiseless case, the theory specialises to the classical statement that a square invertible mixture is identifiable if and only if at most one source is Gaussian (Corollary 7.4), and the same machinery proves LiNGAM's causal order is identified.","pith_inferences":["The Gaussian-free hypothesis is strictly stronger than non-Gaussianity—mixtures such as ½N(−1,1)+½N(1,1) are non-Gaussian but not Gaussian-free—so methods that check σmax(Z)=0 (e.g., via characteristic-function decay or support) could be used to certify in advance whether the stronger identifiability conclusion applies to a given dataset.","The splitting theorem suggests a natural quantitative measure of residual ambiguity: the maximal Gaussian scale σmax(Z) acts as a 'Gaussian content' of a source, and one could design partial identifiability statements for misspecified models that bound how much of the source still can be attributed to noise.","The stability quantity ζj for a misspecified score is a computable functional of the source law; one could adaptively choose the nonlinearity per source, tuning it so that ζj>1 holds, which would make the algorithm's convergence guarantee data-dependent rather than assumed.","The finite-difference proof of the column dichotomy is self-contained and may carry over to identifiability questions beyond linear ICA, such as nonlinear or time-varying mixtures, where the same 'every unwanted direction is annihilated by a difference operator' strategy could be replayed."],"forward_implications":["If both candidate source vectors are Gaussian-free, the whole model—mixing matrix, source laws, and noise covariance—is identified up to permutation, scale, and shift, even when the additive Gaussian noise is degenerate and arbitrarily correlated across coordinates.","Merely non-Gaussian sources are not enough: the sources are then determined only up to an additive componentwise Gaussian noise, so the practical lesson is that blindly applying ICA to non-Gaussian but Gaussian-contaminated sources overstates what can be recovered.","In the complete noiseless square case, identifiability holds if and only if at most one source is Gaussian, and the recovered sources are determined up to permutation and sign once centred and scaled.","The online equivariant gradient descent algorithm attains the same boundary: with correctly specified source scores, the separating solution is locally asymptotically stable exactly when the model is identifiable, so identifiability and algorithm stability coincide.","LiNGAM removes the permutation ambiguity for acyclic causal models, so the causal order and coefficients are fully identified from the law of the observations."],"supporting_citations":[{"why":"Supplies the fundamental identifiability dichotomy for independent non-constant sources (Theorem 4.2) that Sections 4–7 branch from.","marker":"Kagan et al., 1973"},{"why":"The theorem that an exponential of a polynomial is a characteristic function only for Gaussian laws; it collapses the finite-difference output to degree two.","marker":"Marcinkiewicz, 1939"},{"why":"The decomposition theorem showing a Gaussian sum has only Gaussian independent factors, used in the Gaussian-splitting and noise-trading arguments.","marker":"Cramér, 1936"},{"why":"One of the two classical sources of the column dichotomy—if two linear forms have independent components then non-proportional coefficients force Gaussianity.","marker":"Darmois, 1953"},{"why":"The other classical source of the same dichotomy, used in the comparison with Theorem 4.2 in Remark B.7.","marker":"Skitovich, 1954"},{"why":"Establishes the complete noiseless ICA identifiability statement (Corollary 7.4) as the square invertible specialisation.","marker":"Comon, 1994"},{"why":"Formulates identifiability, separability and uniqueness of linear ICA models, framing the problem the paper solves.","marker":"Eriksson and Koivunen, 2004"},{"why":"Introduces the relative-gradient learning rule whose dynamics and stability the paper analyses in Section 7.","marker":"Amari et al., 1996"},{"why":"Defines equivariant adaptive source separation, the algorithm class whose stability Theorem 7.20 analyses.","marker":"Cardoso and Laheld, 1996"},{"why":"Provides the LiNGAM causal model whose identifiability is derived as Corollary 7.32.","marker":"Shimizu et al., 2006"}],"fun_headline_variants":["Gaussian-free sources: the key to unique ICA","ICA identifiability solved: one assumption does it","Unique source recovery in ICA, noise and all","Why Gaussian-free is enough for ICA identifiability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof of the strongest theorem (Theorem 6.12) collapses if a source can be written as a Gaussian plus something independent—i.e., if it fails to be Gaussian-free—because then the residual Gaussian ambiguity of the merely non-Gaussian case survives.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian-free sources: the key to unique ICA","ICA identifiability solved: one assumption does it","Unique source recovery in ICA, noise and all","Why Gaussian-free is enough for ICA identifiability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000776,"raw_usage":{"total_tokens":3411,"prompt_tokens":900,"completion_tokens":2511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":516,"completion_tokens_details":{"reasoning_tokens":2449}},"tokens_in":516,"tokens_out":2511,"duration_ms":18736,"temperature":1.0,"reasoning_tokens":2449,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:47:02.489659+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Try to find two full-rank representations of the same observed law, A¹Z¹+E¹ = A²Z²+E², where every source in both representations is Gaussian-free in the sense that no non-degenerate Gaussian convolution factor can be split off, and check whether the sources are related only by permutation, scale, and translation. Exhibiting such a pair without this relation—or computing for a concrete law that the set of splittable Gaussian scales is not a closed interval, contradicting Lemma 6.6(iii)—would refute the paper's central claim.","supporting_citations":[],"review_version":1}