Pith. sign in

REVIEW 1 major objections 4 minor 58 references

A deep generative model with a piecewise-affine decoder and a Gaussian-mixture prior is identifiable up to a global affine reparametrization of the latent space, without requiring the decoder to be smooth or injective.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 03:22 UTC pith:LO5INULR

load-bearing objection A genuinely new symmetry-based identifiability hierarchy for PWA-GMM models, but the headline result is narrower than advertised: it only covers estimators that satisfy global PI, which is an estimator-side condition not implied by the data-generating assumptions. the 1 major comments →

arxiv 2607.23182 v1 pith:LO5INULR submitted 2026-07-25 stat.ML cs.LGmath.STstat.TH

Beyond ICA: Identifiability by Symmetry Breaking

classification stat.ML cs.LGmath.STstat.TH MSC 62E10
keywords identifiabilitydeep generative modelspiecewise-affine mapsGaussian mixture modelssymmetry breakingnonlinear ICAnon-injective decodersparameter injectivity
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper proves that a piecewise-affine (PWA) decoder and a Gaussian-mixture latent law are identifiable from unlabeled observations alone, up to a single global affine reparametrization. The engine is algebraic: instead of requiring the decoder to be smooth or injective, the paper imposes three genericity conditions that force the affine symmetry group of the latent law to collapse, so that any two models producing the same distribution must coincide up to that affine map. This matters because it says unsupervised representation learning in this model class is well-posed: distinct mechanisms and distinct latent laws cannot silently generate the same data. The result covers discontinuous and many-to-one decoders, and yields a hierarchy from latent-law identifiability to map, posterior, and pointwise identifiability; injectivity is needed only for the last, pointwise step.

Core claim

The central claim is Theorem 8: if f(Z) and g(Z') induce the same observable distribution, both pairs satisfy parameter injectivity, and the truth pair satisfies universal simple boundary (every branch is witnessed at its own boundary) and mixture symmetry triviality (no nontrivial affine bijection preserves the latent GMM law), then there is a global affine bijection α with α(Z') distributed as Z and f = g ∘ α^{-1}. In words, equality of observables pins down the decoder and the latent distribution up to the same affine reparametrization, and the posterior is identified up to the same map. Smoothness, continuity, and injectivity are not part of the argument; they are replaced by the algebra

What carries the argument

The load-bearing object is the affine symmetry group of the latent GMM: the set of affine bijections T with T♯p = p. Identifiability is achieved by conditions that trivialize this group at three levels. 'Domain contrast' (mixture symmetry triviality) removes nontrivial affine self-symmetries of the latent law; 'mechanism contrast' (universal simple boundary) ensures every decoder branch is witnessed at a boundary where it toggles alone; 'interaction contrast' (parameter injectivity) makes the map from branch–component pairs to pushforward Gaussian parameters injective, so density equality yields a global bijection between the two models' branch-component pairs. Matching at boundaries then st

Load-bearing premise

The assumption that carries the whole result is parameter injectivity on the estimator side: a learned model must have no two branch–component pairs pushing forward to the same Gaussian mean–covariance; if that fails, the global matching bijection breaks and the identifiability hierarchy no longer applies.

What would settle it

Take a truth pair (f,Z) in R^2 with a three-branch PWA decoder and a three-component GMM chosen generically, so parameter injectivity, universal simple boundary, and mixture symmetry triviality hold almost surely, and search for any reduced PWA-GMM pair (g,Z') with the same observable distribution and satisfying parameter injectivity for which no affine bijection α satisfies α(Z')∼Z and f=g∘α^{-1}. The theorem predicts no such pair exists; exhibiting one would refute it. As a cheaper check, fitting (g,Z') from samples and testing whether two distinct branches share a pushforward Gaussian param

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Latent laws and decoders are recoverable up to a global affine map in a purely unsupervised setting, so two fits of the same data cannot disagree by arbitrary nonlinear transformations.
  • Discontinuous and non-injective decoders are harmless for structural identifiability; the posterior remains identifiable up to the affine link, which matters when many latent codes map to one observation.
  • No latent independence assumption is required: arbitrary component covariances are enough; diagonal covariances refine the ambiguity to the ICA form (permutation, scaling, shift).
  • Estimators only need to satisfy parameter injectivity; the other conditions can be guaranteed on the data-generating side, decoupling identifiability theory from architecture constraints like invertibility.
  • Pointwise recovery of the true latent value remains a separate, optional step that requires global injectivity—classical ICA's goal is thus cleanly separated from representation identification.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the framework: train two PWA-decoder GMM models on the same data and check whether their decoders differ by a global affine map; if they do not, one of the three genericity conditions is being violated in the fitted models.
  • The symmetry-collapse mechanism should transfer to other component families, such as exponential families, where the affine group is replaced by the family's natural symmetry group; analogues of (MST), (USB), and (PI) would mark the identifiability boundary.
  • If the theorem holds, multi-mechanism causal discovery becomes feasible: multiple regimes sharing a PWA mechanism could be identified up to one affine ambiguity, giving a nonlinear counterpart to what linear ICA enabled for causal discovery from non-Gaussian data.
  • The estimator-side PI condition is a practicality warning: since it is not implied by the data-generating process, fitting methods that allow branch-component collisions (e.g., degenerate branches or over-parameterized decoders) may fall outside the theorem in practice, and monitoring pushforward parameter collisions is a cheap diagnostic.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. The paper studies identifiability of piecewise-affine (PWA) decoder / Gaussian-mixture (GMM) generative models X=f(Z) from the law of X alone, without observed auxiliary labels. It introduces three conditions: parameter injectivity (PI), simple/universal simple boundary structure (SB/USB), and mixture symmetry triviality (MST). Under distributional equality f(Z)∼g(Z'), with PI on both sides and SB/USB+MST on the truth side, it proves a hierarchy: law identifiability (Theorem 5), map identifiability (Theorem 8, giving an affine α with f=g∘α^{-1} and α(Z')∼Z), posterior identifiability (Prop. 9), pointwise identifiability under injectivity (Prop. 10), and ICA-form refinement under diagonal covariances (Cor. 11). The proofs decompose the observation space into chambers, use analytic continuation to obtain global Gaussian-mixture identities, and then use PI to construct a global branch-component bijection Ψ; boundary richness and MST collapse the remaining branch-wise symmetries. The central result is Theorem 8, which replaces continuity and injectivity with algebraic symmetry conditions.

Significance. If correct, this is a substantial contribution: it moves nonlinear identifiability from differential/continuity arguments to algebraic symmetry collapse, it admits discontinuous and fully non-injective decoders, and it separates law, map, posterior, and pointwise identifiability into a modular hierarchy. The appendix is unusually complete: full proofs of the main results, worked examples showing necessity of (SB) and (MST), explicit comparisons with Kivva et al., and a counterexample (Prop. 18) to several natural strengthenings. The main qualification is that the identifiability theorem is conditional on global PI on both the truth and the estimator side; the paper explicitly acknowledges this in the abstract and in Remark E.1, but the advertised scope is broader than the formal theorem.

major comments (1)
  1. [§4.1 (Condition (PI), Lemma 2), Prop. 18; Abstract/Contributions] The global bijection Ψ in Lemma 2 — and therefore Theorem 8 — requires global (PI) on the estimator side as well as the truth side. This estimator-side condition is not implied by distributional equality, by (USB)+(MST), or by injectivity: Prop. 18 gives an injective reduced PWA map whose pair violates global (PI) while satisfying chamber-wise (PI). The manuscript is transparent in the abstract's 'except for the interaction contrast' and in Remark E.1, but the title and Contributions still advertise identifiability of DGMs / unsupervised identifiability without stating that the model class is restricted to PI-satisfying estimators. This is a load-bearing scope restriction: if a learned or candidate estimator violates global (PI), the global bijection and the entire hierarchy collapse at Lemma 2, even though the truth-side conditions hold. Please revise the central claims to state identif
minor comments (4)
  1. [§5 vs. Appendix L.7] The conclusion that all assumptions 'can simultaneously be argued to be generic' overstates the support for (SB)/(USB). Appendix L.7 explicitly says there is no genericity proof for continuous decoders and that in 1D the failure set of (SB) can have positive measure. The genericity claim should be qualified to the discontinuous/unconstrained case or to the stated open problem.
  2. [§5 and Appendix O] The claimed structural necessity of (MST) relies on localizing finite-order affine symmetries into strict PWA self-maps, but Remark O.1 says the localization proof is 'omitted for scope'. Please either include the proof or explicitly mark that part of the necessity claim as conjectural.
  3. [Cor. 11 / Appendix J.2] The statement assumes (DD) on both Z and Z′, but the proof appears to require only diagonal covariances on the estimator side; distinctness of the spectral ratios is used on the truth side and then transferred via similarity. Consider stating the weaker sufficient condition or explaining why the stronger side is needed.
  4. [Abstract and Related Work] The phrase 'the first to admit discontinuous decoders' should be qualified relative to LID, since the related-work discussion credits Kivva et al. with allowing discontinuity for LID; the genuine novelty is discontinuous decoders for MID.

Circularity Check

0 steps flagged

No circular derivation: the identifiability hierarchy is a conditional theorem with explicit hypotheses; the only self-citation is an application note and is not load-bearing.

full rationale

Walking the derivation chain, I find no step in which a claimed output is identical to an input by construction. Theorem 5 constructs the affine link α explicitly from the parameter equality at a simple boundary, and Theorem 8 reduces the two-map problem to Lemma 7 by defining h = g∘α^{-1} and verifying that (PI) is inherited. The conclusions are conditional on (PI), (SB)/(USB), and (MST), all stated as hypotheses; none is fitted to data and none is derived from the conclusion. The estimator-side (PI) is indeed a restriction on the candidate class—the paper itself says 'Assumptions are only on the data-generating process, not on learning methods, except for the interaction contrast' and later 'Estimators need satisfy only (PI); all other conditions live on the truth side'—but this narrows the quantification of the theorem; it is not circularity, because the theorem's statement includes that hypothesis. The 'uniqueness' step in Lemma 7 uses (MST) directly: the branch-wise difference map is shown to lie in G_mix(p), and (MST) is exactly the assertion that this group is trivial. This is a legitimate application of an assumption, not a renaming of the conclusion. The only self-citation is [40] (Pengzhou Abel Wu & Fukumizu) in the conclusion's causal-inference outlook; it is not load-bearing. External citations such as [38] for Gaussian linear independence are classical mathematical facts. Therefore no circular step is identified; the score reflects only the minor non-load-bearing self-citation.

Axiom & Free-Parameter Ledger

0 free parameters · 8 axioms · 0 invented entities

The central claims rest on standard analytic facts (Gaussian independence, analytic continuation), the restricted model class (reduced PWA + irredundant GMM), and three explicit ad-hoc conditions (PI, USB, MST) that are introduced and assumed by the paper. No free parameters are fitted to data and no new physical or mathematical entities are postulated. The heaviest burden is the estimator-side PI assumption and the unproven localization claim in the necessity discussion.

axioms (8)
  • standard math Gaussian densities with distinct parameters are linearly independent (Lemma 17).
    Used in Lemma 2 to force a bijection between branch-component pairs on each chamber.
  • standard math Real-analytic functions that agree on an open set agree everywhere (Remark 4.1).
    Extends chamber-wise Gaussian-mixture identities from a fine chamber to all of R^d.
  • standard math Finite irredundant Gaussian mixtures have unique decompositions into component parameters and weights.
    Underpins the parameter/weight matching in Lemmas 2 and 13; standard GMM identifiability.
  • domain assumption Reduced PWA maps consist of finitely many invertible affine branches whose domains partition R^d up to measure zero (Definitions 3.1-3.3).
    The entire model class is restricted to invertible affine local rules; non-invertible branch rules are not covered.
  • ad hoc to paper Both truth and estimator satisfy Parameter Injectivity (PI).
    Required for Lemma 2's global bijection; explicitly acknowledged as the one estimator-side assumption.
  • ad hoc to paper Truth latent GMM satisfies Mixture Symmetry Triviality (MST).
    Kills the residual branch-wise affine ambiguity in Theorem 8; argued to be generic but is a substantive condition.
  • ad hoc to paper Truth PWA map satisfies Universal Simple Boundary (USB), hence Simple Boundary (SB).
    Provides a unique toggling branch witness for every branch; generic for discontinuous maps, but open for continuous maps per Appendix L.7.
  • standard math Every nontrivial compact Lie group contains a nontrivial finite-order element.
    Used in Proposition 30 to reduce affine symmetries of GMMs to finite-order symmetries.

pith-pipeline@v1.3.0-alltime-deepseek · 39851 in / 21324 out tokens · 204908 ms · 2026-08-01T03:22:55.095062+00:00 · methodology

0 comments
read the original abstract

We prove the identifiability of deep generative models (DGMs) with piecewise-affine (PWA) decoders and Gaussian mixture model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: domain contrast, which trivializes the mixture symmetry group; mechanism contrast, which ensures every decoder branch is witnessed by a unique boundary; and interaction contrast, which forbids parameter conspiracies between latent components and decoder branches. Together they exploit the interplay between the discrete combinatorics of the PWA map and the continuous symmetry structure of the latent GMM. Continuity is replaced by algebraic symmetry conditions; injectivity is decoupled from structural identification and required only for pointwise inversion. Our results form a hierarchy: from law identifiability (LID; latent distribution up to a global affine map) through map identifiability (MID; decoder up to the same map) to posterior and pointwise identifiability. The ICA-form ambiguity emerges under conditions on diagonal component covariances. Assumptions are only on the data-generating process, not on learning methods, except for the interaction contrast. To our knowledge this is the first to make algebraic symmetry-breaking the engine of nonlinear identifiability, the first to admit discontinuous decoders, and the first to handle fully non-injective decoders, where every observation admits multiple latent codes.

Figures

Figures reproduced from arXiv: 2607.23182 by Pengzhou Wu.

Figure 1
Figure 1. Figure 1: Example 1: f = |x| collapses the symmetric two-component truth Z and the single Gaussian estimator Z ′ to the same observed law. LID fails; the two latents are not related by any affine map. Example 1 (Fold-collapse; LID fails catastrophically). Setup. Take f = g = |x| with branches P1 = [0, ∞), x 7→ x and P2 = (−∞, 0), x 7→ −x. Truth latent: Z ∼ 1 2N (−2, 1) + 1 2N (2, 1). Estimator latent: Z ′ ∼ N (2, 1)… view at source ↗
Figure 2
Figure 2. Figure 2: Example 2 (Kivva D.2): two different PWA maps [PITH_FULL_IMAGE:figures/full_fig_p030_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Example 3 (“Example 101”): both f and g = Id are injective, and f satisfies (USB), yet MID fails. The symmetric latent Z and the (PI) failure of (f, Z) conspire to hide the difference between the two maps. Equal pushforwards. Since Z is symmetric under x 7→ −x, the measure of Z on [−1, 1] is invariant under the reflection performed by branch M. Combined with the identity action of branch O on the tails, f(… view at source ↗
Figure 4
Figure 4. Figure 4: Example 4: the fold f = |x| maps the truth Z and the one-component-reflected Y to the same observed law. (PI) and (MST) hold for both latents, but f fails (SB) and LID fails—showing (SB) cannot be omitted even when the other conditions are in force. Example 4 (Single-flip fold; (PI) and (MST) hold, (SB) fails, LID fails). Setup. Let f(x) = |x| (same (SB)/(USB) analysis as Example 1), and define Z ∼ 1 2N (1… view at source ↗
Figure 5
Figure 5. Figure 5: Example 5: a 2D four-component GMM satisfying (JST) but failing (MST). Components 1 [PITH_FULL_IMAGE:figures/full_fig_p032_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Example A: a continuous, weakly injective PWA map that violates (SB). [PITH_FULL_IMAGE:figures/full_fig_p035_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Example B: a discontinuous 4-branch PWA map satisfying (USB) with no 1-active chamber. [PITH_FULL_IMAGE:figures/full_fig_p036_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Example C: a continuous 4-branch PWA map satisfying (USB) with no 1-active chamber. [PITH_FULL_IMAGE:figures/full_fig_p036_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

58 extracted references · 4 linked inside Pith

  1. [1]

    Properties from mechanisms: An equiv- ariance perspective on identifiable representation learning

    Kartik Ahuja, Jason Hartford, and Yoshua Bengio. Properties from mechanisms: An equiv- ariance perspective on identifiable representation learning. InInternational Conference on Learning Representations, 2022

  2. [2]

    Allman, Catherine Matias, and John A

    Elizabeth S. Allman, Catherine Matias, and John A. Rhodes. Identifiability of parameters in latent structure models with many observed variables.Annals of Statistics, 37(6A):3099–3132, 2009

  3. [3]

    T. W. Anderson and Herman Rubin. Statistical inference in factor analysis.Proceedings of the Third Berkeley Symposium on Mathematical Statistics and Probability, 5:111–150, 1956

  4. [4]

    Zimmermann, Yash Sharma, Bernhard Schölkopf, Julius von Kügelgen, and Wieland Brendel

    Jack Brady, Roland S. Zimmermann, Yash Sharma, Bernhard Schölkopf, Julius von Kügelgen, and Wieland Brendel. Provably learning object-centric representations. InInternational Conference on Machine Learning, 2023

  5. [5]

    Interaction asymmetry: A general principle for learning composable abstrac- tions

    Jack Brady, Julius von Kügelgen, Sébastien Lachapelle, Simon Buchholz, Thomas Kipf, and Wieland Brendel. Interaction asymmetry: A general principle for learning composable abstrac- tions. InInternational Conference on Learning Representations, 2025. arXiv:2411.07784

  6. [6]

    Function classes for identifiable nonlinear independent component analysis

    Simon Buchholz, Michel Besserve, and Bernhard Schölkopf. Function classes for identifiable nonlinear independent component analysis. InAdvances in Neural Information Processing Systems, 2022

  7. [7]

    Independent component analysis, a new concept?Signal Processing, 36(3): 287–314, 1994

    Pierre Comon. Independent component analysis, a new concept?Signal Processing, 36(3): 287–314, 1994

  8. [8]

    Analyse générale des liaisons stochastiques

    Georges Darmois. Analyse générale des liaisons stochastiques. étude particulière de l’analyse factorielle linéaire.Revue de l’Institut International de Statistique, 21(1/2):2–8, 1953

  9. [9]

    Daniel Durstewitz. A state space approach for piecewise-linear recurrent neural networks for identifying computational dynamics from neural measurements.PLoS Computational Biology, 13(6):e1005542, 2017

  10. [10]

    Eaton.Group Invariance Applications in Statistics, volume 1 ofRegional Conference Series in Probability and Statistics

    Morris L. Eaton.Group Invariance Applications in Statistics, volume 1 ofRegional Conference Series in Probability and Statistics. Institute of Mathematical Statistics and American Statistical Association, 1989

  11. [11]

    Fisher III

    Oren Freifeld, Søren Hauberg, Kayhan Batmanghelich, and John W. Fisher III. Transformations based on continuous piecewise-affine velocity fields.IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12):2496–2509, 2017

  12. [12]

    Giri.Group Invariance in Statistical Inference

    Narayan C. Giri.Group Invariance in Statistical Inference. World Scientific, 1996

  13. [13]

    Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Schölkopf

    Luigi Gresele, Paul K. Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Schölkopf. The incomplete Rosetta Stone problem: Identifiability results for multi-view nonlinear ICA. InProceedings of the 35th Conference on Uncertainty in Artificial Intelligence (UAI), 2019. 10

  14. [14]

    Independent mechanism analysis, a new concept? InAdvances in Neural Information Processing Systems, 2021

    Luigi Gresele, Julius von Kügelgen, Vincent Stimper, Bernhard Schölkopf, and Michel Besserve. Independent mechanism analysis, a new concept? InAdvances in Neural Information Processing Systems, 2021

  15. [15]

    Identifiability of models for clusterwise linear regression.Journal of Classification, 17(2):273–296, 2000

    Christian Hennig. Identifiability of models for clusterwise linear regression.Journal of Classification, 17(2):273–296, 2000

  16. [16]

    Towards a definition of disentangled representations

    Irina Higgins, David Amos, David Pfau, Sébastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner. Towards a definition of disentangled representations. InarXiv preprint arXiv:1812.02230, 2018

  17. [17]

    Generalization of the concept of identification

    Leonid Hurwicz. Generalization of the concept of identification. In Tjalling C. Koopmans, editor,Statistical Inference in Dynamic Economic Models, number 10 in Cowles Commission Monograph, pages 245–257. Wiley, 1950

  18. [18]

    Unsupervised feature extraction by time-contrastive learning and nonlinear ICA

    Aapo Hyvärinen and Hiroshi Morioka. Unsupervised feature extraction by time-contrastive learning and nonlinear ICA. InAdvances in Neural Information Processing Systems, 2016

  19. [19]

    Nonlinear independent component analysis: Existence and uniqueness results.Neural Networks, 12(3):429–439, 1999

    Aapo Hyvärinen and Petteri Pajunen. Nonlinear independent component analysis: Existence and uniqueness results.Neural Networks, 12(3):429–439, 1999

  20. [20]

    Deep neural networks learn non-smooth functions effectively

    Masaaki Imaizumi and Kenji Fukumizu. Deep neural networks learn non-smooth functions effectively. InProceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS), 2019

  21. [21]

    Kingma, Ricardo P

    Ilyes Khemakhem, Diederik P. Kingma, Ricardo P. Monti, and Aapo Hyvärinen. Variational au- toencoders and nonlinear ICA: A unifying framework. InProceedings of the 23rd International Conference on Artificial Intelligence and Statistics (AISTATS), 2020

  22. [22]

    Identifiability of deep generative models without auxiliary information

    Bohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, and Bryon Aragam. Identifiability of deep generative models without auxiliary information. InAdvances in Neural Information Processing Systems, 2022

  23. [23]

    Joseph B. Kruskal. Three-way arrays: Rank and uniqueness of trilinear decompositions, with application to arithmetic complexity and statistics.Linear Algebra and its Applications, 18(2): 95–138, 1977

  24. [24]

    Additive decoders for latent variables identification and cartesian-product extrapolation

    Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, and Simon Lacoste-Julien. Additive decoders for latent variables identification and cartesian-product extrapolation. InAdvances in Neural Information Processing Systems, 2023

  25. [25]

    Linderman, Matthew J

    Scott W. Linderman, Matthew J. Johnson, Andrew C. Miller, Ryan P. Adams, David M. Blei, and Liam Paninski. Bayesian learning and inference in recurrent switching linear dynamical systems. InProceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), 2017

  26. [26]

    Challenging common assumptions in the unsupervised learning of disentangled representations

    Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem. Challenging common assumptions in the unsupervised learning of disentangled representations. InInternational Conference on Machine Learning, 2019

  27. [27]

    Pritchard, and Aviv Regev

    Romain Lopez, Jan-Christian Huetter, Ehsan Hajiramezanali, Jonathan K. Pritchard, and Aviv Regev. Toward the identifiability of comparative deep generative models. InConference on Causal Learning and Reasoning (CLeaR), 2024

  28. [28]

    Mechanistic independence: A principle for identifiable disentangled representations.arXiv preprint arXiv:2509.22196, 2025

    Stefan Matthes, Zhiwei Han, and Hao Shen. Mechanistic independence: A principle for identifiable disentangled representations.arXiv preprint arXiv:2509.22196, 2025

  29. [29]

    Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio

    Guido F. Montúfar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio. On the number of linear regions of deep neural networks. InAdvances in Neural Information Processing Systems, 2014

  30. [30]

    Juloski, Giancarlo Ferrari-Trecate, and René Vidal

    Simone Paoletti, Aleksandar Lj. Juloski, Giancarlo Ferrari-Trecate, and René Vidal. Identifica- tion of hybrid systems: A tutorial.European Journal of Control, 13(2–3):242–260, 2007. 11

  31. [31]

    Geoffrey Roeder, Luke Metz, and Diederik P. Kingma. On linear identifiability of learned representations. InInternational Conference on Machine Learning, 2021

  32. [32]

    Daniel Rueckert, L. I. Sonoda, C. Hayes, D. L. G. Hill, M. O. Leach, and D. J. Hawkes. Nonrigid registration using free-form deformations: Application to breast MR images.IEEE Transactions on Medical Imaging, 18(8):712–721, 1999

  33. [33]

    Toward causal representation learning.Proceedings of the IEEE, 109(5):612–634, 2021

    Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. Toward causal representation learning.Proceedings of the IEEE, 109(5):612–634, 2021

  34. [34]

    Hoyer, Aapo Hyvärinen, and Antti Kerminen

    Shohei Shimizu, Patrik O. Hoyer, Aapo Hyvärinen, and Antti Kerminen. A linear non-gaussian acyclic model for causal discovery.Journal of Machine Learning Research, 7(72):2003–2030,

  35. [35]

    V . P. Skitovich. On a property of the normal distribution.Doklady Akademii Nauk SSSR, 89: 217–219, 1953

  36. [36]

    Deformable medical image registration: A survey.IEEE Transactions on Medical Imaging, 32(7):1153–1190, 2013

    Aristeidis Sotiras, Christos Davatzikos, and Nikos Paragios. Deformable medical image registration: A survey.IEEE Transactions on Medical Imaging, 32(7):1153–1190, 2013

  37. [37]

    Non-rigid registration for two-photon imaging using triangulation and piecewise affine transformation.Neuroscience, 491:86–99, 2022

    Feng Su and Yonglu Tian. Non-rigid registration for two-photon imaging using triangulation and piecewise affine transformation.Neuroscience, 491:86–99, 2022

  38. [38]

    Identifiability of finite mixtures.Annals of Mathematical Statistics, 34(4): 1265–1269, 1963

    Henry Teicher. Identifiability of finite mixtures.Annals of Mathematical Statistics, 34(4): 1265–1269, 1963

  39. [39]

    Generalized principal component analysis (GPCA)

    René Vidal, Yi Ma, and Shankar Sastry. Generalized principal component analysis (GPCA). IEEE Transactions on Pattern Analysis and Machine Intelligence, 27(12):1945–1959, 2005

  40. [40]

    $\beta$-intact-V AE: Identifying and estimating causal effects under limited overlap

    Pengzhou Abel Wu and Kenji Fukumizu. $\beta$-intact-V AE: Identifying and estimating causal effects under limited overlap. InInternational Conference on Learning Representations, 2022. URLhttps://openreview.net/forum?id=q7n2RngwOM

  41. [41]

    Indeterminacy in generative models: Characterization and strong identifiability

    Quanhan Xi and Benjamin Bloem-Reddy. Indeterminacy in generative models: Characterization and strong identifiability. InProceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS), 2023

  42. [42]

    Identifiability of potentially degenerate gaussian mixture models with piecewise affine mixing, 2026

    Danru Xu, Sébastien Lachapelle, and Sara Magliacane. Identifiability of potentially degenerate gaussian mixture models with piecewise affine mixing, 2026. arXiv:2604.13218

  43. [43]

    Yakowitz and John D

    Sidney J. Yakowitz and John D. Spragins. On the identifiability of finite mixtures.Annals of Mathematical Statistics, 39(1):209–214, 1968

  44. [44]

    in particular

    Yujia Zheng and Kun Zhang. Generalizing nonlinear ICA beyond structural sparsity. In Advances in Neural Information Processing Systems, 2023. 12 A Proofs of main results Lemma 12 (Single-chamber PW A is globally affine) Lemma 12.Let h:R d →R d be a reduced PWA map with invertible affine branch rules. If h has exactly one chamber, then h has exactly one br...

  45. [46]

    almost arbitrary support

    (Type H) C2-diffeo additive de- coder + block latents Map (additivity + suffi- cient nonlinearity) differential (block-wise)× × ×(C 2 diffeo)✓ (“almost arbitrary support”) ✓ Brady et al. [4] (Type D/S) diffeo compositional de- coder + slot latents Interaction (compositional- ity + irreducibility onJ f ) differential (Jacobian spar- sity) ×(slot blocks)× ×...

  46. [47]

    2) n/a× (multi-env

    (Type M) smooth, possibly non- bijective + grouped la- tents Map (structural / partial sparsity); Law (partial dep.) differential + sparsity× × × (undercomplete, still inj.) ✓(explicit)✓ Xi and Bloem- Reddy [41] injective f + unspecified prior family Taxonomic: A(F)∩ A(Pz) algebraic taxonomy; measure-theoretic intersec- tion ×n/a×(Asn. 2) n/a× (multi-env....

  47. [48]

    hidden duplicate parameter in a different branch

    (iV AE) smooth invertible f + EF prior cond. on aux Map (smooth/inv.); Law (EF + aux variability) differential (Jacobian + suf- ficient stats) × × × × (cond. indep. given aux) ×(aux observed) Gresele et al. [13] (Multi-View NICA) smooth invertible views + indep. noise Law (indep. + SDV on cond. log-density); Map (smooth/inv.) differential (log-density fac...

  48. [49]

    But the reflection T(x) =−x+ (µ 1 +µ 2) swaps the two components; with equal weights, T♯p=p , so (MST) fails

    Then Sym1 ∩Sym 2 ={Id} : (JST) holds. But the reflection T(x) =−x+ (µ 1 +µ 2) swaps the two components; with equal weights, T♯p=p , so (MST) fails. If a decoder branch is post-composed with T , the output distribution is unchanged but the map has been modified. (JST) cannot detect this switcheroo; (MST) can. This is why our main-text MID result uses (MST)...

  49. [50]

    (DW) (distinct weights).Under (DW), any affine symmetry must fix every component individually (Proposition 20), so (MST) reduces to (JST)

  50. [51]

    ,det ΣK are pairwise distinct, any affine symmetry inducesσ= id, again reducing (MST) to (JST)

    Distinct covariance determinants.By Lemma 21, if det Σ1, . . . ,det ΣK are pairwise distinct, any affine symmetry inducesσ= id, again reducing (MST) to (JST). I.2 Sufficient conditions for (JST) Two sufficient conditions for (JST): 1.Affinely independent means. 25 Lemma 22(Affinely independent means ⇒ (JST)).If the means {µk}K k=1 are in affinely general ...

  51. [52]

    distinct det(Σ) + 2-component generic position

    Set O:= Σ −1/2 1 AΣ1/2 1 . Then O∈O(d) , and b= (I−A)µ 1. Fixing the second component gives OHO ⊤ =H and Ov=v . Since O−1 =O ⊤, the first identity implies OH=HO . Because H has simple spectrum, every orthogonal matrix commuting with H is diagonal with entries ±1 in an eigenbasis of H. In that basis, v has no zero coordinate; therefore Ov=v forces every si...

  52. [53]

    4) forbid component-swapping, and no single reflection fixes both components individually

    Both Z and Y satisfy (MST): the distinct variances (1 vs. 4) forbid component-swapping, and no single reflection fixes both components individually

  53. [54]

    Both pairs (f, Z)and (f, Y)satisfy (PI): the four pushforward parameters N(±1,1) , N(±3,4) are pairwise distinct

  54. [55]

    4) forbid component-swapping, and neither α(x) =x+2 nor α(x) =−x+2 is consistent across both components simultaneously

    No affine bijection α satisfies α(Y)∼Z : distinct variances (1 vs. 4) forbid component-swapping, and neither α(x) =x+2 nor α(x) =−x+2 is consistent across both components simultaneously. Hence LID fails despite (PI) and (MST) holding on both sides. The obstruction is entirely on the map side: f=|x| has no simple boundary (point (1) of Appendix K.1’s proof...

  55. [56]

    Both latent mixtures satisfy (MST)

  56. [57]

    Both pairs(f, Z)and(f, Y)satisfy (PI)

  57. [58]

    weakly-injective tail attached to a co-toggling core

    There is no affine bijectionα∈Aff(R)such thatα(Y)∼Z. Consequently, law identifiability fails although latent (MST) and (PI) both hold. Proof. Step 1: equal pushforwards.Because f(x) =|x| , f♯ N(1,1) =f ♯ N(−1,1) . The second componentN(3,4)is the same in both mixtures, sof(Z)∼f(Y). Step 2: (SB) fails for f.The two branch images are both [0,∞) . Hence the ...

  58. [2006]

    URLhttp://jmlr.org/papers/v7/shimizu06a.html