Pith. sign in

REVIEW 2 major objections 3 minor 61 references

Exogenous Isomorphism for Counterfactual Identifiability

T0 review · 2 major / 3 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read This paper proves that Markovian triangular monotonic SCMs sharing a causal order and an observational distribution are indistinguishable at every level of the Pearl causal hierarchy, making all counterfactual questions answerable from…

desk verdict Solid theoretical advance on full L3 identifiability for triangular monotone SCMs, but the experiments leak ground-truth counterfactuals into model selection. read the letter →

arxiv 2505.02212 v1 pith:ZDRFFBBI submitted 2025-05-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords counterfactualidentifiabilityPearlcausalhierarchyexogenousisomorphismtriangularmonotonicSCMtransportKRneuralmodelsMarkovianassumption
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper targets complete counterfactual identification: guaranteeing that every structural causal model consistent with the stated assumptions returns the same answer to every Pearl-hierarchy level-3 query. To make the problem tractable it introduces exogenous isomorphism, a componentwise bijection between noise variables that preserves both mechanisms and the noise distribution, and shows that exogenously isomorphic models are indistinguishable at level 3. The central result, Corollary 5.4, is that any Markovian triangular monotonic SCM is identifiable from its causal order and observational distribution up to this isomorphism. This means two such models that match the observable world cannot disagree on any what-if question, which extends and unifies earlier identifiability theorems for bijective, monotone, and fixed-point causal models. A reader should care because it identifies a practical class of neural causal models on which every counterfactual answer, not just a single outcome, can be trusted from observational data alone.

What carries the argument

The load-bearing object is the triangular monotonic (TM) mapping: a coordinate-wise transformation in which component $j$ is strictly monotone in coordinate $j$ and depends only on earlier coordinates. TM maps are bijective, and they are closed under inversion, composition, and taking contiguous subcomponents, with a well-defined monotonicity signature. The crucial identity is that composing one TM map with the inverse of another of the same signature always gives a strictly increasing triangular map, and for distributions with strictly positive density such maps coincide almost surely with the KR transport. In a TM-SCM every counterfactual transport is therefore an increasing triangular map; Markovianity ties the relevant conditional distributions to the observed ones; and KR transport's uniqueness forces all models in the class to share the same counterfactual transport. That shared transport induces exogenous isomorphism, which Theorem 3.2 converts into full level-3 consistency.

What would settle it

Generate pairs of Markovian TM-SCMs with vector-valued variables from the same causal order and identical conditional densities evaluated on a fine grid, then compute the same nested counterfactual probability, such as $P(V_1[x_1]\in A, V_3[x_3]\in B \mid V_2=v)$, for each pair by exact simulation. Corollary 5.4 predicts the two values coincide up to numerical tolerance; the central claim is refuted if any such pair differs beyond tolerance. The most informative search varies the monotone mechanisms while holding all conditional densities fixed, since that is the only degree of freedom the theory says does not matter.

Watch

Extended reading notes

Core claim

Within the Pearl Causal Hierarchy the counterfactual layer encodes all causal information, so two SCMs that agree on all level-3 statements are indistinguishable for every causal question. The paper's discovery is an equivalence relation, called exogenous isomorphism, that captures this kind of identifiability without forcing a unique latent representation: it requires only a componentwise bijection between the exogenous variables, with the exogenous distribution and each causal mechanism preserved through that bijection. Theorem 3.2 shows this relation implies level-3 consistency for arbitrary recursive SCMs. For bijective SCMs, Theorem 4.6 shows that fixing counterfactual transports is enough to force exogenous isomorphism, and the KR-transport version, Theorem 4.8, reduces the needed data to conditional distributions. Corollary 5.4 then gives the headline condition: a Markovian triangular monotonic SCM is identifiable, and hence completely counterfactually identifiable, from its causal order and observational distribution alone, a guarantee that goes beyond earlier counterfactual-outcome identifiability results.

Load-bearing premise

The guarantee rests on knowing the true causal order and the exact coordinate alignment of each variable, and on the true model being Markovian with every mechanism strictly monotone in its own noise; if any of these is misspecified, the transport recovered from the observational distribution is not the true counterfactual transport and counterfactual answers can be wrong.

Editorial extensions

If this is right

  • Two Markovian TM-SCMs with the same causal order and the same observational distribution must give identical answers to every level-3 query, including nested counterfactual conjunctions, not just single counterfactual outcomes.
  • Earlier identifiability theorems for monotone state transitions, bijective causal mechanisms, and fixed-point causal generative models become special cases of Corollary 5.4, with endogenous variables no longer required to be scalar.
  • For a neural TM-SCM trained by maximum likelihood, convergence to the observed distribution is enough to guarantee that its counterfactual predictions match those of the true model, as long as the four structural assumptions hold.
  • Counterfactual systems do not need to recover the true noise variables; any componentwise bijective encoding that preserves mechanisms and exogenous distribution is sufficient for complete counterfactual identifiability.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same argument should extend to any other query expressible in level 3, such as path-specific effects and probabilities of causation, because the guarantee concerns the whole theory of the model rather than a selected query; the paper states this consequence but does not single out those applications.
  • A testable practical extension is to estimate the causal order from data and then feed it to a neural TM-SCM; the paper's order-reversal ablation predicts that order misspecification will degrade counterfactual accuracy even when observational fit is good, so an uncertainty-aware order selection step is a natural next component.
  • Because exogenous isomorphism only needs the existence of a componentwise bijection, a practitioner can deliberately relabel noise variables, for instance by standardizing them, without changing any causal answer; the normalizing-flow choice in the neural models exploits precisely this freedom.
  • The boundary of the guarantee can be probed by replacing triangular monotone mechanisms with general bijective autoregressive mechanisms: monotonicity is what makes the counterfactual transport agree with KR transport, so outside the TM class two models with the same observational distribution can be expected to diverge on counterfactual queries.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The paper studies full counterfactual identifiability, denoted ~L3-identifiability, within the Pearl Causal Hierarchy. It introduces a model-level equivalence relation called exogenous isomorphism (~EI) and proves that ~EI-identifiability implies ~L3-identifiability. The authors then give sufficient conditions for ~EI-identifiability in two classes: bijective SCMs (BSCMs), via the concept of counterfactual transport and optimal-transport uniqueness, and triangular monotonic SCMs (TM-SCMs), where the counterfactual transport is shown to coincide with the Knothe-Rosenblatt transport. The main theoretical corollary states that any Markovian TM-SCM is ~EI-identifiable from the causal order, the Markov assumption, and the observational distribution alone, which strengthens known counterfactual-outcome identifiability results to full L3-consistency. The paper also proposes neural TM-SCM implementations (DNME, TNME, CMSM, TVSM) and reports synthetic experiments that claim to validate the theory.

Significance. If the theory is correct, Corollary 5.4 together with Theorem 3.2 is a substantial advance: it unifies and generalizes several prior results (Lu et al., 2020; Nasr-Esfahany et al., 2023; Scetbon et al., 2024) by extending identifiability from counterfactual outcomes to the entire L3 layer. The formal apparatus is a strength: the definitions are measure-theoretically explicit, the appendix provides a detailed dependency graph and proof structure, and the authors make a credible case that their results subsume earlier counterfactual-equivalence notions while being weaker than full model identifiability. The KR-transport and TM-SCM arguments form a coherent chain. However, the empirical validation, which is advertised as supporting the practical claim of observational-only counterfactual consistency, is undermined by oracle-based model selection; the experiments therefore do not currently provide the advertised evidence. The theoretical contribution is valuable and likely defensible, but the experimental section needs substantive repair.

major comments (2)
  1. [Appendix D.3 and Section 8] The experiments do not currently establish that observational training alone achieves L3-consistency. Appendix D.3 explicitly states: 'The model weights corresponding to the epoch with the lowest CTF RMSE on the validation set are saved for testing.' CTF RMSE (Appendix D.2) is computed against ground-truth counterfactual outcomes. Therefore the test CTF RMSE in Table 1, Table 6, and Figure 1 is obtained by oracle selection over training epochs, which leaks counterfactual information into model selection. The non-ablated models and all ablations are selected by the same oracle criterion, so the comparisons in Table 1 do not isolate the effect of the assumed structural condition on L3-consistency from the effect of selection. Please re-run the experiments with a model-selection rule based only on observational data (for example, validation OBS WD or the NLL), or explicitly report the performance at a fixed epoch without oracle selection.
  2. [Appendix A.3, Theorem 4.3] The proof of the reverse direction of Theorem 4.3 is incomplete as written. It states 'By Lemma A.8 and Lemma A.11, we have Γ(1)=Γ(2)∘h', but Lemmas A.8 and A.11 are proved under the full exogenous-isomorphism assumptions, which include the mechanism isomorphism that the reverse direction is supposed to establish. What is needed is an induction over the common causal order, propagating the component-wise relation (f_i^(2)(v,·))^{-1}∘(f_i^(1)(v,·))=h_i to prefix potential responses; this gap appears fillable but should be spelled out. The statement should also clarify how the 'almost surely' quantification interacts with the quantification over all v∈Ω_V.
minor comments (3)
  1. [Section 8 and Figure 1] The text says that the w/o O and w/o T configurations 'fail to converge', but in Figure 1 the w/o O curves appear flat with a high CTF RMSE rather than divergent; please describe the behavior more precisely, e.g., 'the CTF RMSE does not decrease as the observational fit improves'.
  2. [Table 1] The table has formatting issues: the best model entries (for example, DNME '-0.53 ±0.05') contain a stray leading dash that is likely a LaTeX minus-sign artifact. In addition, some confidence intervals are very wide (e.g., TNME w/o O, 11.24±20.98); the paper should comment on the instability behind these intervals rather than only reporting the means.
  3. [Definition 4.4] In Definition 4.4 the notation K_M(·,v,v′) overloads the first argument: for a fixed pair (v,v′), the displayed object is a map on the third Ω_V factor. Using distinct symbols or explicitly distinguishing the three factors would improve readability.

Circularity Check

1 steps flagged · score 5.0 of 10

Theoretical identifiability chain is self-contained, but the experimental validation is circular: the saved epoch is selected by the same ground-truth counterfactual metric that is then reported as consistency.

  1. fitted input called prediction [Section 8 (Experiments); Appendix D.2 (Metrics) and D.3 (Execution)]
    "CTF RMSE measures the error between the counterfactual results inferred by the model and the ground truth, assessing the accuracy of the trained model in counterfactual reasoning. ... The model weights corresponding to the epoch with the lowest CTF RMSE on the validation set are saved for testing."

    The paper claims to demonstrate counterfactual consistency using only observational samples as the training set, but the checkpoint is chosen by minimizing CTF RMSE on the validation set, and CTF RMSE is computed against ground-truth counterfactual outcomes. Thus the reported test CTF RMSE, and the ablation comparisons built on it, evaluate a model selected by the very counterfactual-consistency target under test. The empirical claim therefore reduces partly to the availability of counterfactual labels for model selection, rather than being derived from observational data and the identifiability theorem alone. This does not refute the mathematical derivation, but it invalidates the experimental support for the claim that observational training alone yields the reported L3-consistency.

full rationale

The mathematical derivation chain is not circular. Theorem 3.2 is proved by unrolling causal mechanisms into potential responses and then using the exogenous-distribution isomorphism to equate L3-evaluations; it does not assume L3-consistency. Corollary 5.4 is obtained by showing that in a Markovian TM-SCM each counterfactual transport component is a TMI mapping, and by Lemma 5.1 such a mapping coincides with the KR transport, which is uniquely determined by the observational conditional distributions. Thus the final identifiability result follows from the assumed observational distribution plus structural assumptions, not from the conclusion. No load-bearing self-citation chain appears: the uniqueness lemmas used are external (Santambrogio; Jaini et al.) and the cited prior counterfactual-identifiability results are presented as special cases rather than as premises. The one exhibited circular step is in the experimental validation: the model is selected by validation CTF RMSE, which requires ground-truth counterfactual labels, and the same metric is then reported as evidence of counterfactual consistency. This is a moderate empirical circularity, but because the central theoretical identifiability result is independent and self-contained, the overall score is 5 rather than higher.

Assumptions & free parameters 0 free parameters · 10 assumptions · 0 invented entities

The theory introduces no fitted free parameters: the observational distribution PV is an input, and the counterfactual transport is uniquely derived from it. The axioms are standard measure-theoretic, optimal transport, and domain assumptions about the SCM class. Exogenous isomorphism and counterfactual transport are mathematical relations, not new physical entities, so no invented entities appear.

assumptions (10)
  • standard math KR transport between any two distributions on Rd with strictly positive densities exists and is a.s. unique (Lemma 4.7, citing Santambrogio 2015).
    Used in Theorem 4.8 and Corollary 5.4 to fix the counterfactual transport from conditional distributions.
  • standard math A TMI mapping pushing one distribution to another is a.s. the KR transport (Lemma 5.1, citing Jaini et al. 2019).
    Converts triangular monotone causal mechanisms into the unique KR transport.
  • standard math Regular conditional distributions exist and are a.s. unique for standard Borel spaces (Theorem A.12).
    Justifies that PV determines the conditional distributions used to define KR transports.
  • standard math ODE flows with Lipschitz triangular velocity fields are TMI (Lemma A.21, restated from Khoa Le et al. 2025).
    Used to justify the TVSM neural architecture.
  • domain assumption All measurable spaces are standard Borel (Section 2 notation).
    Ensures conditional distributions and pushforwards are well behaved.
  • domain assumption SCMs are recursive and solvable (Definition 2.1, Section 2.1).
    Ensures unique solution mapping Γ and submodel solutions.
  • domain assumption Observational distribution PV has strictly positive density (APV).
    Needed for KR transport existence and uniqueness.
  • domain assumption The model is Markovian (AM) with known causal order and vectorization (A≤ and ATM-SCM).
    These are the defining conditions of the model class for Corollary 5.4.
  • domain assumption Bijective causal mechanisms (BSCM) for Section 4 results (Definition 4.1).
    Bijectivity lets inverses of mechanisms be composed into counterfactual transport.
  • domain assumption Exogenous and endogenous variables have matching dimensional indexing with a fixed coordinate order (Section 4).
    The TM and KR definitions rely on the lexicographic order of coordinates.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Exogenous Isomorphism for Counterfactual Identifiability." pith.science (2026). https://pith.science/paper/ZDRFFBBI

@misc{pith2026250502212,
  author       = {Pith},
  title        = {Pith review of: Exogenous Isomorphism for Counterfactual Identifiability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZDRFFBBI}},
  note         = {Machine review of arXiv:2505.02212}
}
abstract

This paper investigates $\sim_{\mathcal{L}_3}$-identifiability, a form of complete counterfactual identifiability within the Pearl Causal Hierarchy (PCH) framework, ensuring that all Structural Causal Models (SCMs) satisfying the given assumptions provide consistent answers to all causal questions. To simplify this problem, we introduce exogenous isomorphism and propose $\sim_{\mathrm{EI}}$-identifiability, reflecting the strength of model identifiability required for $\sim_{\mathcal{L}_3}$-identifiability. We explore sufficient assumptions for achieving $\sim_{\mathrm{EI}}$-identifiability in two special classes of SCMs: Bijective SCMs (BSCMs), based on counterfactual transport, and Triangular Monotonic SCMs (TM-SCMs), which extend $\sim_{\mathcal{L}_2}$-identifiability. Our results unify and generalize existing theories, providing theoretical guarantees for practical applications. Finally, we leverage neural TM-SCMs to address the consistency problem in counterfactual reasoning, with experiments validating both the effectiveness of our method and the correctness of the theory.

Figures

Figures reproduced from arXiv: 2505.02212 by the authors.

Figure 1
Figure 1. Ablation results of neural TM-SCMs on TM-SCM-SYM. Colored curves depict sliding-window predictions, with shaded areas showing 95% CI. (a) DNME for BARBELL; (b) TNME for STAIR; (c) CMSM for FORK; (d) TVSM for BACKDOOR. Most w/o M settings result in higher CTFRMSE, emphasizing the role of AM. CMSM shows slightly lower stability than other methods. Additional results are in Section D.4. Ablation on ER-DIAG-50 and ER-TR… view at source ↗
Figure 2
Figure 2. Overview of all theorems discussed in the main text and appendix, along with their dependency graph. Nodes represent theorems, and edges indicate dependencie, directed from top to bottom. Different colors denote different topics: theorems related to recursive SCMs are marked in blue; theorems related to exogenous isomorphism are marked in green; theorems related to BSCMs are marked in yellow; and theorems related to… view at source ↗
Figure 3
Figure 3. (a) The objects of study in Theorem 4.3, (f (2) i (v, ·))−1 ◦(f (1) i (v, ·)) (green) and (f (2) i (v ′ , ·))−1 ◦(f (1) i (v ′ , ·)) (red), constructed across different BSCMs; (b) The objects of study in Definition 4.4, (f (1) i (v ′ , ·)) ◦ (f (1) i (v, ·))−1 (blue) and (f (2) i (v ′ , ·)) ◦ (f (2) i (v, ·))−1 (yellow), constructed within the same BSCM. Proposition 4.5. If the BSCM M is Markovian, then for almost a… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: , which are obtained from the validation results of experiments conducted under 10 different random seeds. TM-SCM w/o O w/o M w/o T 0.5 1.0 1.5 2.0 2.5 0.0 0.2 0.4 0.6 0.8 0.15 0.20 0.25 0.30 0.35 0.40 0.0 0.2 0.4 0.6 0.8 0.3 0.4 0.5 0.6 0.7 0.8 0.0 0.1 0.2 0.3 0.4 0.5…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 49 canonical work pages

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    Identifiability of path-specific effects

    Avin, C., Shpitser, I., and Pearl, J. Identifiability of path-specific effects. In Proceedings of the 19th International Joint Conference on Artificial Intelligence, IJCAI'05, pp.\ 357–363, San Francisco, CA, USA, 2005. Morgan Kaufmann Publishers Inc

  3. [3]

    D., Ibeling, D., and Icard, T

    Bareinboim, E., Correa, J. D., Ibeling, D., and Icard, T. On Pearl’s Hierarchy and the Foundations of Causal Inference, pp.\ 507–556. Association for Computing Machinery, New York, NY, USA, 1 edition, 2022. ISBN 9781450395861

  4. [4]

    Representation learning: A review and new perspectives

    Bengio, Y., Courville, A., and Vincent, P. Representation learning: A review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell., 35 0 (8): 0 1798–1828, August 2013. ISSN 0162-8828. doi:10.1109/TPAMI.2013.50

  5. [5]

    Bongers, S., Forr'e, P., Peters, J., and Mooij, J. M. Foundations of structural causal models with cycles and latent variables. The Annals of Statistics, 2016

  6. [6]

    Brehmer, J., de Haan, P., Lippe, P., and Cohen, T. S. Weakly supervised causal representation learning. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp.\ 38319--38331. Curran Associates, Inc., 2022

  7. [7]

    K., and Kasiviswanathan, S

    Chao, P., Bl \"o baum, P., Patel, S. K., and Kasiviswanathan, S. Modeling causal mechanisms with diffusion models for interventional and counterfactual queries. Transactions on Machine Learning Research, 2024. ISSN 2835-8856

  8. [8]

    I., Gao, X., Baptista, R., and Krishnan, R

    Chen, A., Shi, R. I., Gao, X., Baptista, R., and Krishnan, R. G. Structured neural networks for density estimation and causal inference. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp.\ 66438--66450. Curran Associates, Inc., 2023

Show all 61 references
  1. [9]

    Chen, R. T. Q. torchdiffeq, 2018

  2. [10]

    Chen, R. T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K. Neural ordinary differential equations. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 31. Curran As...

  3. [11]

    Nested counterfactual identification from arbitrary surrogate experiments

    Correa, J., Lee, S., and Bareinboim, E. Nested counterfactual identification from arbitrary surrogate experiments. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp.\ 6856--6867....

  4. [12]

    Density estimation using real NVP

    Dinh, L., Sohl-Dickstein, J., and Bengio, S. Density estimation using real NVP . In International Conference on Learning Representations, 2017

  5. [13]

    Neural spline flows

    Durkan, C., Bekasov, A., Murray, I., and Papamakarios, G. Neural spline flows. In Wallach, H., Larochelle, H., Beygelzimer, A., d Alch\' e -Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019

  6. [14]

    Interpolating between optimal transport and mmd using sinkhorn divergences

    Feydy, J., S \'e journ \'e , T., Vialard, F.-X., Amari, S.-i., Trouve, A., and Peyr \'e , G. Interpolating between optimal transport and mmd using sinkhorn divergences. In The 22nd International Conference on Artificial Intelligence and Statistics, pp.\ 2681--2690, 2019

  7. [15]

    Review of causal discovery methods based on graphical models

    Glymour, C., Zhang, K., and Spirtes, P. Review of causal discovery methods based on graphical models. Frontiers in Genetics, 10, 2019. ISSN 1664-8021. doi:10.3389/fgene.2019.00524

  8. [16]

    Grathwohl, W., Chen, R. T. Q., Bettencourt, J., and Duvenaud, D. Scalable reversible generative models with free-form continuous dynamics. In International Conference on Learning Representations, 2019

  9. [17]

    u gelgen, J., Stimper, V., Sch\

    Gresele, L., von K\" u gelgen, J., Stimper, V., Sch\" o lkopf, B., and Besserve, M. Independent mechanism analysis, a new concept? In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, ...

  10. [18]

    M., Peters, J., and Sch\" o lkopf, B

    Hoyer, P., Janzing, D., Mooij, J. M., Peters, J., and Sch\" o lkopf, B. Nonlinear causal discovery with additive noise models. In Koller, D., Schuurmans, D., Bengio, Y., and Bottou, L. (eds.), Advances in Neural Information Processing Systems, volume 21. Curran Associates, Inc., 2008

  11. [19]

    Nonlinear independent component analysis for principled disentanglement in unsupervised deep learning

    Hyv\" a rinen, A., Khemakhem, I., and Morioka, H. Nonlinear independent component analysis for principled disentanglement in unsupervised deep learning. Patterns, 4 0 (10): 0 100844, 2023. ISSN 2666-3899. doi:https://doi.org/10.1016/j.patter.2023.100844

  12. [20]

    o lkopf, B., B\

    Immer, A., Schultheiss, C., Vogt, J. E., Sch\" o lkopf, B., B\" u hlmann, P., and Marx, A. On the identifiability and estimation of causal location-scale noise models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of th...

  13. [21]

    A., and Yu, Y

    Jaini, P., Selby, K. A., and Yu, Y. Sum-of-squares polynomial flow. In Chaudhuri, K. and Salakhutdinov, R. (eds.), Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, pp.\ 3009--3018. PMLR, 09--15 Jun 2019

  14. [22]

    Causal normalizing flows: from theory to practice

    Javaloy, A., Sanchez-Martin, P., and Valera, I. Causal normalizing flows: from theory to practice. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp.\ 58833--58864. Curran Assoc...

  15. [23]

    u gelgen, J., Sch\

    Karimi, A.-H., von K\" u gelgen, J., Sch\" o lkopf, B., and Valera, I. Algorithmic recourse under imperfect causal knowledge: a probabilistic approach. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing System...

  16. [24]

    Variational autoencoders and nonlinear ica: A unifying framework

    Khemakhem, I., Kingma, D., Monti, R., and Hyvarinen, A. Variational autoencoders and nonlinear ica: A unifying framework. In Chiappa, S. and Calandra, R. (eds.), Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics, volume 108 of P...

  17. [25]

    Learning structural causal models from ordering: Identifiable flow models, Apr

    Khoa Le, M., Do, K., and Tran, T. Learning structural causal models from ordering: Identifiable flow models, Apr. 2025

  18. [26]

    Identifiability of deep generative models without auxiliary information

    Kivva, B., Rajendran, G., Ravikumar, P., and Aragam, B. Identifiability of deep generative models without auxiliary information. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp....

  19. [27]

    J., Loftus, J., Russell, C., and Silva, R

    Kusner, M. J., Loftus, J., Russell, C., and Silva, R. Counterfactual fairness. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017

  20. [28]

    D., Gonz \'a lez-Sanz, A., Asher, N., Risser, L., and Loubes, J.-M

    Lara, L. D., Gonz \'a lez-Sanz, A., Asher, N., Risser, L., and Loubes, J.-M. Transport-based counterfactual models. Journal of Machine Learning Research, 25 0 (136): 0 1--59, 2024

  21. [29]

    M., Zhang, K., and Sch \" o lkopf, B

    Lu, C., Huang, B., Wang, K., Hern \' a ndez - Lobato, J. M., Zhang, K., and Sch \" o lkopf, B. Sample-efficient reinforcement learning via counterfactual-based data augmentation. CoRR, abs/2012.09092, 2020

  22. [30]

    and Kiciman, E

    Nasr-Esfahany, A. and Kiciman, E. Counterfactual (non-)identifiability of learned structural causal models, 2023

  23. [31]

    Counterfactual identifiability of bijective causal models

    Nasr-Esfahany, A., Alizadeh, M., and Shah, D. Counterfactual identifiability of bijective causal models. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learning, volume 202...

  24. [32]

    Masked autoregressive flow for density estimation

    Papamakarios, G., Pavlakou, T., and Murray, I. Masked autoregressive flow for density estimation. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 30. Curran A...

  25. [33]

    Deep structural causal models for tractable counterfactual inference

    Pawlowski, N., Coelho de Castro, D., and Glocker, B. Deep structural causal models for tractable counterfactual inference. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, pp.\ 857--869. ...

  26. [34]

    Causality: Models, Reasoning and Inference

    Pearl, J. Causality: Models, Reasoning and Inference. Cambridge University Press, USA, 2nd edition, 2009. ISBN 052189560X

  27. [35]

    and Mackenzie, D

    Pearl, J. and Mackenzie, D. The Book of Why: The New Science of Cause and Effect. Basic Books, Inc., USA, 1st edition, 2018. ISBN 046509760X

  28. [36]

    Elements of Causal Inference: Foundations and Learning Algorithms

    Peters, J., Janzing, D., and Schlkopf, B. Elements of Causal Inference: Foundations and Learning Algorithms. The MIT Press, 2017. ISBN 0262037319

  29. [37]

    Causal Models for Dynamical Systems, pp.\ 671–690

    Peters, J., Bauer, S., and Pfister, N. Causal Models for Dynamical Systems, pp.\ 671–690. Association for Computing Machinery, New York, NY, USA, 1 edition, 2022. ISBN 9781450395861

  30. [38]

    Richens, J., Beard, R., and Thompson, D. H. Counterfactual harm. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, pp.\ 36350--36365. Curran Associates, Inc., 2022

  31. [39]

    Zuko: Normalizing Flows in PyTorch , 2024

    Rozet, F., Divo, F., and Schnake, S. Zuko: Normalizing Flows in PyTorch , 2024

  32. [40]

    and Tsaftaris, S

    Sanchez, P. and Tsaftaris, S. A. Diffusion causal models for counterfactual estimation. In Schölkopf, B., Uhler, C., and Zhang, K. (eds.), Proceedings of the First Conference on Causal Learning and Reasoning, volume 177 of Proceedings of Machine Learning Research, pp.\ 647--66...

  33. [41]

    Vaca: Designing variational graph autoencoders for causal queries

    S \'a nchez-Martin, P., Rateike, M., and Valera, I. Vaca: Designing variational graph autoencoders for causal queries. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pp.\ 8159--8168, 2022

  34. [42]

    One-dimensional issues, pp.\ 59--85

    Santambrogio, F. One-dimensional issues, pp.\ 59--85. Springer International Publishing, Cham, 2015. ISBN 978-3-319-20828-2. doi:10.1007/978-3-319-20828-2_2

  35. [43]

    A fixed-point approach for causal generative modeling

    Scetbon, M., Jennings, J., Hilmkil, A., Zhang, C., and Ma, C. A fixed-point approach for causal generative modeling. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference ...

  36. [44]

    R., Kalchbrenner, N., Goyal, A., and Bengio, Y

    Sch \"o lkopf, B., Locatello, F., Bauer, S., Ke, N. R., Kalchbrenner, N., Goyal, A., and Bengio, Y. Toward causal representation learning. Proceedings of the IEEE, 109 0 (5): 0 612--634, 2021

  37. [45]

    O., Hyv\" a rinen, A., and Kerminen, A

    Shimizu, S., Hoyer, P. O., Hyv\" a rinen, A., and Kerminen, A. A linear non-gaussian acyclic model for causal discovery. Journal of Machine Learning Research, 7 0 (72): 0 2003--2030, 2006

  38. [46]

    and Pearl, J

    Shpitser, I. and Pearl, J. Complete identification methods for the causal hierarchy. Journal of Machine Learning Research, 9 0 (64): 0 1941--1979, 2008

  39. [47]

    and Zhang, K

    Spirtes, P. and Zhang, K. Causal discovery and inference: concepts and recent methodological advances. Applied Informatics, 3 0 (1): 0 3, Feb 2016. ISSN 2196-0089. doi:10.1186/s40535-016-0018-x

  40. [48]

    and Pearl, J

    Tian, J. and Pearl, J. Probabilities of causation: Bounds and identification. Annals of Mathematics and Artificial Intelligence, 28 0 (1): 0 287--313, 2000

  41. [49]

    and Rodriguez, M

    Tsirtsis, S. and Rodriguez, M. Finding counterfactually optimal action sequences in continuous state spaces. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, pp.\ 3220--3247. Curr...

  42. [50]

    u gelgen, J., Besserve, M., Wendong, L., Gresele, L., Keki\' c , A., Bareinboim, E., Blei, D., and Sch\

    von K\" u gelgen, J., Besserve, M., Wendong, L., Gresele, L., Keki\' c , A., Bareinboim, E., Blei, D., and Sch\" o lkopf, B. Nonparametric identifiability of causal representations from unknown interventions. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Le...

  43. [51]

    J., Camgoz, N

    Vowels, M. J., Camgoz, N. C., and Bowden, R. D’ya like dags? a survey on structure learning and causal discovery. ACM Comput. Surv., 55 0 (4), November 2022. ISSN 0360-0300. doi:10.1145/3527154

  44. [52]

    and Louppe, G

    Wehenkel, A. and Louppe, G. Unconstrained monotonic neural networks. In Wallach, H., Larochelle, H., Beygelzimer, A., d Alch\' e -Buc, F., Fox, E., and Garnett, R. (eds.), Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019

  45. [53]

    Learning counterfactual outcomes under rank preservation, 2025

    Wu, P., Li, H., Zheng, C., Zeng, Y., Chen, J., Liu, Y., Guo, R., and Zhang, K. Learning counterfactual outcomes under rank preservation, 2025

  46. [54]

    and Bloem-Reddy, B

    Xi, Q. and Bloem-Reddy, B. Indeterminacy in generative models: Characterization and strong identifiability. In Ruiz, F., Dy, J., and van de Meent, J.-W. (eds.), Proceedings of The 26th International Conference on Artificial Intelligence and Statistics, volume 206 of Proceeding...

  47. [55]

    The causal-neural connection: Expressiveness, learnability, and inference

    Xia, K., Lee, K.-Z., Bengio, Y., and Bareinboim, E. The causal-neural connection: Expressiveness, learnability, and inference. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, pp.\...

  48. [56]

    M., Pan, Y., and Bareinboim, E

    Xia, K. M., Pan, Y., and Bareinboim, E. Neural causal models for counterfactual identification and estimation. In The Eleventh International Conference on Learning Representations, 2023

  49. [57]

    S., Veli c kovi \'c , P., and Kersting, K

    Ze c evi \'c , M., Dhami, D. S., Veli c kovi \'c , P., and Kersting, K. Relating graph neural networks to structural causal models, 2021

  50. [58]

    and Bareinboim, E

    Zhang, J. and Bareinboim, E. Fairness in decision-making — the causal explanation formula. Proceedings of the AAAI Conference on Artificial Intelligence, 32 0 (1), Apr. 2018. doi:10.1609/aaai.v32i1.11564

  51. [59]

    and Hyv\" a rinen, A

    Zhang, K. and Hyv\" a rinen, A. On the identifiability of the post-nonlinear causal model. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, UAI '09, pp.\ 647–655, Arlington, Virginia, USA, 2009. AUAI Press. ISBN 9780974903958

  52. [60]

    Causally consistent normalizing flow, Apr

    Zhou, Q., Lu, K., and Xu, M. Causally consistent normalizing flow, Apr. 2025

  53. [61]

    Zhou, Z., Bai, R., Kulinski, S., Kocaoglu, M., and Inouye, D. I. Towards characterizing domain counterfactuals for invertible latent causal models. In The Twelfth International Conference on Learning Representations, 2024

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.