REVIEW 3 major objections 3 minor 18 references
Identifiability in Causal Abstractions: A Hierarchy of Criteria
T0 review · 3 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read To decide whether a causal query is identifiable from a collection of possible diagrams, it is enough to check the maximal diagrams.
desk verdict Useful conceptual hierarchy for identifiability under causal abstraction, but the main maximal-graph reduction is only proven modulo a repairable gap about bidirected edges. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the identification object—a uniform name for a causal estimand, a do-calculus proof, or a graphical criterion—together with the subgraph-inclusion monotonicity of Lemma 1: because edges represent dependencies, removing an edge preserves d-separation, so anything that identifies a query in a super-graph also works in every sub-graph. Theorem 2 uses this monotonicity to shrink a collection of diagrams to its inclusion-maximal elements. The hierarchy Theorem 3 then relates the notions by logical implication, and its proof uses the completeness of do-calculus (every single-graph identification has a do-calculus proof) to equate common graphical criteria with common do-calculus.
What would settle it
Exhibit a causal diagram $G_1$, a subgraph $G_2$, and an SCM $M_2$ inducing $G_2$ for which no SCM $M_1$ inducing $G_1$ has the same observational and interventional distributions; Lemma 1 would then fail. Concretely, try adding a bidirected edge in $G_1$ between two variables that share no common exogenous parent in $M_2$, and check whether the required latent variable can be added without altering $P(y|\mathrm{do}(x))$.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a structural simplification plus a hierarchy. Given a class of causal diagrams, define an 'identification object' as a valid estimand, a do-calculus proof, or a graphical criterion; if such an object applies to every maximal graph of the class, it applies to every graph in the class (Lemma 1, Theorem 2). Consequently, for the identifiability notions IG, ICD, ICGC, ICB, and ICF, a query is identifiable in the collection exactly when it is identifiable in the subcollection of maximal elements. The hierarchy (Theorem 3) further shows that identifiability by a common specific criterion (e.g., common backdoor or frontdoor) implies identifiability by common graphical criterion, which is equivalent to identifiability by common do-calculus; that in turn implies identifiability through graphs, which implies identifiability knowing the true observed density. The missing arrow—whether identifiability through graphs forces a single common do-calculus proof—is left as an explicit conjecture.
Load-bearing premise
The reduction to maximal graphs rests on the assumption that any causal diagram can be enlarged to a given super-diagram by adding edges—including new latent confounding shown as bidirected edges—without changing the observational or interventional distributions, just by making the new structural functions ignore the added arguments.
Editorial extensions
If this is right
- If a collection of diagrams has a single greatest element under inclusion, identifiability through graphs and identifiability by common do-calculus coincide (Corollary 1), so checking one graph settles the whole class.
- Designing causal abstractions with few maximal elements—ideally one—makes identifiability verification computationally tractable, since only the maximal diagrams need to be inspected.
- Because identifiability by common graphical criterion is equivalent to identifiability by common do-calculus, no criterion known to be incomplete on single graphs (backdoor, frontdoor) can be complete for common do-calculus in general classes.
- If the conjecture holds, some causal queries are identifiable in the per-graph sense but not provable by any single uniform do-calculus derivation; if it fails, the two notions collapse and any atomically complete calculus is fully complete.
Reading between the lines
- A practical reading of Theorem 2 is algorithmic: compute the inclusion-maximal elements of an abstraction (or an over-approximation of them) and run any single-graph identification routine there; if the resulting estimand is the same across maximal elements, it is valid for the entire class.
- The open IG-versus-ICD gap is likely testable on small graphs: one can search for two ADMGs with the same marginal query formula but no shared do-calculus proof, following the paper's proof strategy of exhibiting two proofs that yield the same formula.
- The equivalence ICGC ⇔ ICD suggests a design target for abstraction languages: a complete adjustment criterion for a class automatically yields uniform do-calculus proofs, so developing such criteria is exactly as hard as common do-calculus.
- For neighboring problems, the same maximal-element reduction may apply to other query types (e.g., conditional effects or counterfactuals), whenever the query is monotone under edge removal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes identifiability of causal queries over collections of causal diagrams, called causal abstractions. It introduces several notions of identifiability in such collections—Identifiability through Graphs (IG), Identifiability through Graphs knowing P* (IGP), Identifiability by Common Do-Calculus (ICD), Identifiability by Common Graphical Criterion (ICGC), and common specific criteria such as backdoor and frontdoor (ICB, ICF)—and studies the logical relations among them. The main technical claim is Theorem 2, which asserts that for these notions it suffices to consider only the maximal elements of the collection under graph inclusion. The paper also states a hierarchy of implications among the notions and leaves open a conjecture separating IG from ICD.
Significance. If established, the maximal-element reduction of Theorem 2 would be practically useful: it would allow identifiability questions over very large or infinite collections of graphs to be reduced to a usually smaller subcollection. The taxonomy of identifiability notions is also a useful conceptual contribution, and Example 2 correctly illustrates the difference between IG and IGP. The paper is honest about the open IG-versus-ICD question and gives concrete strategies for attacking it. However, the central reduction proof has a genuine gap for graphs with bidirected edges, and the equivalence ICGC iff ICD is not rigorously defined. These issues are load-bearing for the paper's main claims, so the manuscript needs substantial revision.
major comments (3)
- [§5.1, Lemma 1(1)] The construction of M1 from M2 by adding ignored arguments to the structural functions cannot add bidirected edges. A bidirected edge Vi<->Vj requires a newly shared exogenous parent, but the construction keeps the exogenous variable set U unchanged. Consequently, the equalities PM1(V)=PM2(V) and PM1(y|do(x),z)=PM2(y|do(x),z) are not justified. For example, if G2 has no edge between X and Y and G1 adds X<->Y, then any SCM inducing G1 must have X and Y share an exogenous U that continues to influence Y under do(X=x); this interventional law cannot be reproduced by the proposed construction from an arbitrary SCM inducing G2. The directed-edge part of the construction is fine, and Lemma 1(2) is correct, but the stated claim that any causal estimand valid in G1 is valid in G2 is not supported for ADMGs with bidirected edges.
- [§5.1, Theorem 2] Theorem 2 relies on Lemma 1 to carry an identification object from each maximal graph to every subgraph. For ICD and ICGC, Lemma 1(2) and (3) do provide that transfer because d-separations are preserved under edge removal. For IG, however, the identification object is an arbitrary causal estimand, and its transfer depends exactly on the unproven Lemma 1(1). A possible repair would use the completeness of do-calculus: if the query is identifiable in a maximal graph, there is a do-calculus proof valid in that graph, and the same proof is valid in every subgraph. But the paper does not supply this argument, and such a repair would establish ICD, not the literal 'any estimand' statement of Lemma 1. As written, the equivalence between identifiability in C and in Cmax for IG is not established.
- [§4.2, Definition 5 and §5.2, Theorem 3] The equivalence ICGC iff ICD cannot be evaluated as stated because Definition 5 does not define what a 'graphical criterion' is, what it means for a criterion to be satisfied by a graph, or whether the criterion is required to be sound. The proof of ICGC⇒ICD asserts that any graphical criterion on a single graph can be established by a do-calculus proof; without a formal definition of the allowed criteria, this is not a theorem. Conversely, the proof of ICD⇒ICGC defines the common graphical criterion as the conjunction of the d-separation conditions appearing in the common do-calculus proof, which makes the direction true by construction but also makes ICGC essentially a restatement of ICD. The hierarchy in Figure 1 is therefore only meaningful once 'graphical criterion' is formalized.
minor comments (3)
- [§7, Conclusion] The conclusion asks whether there exists a class of diagrams in which a query is 'ICD but not IG'; this direction is already ruled out by Theorem 3, and it reverses the conjecture stated in Section 6, which asks whether a query can be IG but not ICD.
- [§4.2, Example 1 and Figure 1] The graph drawings in Example 1 are very hard to read in the arXiv rendering, and Figure 1 does not label the arrows or the '?' in the text, making it difficult to verify the claimed hierarchy against Theorem 3.
- [§5.2, Theorem 3] In the proof of IG⇒IGP, the SCM class in Definition 3 is described as 'strictly smaller'; it is only smaller or equal, and the word 'strictly' is unnecessary and potentially misleading.
Circularity Check
No circularity; the maximal-subcollection reduction is a direct implication of a (gappy but independent) monotonicity lemma, and the hierarchy is definitional rather than self-referential.
full rationale
Walking the derivation chain, Theorem 2 is the central reduction and it is proved directly from Lemma 1 plus the definition of maximal elements: an identification object (estimand, do-calculus proof, or graphical criterion) valid on every maximal graph is carried to each subgraph by Lemma 1. Lemma 1 is an independent monotonicity statement, not an assumption of the target identifiability conclusion, and it is not derived from Jaber et al. or from any fitted data. Theorem 3's implications are consequences of the paper's own definitions of IG, ICD, ICGC, ICB, and ICF; the ICGC/ICD equivalence is close to definitional in the sense that a 'graphical criterion' is any conjunction of d-separation conditions, but the manuscript does not disguise this as an empirical prediction. The paper explicitly leaves the IG-versus-ICD implication open as a conjecture, so it does not claim a result it lacks. The self-citations (Assaad et al. 2024; Yvernes et al. 2025) appear only in the related-work discussion of common backdoor and adjustment criteria and are not used to validate Theorem 2 or Theorem 3. The one genuine issue is in the proof of Lemma 1 point 1: after constructing M1 by adding ignored directed-parent arguments, the proof asserts M1 induces G1, but a bidirected edge in G1 that is absent in G2 requires a newly shared exogenous parent, and the construction introduces none. As written, the equalities P_M1(V)=P_M2(V) and P_M1(y|do(x),z)=P_M2(y|do(x),z) are not established for ADMGs with such bidirected edges. This is a correctness gap (potentially repairable via do-calculus completeness), not a circular reduction; no estimand, proof, or criterion is fitted or assumed equal to its own target. Hence no circularity.
Assumptions & free parameters
assumptions (3)
- standard math do-calculus is sound and complete for causal diagrams (Shpitser and Pearl 2006), so any identifiability result is expressible as a do-calculus proof (Section 3).
- domain assumption A causal abstraction always induces a collection C of causal diagrams over the same variable set V (Section 4).
- domain assumption An SCM M2 inducing G2 can be extended to an SCM M1 inducing a supergraph G1, with the same observational and interventional distributions, by adding ignored arguments to structural functions (Lemma 1).
Cite this review
Pith. "Pith review of Identifiability in Causal Abstractions: A Hierarchy of Criteria." pith.science (2026). https://pith.science/paper/KBFMP6JM
@misc{pith2026250706213,
author = {Pith},
title = {Pith review of: Identifiability in Causal Abstractions: A Hierarchy of Criteria},
year = {2026},
howpublished = {\url{https://pith.science/paper/KBFMP6JM}},
note = {Machine review of arXiv:2507.06213}
}
read the original abstract
Identifying the effect of a treatment from observational data typically requires assuming a fully specified causal diagram. However, such diagrams are rarely known in practice, especially in complex or high-dimensional settings. To overcome this limitation, recent works have explored the use of causal abstractions-simplified representations that retain partial causal information. In this paper, we consider causal abstractions formalized as collections of causal diagrams, and focus on the identifiability of causal queries within such collections. We introduce and formalize several identifiability criteria under this setting. Our main contribution is to organize these criteria into a structured hierarchy, highlighting their relationships. This hierarchical view enables a clearer understanding of what can be identified under varying levels of causal knowledge. We illustrate our framework through examples from the literature and provide tools to reason about identifiability when full causal knowledge is unavailable.
Figures
Reference graph
Works this paper leans on
-
[1]
Tara V. Anand, Adele H. Ribeiro, Jin Tian, and Elias Bareinboim. Causal Effect Identification in Cluster DAGs . Proceedings of the AAAI Conference on Artificial Intelligence, 37 0 (10): 0 12172--12179, June 2023. ISSN 2374-3468. doi:10.1609/aaai.v37i10.26435. URL https://ojs.aaai.org/index.php/AAAI/article/view/26435
-
[2]
Assaad, Emilie Devijver, and Eric Gaussier
Charles K. Assaad, Emilie Devijver, and Eric Gaussier. Discovery of extended summary graphs in time series. In James Cussens and Kun Zhang, editors, Proceedings of the Thirty-Eighth Conference on Uncertainty in Artificial Intelligence, volume 180 of Proceedings of Machine Learning Research, pages 96--106. PMLR, 01--05 Aug 2022. URL https://proceedings.mlr...
work page 2022
-
[3]
Assaad, Emilie Devijver, Eric Gaussier, Gregor G\" o ssler, and Anouar Meynaoui
Charles K. Assaad, Emilie Devijver, Eric Gaussier, Gregor G\" o ssler, and Anouar Meynaoui. Identifiability of total effects from abstractions of time series causal graphs. In Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence. JMLR.org, 2024
work page 2024
-
[4]
Sander Beckers and Joseph Y. Halpern. Abstracting causal models. CoRR, abs/1812.03789, 2018. URL http://arxiv.org/abs/1812.03789
arXiv 2018
-
[5]
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 9574--9586. Curran Associates, Inc., 2021. URL https://proceedings.neurips.cc/paper_files/paper/2...
work page 2021
-
[6]
Causal abstraction: A theoretical foundation for mechanistic interpretability, 2025
Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, and Thomas Icard. Causal abstraction: A theoretical foundation for mechanistic interpretability, 2025. URL https://arxiv.org/abs/2301.04709
arXiv 2025
-
[7]
Causal identification under markov equivalence: Calculus, algorithm, and completeness
Amin Jaber, Adele Ribeiro, Jiji Zhang, and Elias Bareinboim. Causal identification under markov equivalence: Calculus, algorithm, and completeness. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 3679--3690. Curran Associates, Inc., 2022. URL https://proceed...
work page 2022
-
[8]
Learning structures of bayesian networks for variable groups
Pekka Parviainen and Samuel Kaski. Learning structures of bayesian networks for variable groups. International Journal of Approximate Reasoning, 88: 0 110--127, 2017. ISSN 0888-613X. doi:https://doi.org/10.1016/j.ijar.2017.05.006. URL https://www.sciencedirect.com/science/article/pii/S0888613X17303134
Show all 18 references
-
[9]
Causality
Judea Pearl. Causality. Cambridge University Press, Cambridge, 2 edition, 2009. ISBN 978-0-521-89560-6. doi:10.1017/CBO9780511803161. URL https://www.cambridge.org/core/books/causality/B0046844FAE10CBF274D4ACBDAEB5F5B
2009 doi
-
[10]
Maathuis
Emilija Perković, Johannes Textor, Markus Kalisch, and Marloes H. Maathuis. Complete Graphical Characterization and Construction of Adjustment Sets in Markov Equivalence Classes of Ancestral Graphs . Journal of Machine Learning Research, 18: 0 1--62, May 2018. ISSN 1532-4435. ...
2018 doi
-
[11]
Rischel and Sebastian Weichwald
Eigil F. Rischel and Sebastian Weichwald. Compositional abstraction error and a category of causal models. In Cassio de Campos and Marloes H. Maathuis, editors, Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, volume 161 of Proceedings of...
2021
-
[12]
Identification of joint interventional distributions in recursive semi- Markovian causal models
Ilya Shpitser and Judea Pearl. Identification of joint interventional distributions in recursive semi- Markovian causal models. In proceedings of the 21st national conference on Artificial intelligence - Volume 2 , pages 1219--1226, Boston, Massachusetts, July 2006. AAAI Press...
2006
-
[13]
Foundations of causal discovery on groups of variables
Jonas Wahl, Urmi Ninad, and Jakob Runge. Foundations of causal discovery on groups of variables. Journal of Causal Inference, 12 0 (1): 0 20230041, 2024. doi:doi:10.1515/jci-2023-0041. URL https://doi.org/10.1515/jci-2023-0041
2024 doi
-
[14]
Weinstein and David M
Eli N. Weinstein and David M. Blei. Hierarchical causal models, 2024. URL https://arxiv.org/abs/2401.05330
2024 arXiv
-
[15]
Complete characterization for adjustment in summary causal graphs of timeseries
Clément Yvernes, Emilie Devijver, and Eric Gaussier. Complete characterization for adjustment in summary causal graphs of timeseries. In UAI2025, 2025
2025
-
[16]
Towards computing an optimal abstraction for structural causal models, 2022
Fabio Massimo Zennaro, Paolo Turrini, and Theodoros Damoulas. Towards computing an optimal abstraction for structural causal models, 2022. URL https://arxiv.org/abs/2208.00894
2022 arXiv
-
[17]
Quantifying consistency and information loss for causal abstraction learning
Fabio Massimo Zennaro, Paolo Turrini, and Theodoros Damoulas. Quantifying consistency and information loss for causal abstraction learning. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI '23, 2023. ISBN 978-1-956792-03-4. d...
2023 doi
-
[18]
Causally abstracted multi-armed bandits
Fabio Massimo Zennaro, Nicholas Bishop, Joel Dyer, Yorgos Felekis, Anisoara Calinescu, Michael Wooldridge, and Theodoros Damoulas. Causally abstracted multi-armed bandits. In Proceedings of the Fortieth Conference on Uncertainty in Artificial Intelligence, UAI '24. JMLR.org, 2024
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.