Pith. sign in

REVIEW 4 major objections 4 minor 66 references

Uncovering Bias Mechanisms in Observational Studies

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper claims that the sign pattern of three covariances—between the size of the bias and the conditional variance of selection, treatment, and outcome—uniquely identifies which of four common mechanisms is driving the bias in an…

desk verdict New covariance fingerprint for bias mechanisms, novel and well-motivated, but a mismatch between the defined estimator and its consistency proof needs fixing before the results are usable. read the letter →

arxiv 2506.01191 v1 pith:7VWAG5EG submitted 2025-06-01 stat.ME stat.ML

classification stat.MEstat.ML
keywords biasmechanismsobservationalstudiesunmeasuredconfoundingselectiontransportabilitycovariancesignalsnuisancefunctionestimatorsrandomizedcontrolledtrials
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Observational studies often disagree with randomized trials, but the usual methods only say that a bias exists, not why. This paper claims that the reason can be read from how the bias varies across patients: when an unmeasured factor drives a bias, the regions where the bias is largest also show more unexplained variability in the variables that factor influences. The authors define three covariance signals—between the absolute bias and the conditional variance of study selection, treatment, and outcome—and prove that their zero/positive sign pattern uniquely characterizes transportability bias, hidden confounding, and type-1 selection bias, while type-2 selection bias yields nonzero signals in general. They also construct consistent estimators from the squared errors of nuisance functions fit on the observational study, and demonstrate the diagnostic on the Women's Health Initiative hormone-therapy controversy, where it separates an immortal-time selection mechanism from residual transportability bias.

What carries the argument

The load-bearing object is the covariance signal $\bar\rho(b_1,T)=\operatorname{Cov}(|b_1(X)|,\operatorname{Var}(T\mid X,\cdot))$ for $T\in\{S,A,Y\}$, estimated as the covariance between the absolute estimated bias and the squared error of a consistent nuisance-function estimator. The signal works because the paper's generative model builds in "contextual independencies": an unmeasured variable $U$ affects a downstream variable strongly for some patient contexts $X=x$ and not at all for others, with conditional probabilities drawn away from 0.5 from $F(p)=\text{Uniform}([0.1,p]\cup[1-p,0.9])$. Where $U$ acts, both the bias and the conditional variance grow; where it does not, both vanish, so each mechanism leaves a characteristic sign pattern in the three covariances, summarized in Table 1 and formalized in Theorems 4.4 and 4.6.

What would settle it

Simulate a confounding-bias setting in which $U$ affects treatment and outcome but with identical effect sizes for every value of $X$, for instance by holding $p^A_{x,u=1}-p^A_{x,u=0}$ fixed across $x$. The theory predicts $\operatorname{Cov}(|b_1(X)|,\operatorname{Var}(A\mid X,S=1))>0$; if the estimated signal stays near zero across many replications, the contextual-independence mechanism at the core of the paper fails.

Watch

Extended reading notes

Core claim

The paper's central claim is that the mechanism behind an observational study's bias can be identified from the alignment between the bias function and the conditional variance of downstream variables. Writing $b_1(X)=g_1(X)-f_1(X)$ for the difference between the RCT and observational outcome models in the treated group, the paper defines covariance signals $\bar\rho(b_1,S)$, $\bar\rho(b_1,A)$, and $\bar\rho(b_1,Y)$ and shows that, under its generative model, three mechanisms are uniquely characterized by the sign pattern: transportability bias gives $(0,0,+)$, confounding gives $(0,+,+)$, and type-1 selection gives $(+,0,+)$, while type-2 selection gives nonzero signals in general. The mechanism is not recovered from the bias function alone but from the covariance hash table, which converts the question "which bias?" into a pattern-matching problem. The paper also provides consistent estimators of these covariances from squared prediction errors of nuisance functions, so the diagnosis is computable from a paired RCT and observational study, and validates the characterization in synthetic experiments and in a Women's Health Initiative case study.

Load-bearing premise

The load-bearing premise is that the unmeasured factor's influence on each downstream variable varies independently across patient subgroups, with conditional probabilities drawn away from 0.5; if in real data the influence is constant across patient types or the probabilities sit near 0.5, the sign patterns in Table 1 are not guaranteed.

Editorial extensions

If this is right

  • A paired RCT and observational study on overlapping populations can be turned into a bias-mechanism diagnostic by fitting two outcome models, subtracting them, and computing three covariances; the sign pattern then names the mechanism.
  • In the no-bias case all three computed signals are zero, so the same machinery doubles as a falsification check that the observational study is internally valid and transportable.
  • Confounding and type-1 selection, both involving an unmeasured factor that affects the outcome, are told apart by which additional variable aligns with the bias: treatment in the case of confounding, selection in the case of type-1 selection bias.
  • The Women's Health Initiative analysis suggests the diagnostic can separate a collider or immortal-time selection mechanism from residual transportability in real data, thereby pointing to which correction is needed.
  • Because the estimator is consistent whenever the nuisance functions are consistent, the diagnostic remains usable with any modern prediction method, not only parametric models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The exact zeros in Table 1 rely on the independent probability draws across covariate cells in Algorithm 1; in real data those zeros are likely to become near-zero values, so a practical implementation should use thresholds or statistical tests rather than exact equality.
  • The paper's mechanism is essentially a heterogeneity assumption about unmeasured confounding: the effect of $U$ must vary across patient contexts. A direct stress test would hold that effect constant across $X$ in a synthetic study and verify that the covariance signals collapse.
  • Because the case study finds two mechanisms acting simultaneously, a natural next step the paper leaves implicit is to decompose the total bias magnitude into contributions from each mechanism rather than merely naming the dominant one.
  • The same alignment idea could in principle diagnose bias in other paired benchmark settings, such as registry-to-trial emulations, whenever a bias function and consistent nuisance estimators are available.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper develops a method to distinguish among four causal bias mechanisms in observational studies when an RCT benchmark is available. It defines the bias function b_a(x) = g_a(x) - f_a(x) between RCT and OS outcome models, proposes a generative model in Algorithm 1 in which an unmeasured covariate U has context-dependent effects on S, A, and Y, and shows in Lemma 4.2 and Theorem 4.4 that under this model the covariance between |b_1(X)| and conditional variances of S, A, and Y forms a sign pattern (Table 1) that separates transportability, confounding, type-1 selection, and type-2 selection biases. Consistent estimators for the covariances are proposed in Definition 4.7 and Theorem 4.8, and the method is validated on synthetic data and on the WHI hormone-replacement example.

Significance. If the main claims held, the paper would provide a practically attractive, falsifiable diagnostic: from a paired RCT/OS dataset, the sign pattern of three covariances would reveal which bias family is active. The strengths are the explicit generative model, the clean algebraic expressions for the bias in Lemma 4.2, the concrete synthetic experiments including combinations of biases and continuous U, and the WHI case study with a positive-control experiment. However, as detailed below, a central estimator/proof mismatch and the reliance of the sign claims on numerically evaluated integrals currently prevent the theoretical guarantees from supporting the empirical claims at the level stated.

major comments (4)
  1. [Section 4.5, Definition 4.7 and Theorem 4.8] The estimator as written is not the one whose consistency is proved. The second term in Definition 4.7 uses (T_j - \hat{\eta}_T(X_i))^2, but in the proof of Theorem 4.8 (Appendix A.1.4, term (2), Eq. (90)) this is replaced by (Y_j - \hat{\eta}_Y(X_j))^2, and the independence claim for the n^2 - n off-diagonal terms is valid only for that replacement. Under the written definition, for i != j the conditional expectation of |\hat{b}_1(X_i)|(T_j - \hat{\eta}_T(X_i))^2 given X_i and X_j is |\hat{b}_1(X_i)|[Var(T|X_j) + (\eta_T(X_j) - \hat{\eta}_T(X_i))^2], which does not factor into E|b_1| E[Var(T|X)]. Consequently Theorem 4.8 does not establish consistency of the defined estimator, and the synthetic and WHI results computed from that estimator are not covered by the stated theoretical claim. The definition should be corrected to the sample-covariance form with (T_j - \hat{\eta}_T(X_j))^2, or a proof must be supplied for the term as written.
  2. [Appendix A.1.2, proof of Theorem 4.4] Positivity of the covariance signals is asserted for all F(p) in F, but the proof defers to Figure 5 ("which is nonnegative for all F(p) in F (see Figure 5)") rather than giving a closed-form argument. Since the sign pattern in Table 1 is the central fingerprint, and the integrals in Eqs. (48), (56), (60), and (61) depend on the shape of F(p), a numerical figure does not establish the "for all F(p)" claim. The theorem should either be restricted to the subset of F for which analytic sign arguments are given, or the analytic argument should be completed.
  3. [Section 4.1 and Table 1] The exact zero covariances in Table 1 (e.g., \bar{\rho}(b_1,A)=0 for transportability) are derived from the independence of the p^T_{x,u} draws across x in Algorithm 1; they are not general causal properties. If the effect of U on T is constant across X, then |b_1(X)| is constant and all three covariance signals vanish, so the "No Bias" row is indistinguishable from a bias of constant strength. The paper states the varying-effect property as a general clinical property in Section 4.2, but it is an assumption of the generative model. The "uniquely characterized" language in Theorem 4.4 should be qualified to the generative family in Algorithm 1 with non-constant U effects, and the main text should state this limitation explicitly.
  4. [Section A.1.3 and Table 1 (Type 2 selection)] The "non-zero in general" row for type 2 selection bias is weaker than the sign pattern for the other rows, and in some of the paper's own specifications the signal is practically zero: Eq. (85) reports \rho(b_1,A)=0.013 for the selection mechanism used in the main synthetic experiments. A signal of this size is unlikely to be reliably detected in finite samples, so the "!= 0" entry in Table 1 may not be actionable. Please report the expected effect sizes under the mechanisms in Eq. (85)-(88) and discuss the minimum detectable covariance magnitude implied by the proposed estimator.
minor comments (4)
  1. [Appendix A.1.2] In the proof of the selection bias case, the text says "We consider Figure 1b" but the relevant graph is Figure 1c; the equation numbers also refer to the selection-bias graph.
  2. [Figures 3 and 8] The figures plot Pearson's R while Definition 4.3 defines covariances; the normalization should be stated in the main text, along with how the zeros in Table 1 translate to correlations.
  3. [Theorem 4.8 proof] The variance argument treats the estimator as a linear combination of sample means of i.i.d. bounded random variables, but the nuisance estimators are fitted on the same data; the proof should state cross-fitting or uniform consistency assumptions under which this step is valid.
  4. [Section 5] The sentence "p-values clipped at p=10^{-5}" should specify whether clipping is applied before or after multiple-testing corrections and how the percentages in Figure 3 are computed from the 200 runs.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: Table 1 is derived analytically from an explicit generative model, and the estimator is a plug-in for the defined covariance; the only notable issue is a non-circular estimator/proof mismatch.

full rationale

The paper's central derivation is self-contained: Theorem 4.4 computes covariance signs from the explicit generative model of Algorithm 1 and the bias expressions in Lemma 4.2, using the displayed integral representations (e.g., Eqs. 48, 56, 60-61). The sign pattern is a mathematical consequence of the stated assumptions, not a parameter fitted to outcomes and then relabeled as a prediction. Theorem 4.8 is likewise a plug-in consistency argument: if the nuisance estimators converge to the true conditional mean functions, the squared-error terms estimate conditional variances, so the covariance estimator targets the quantity in Definition 4.3. The low-uncertainty distribution F(p) is a stated modeling assumption rather than a hidden reuse of the conclusion; the paper also tests robustness with continuous U and combined biases. A real concern, but not a circularity, is that Definition 4.7's displayed estimator uses (T_j - eta_hat_T(X_i))^2 in its second term, whereas the proof of Theorem 4.8 repeatedly treats the term as (T_j - eta_hat_T(X_j))^2, so the consistency proof as written establishes a different estimator than the one defined. This is an internal correctness/typo issue that does not reduce the derivation to its inputs, and therefore it does not raise the circularity score.

Assumptions & free parameters 2 free parameters · 7 assumptions · 0 invented entities

The central claim rests on a purpose-built generative model with a low-uncertainty distribution F(p) and independent draws per covariate cell. These are modeling axioms, not fitted parameters; no constants are fit to the real data, but the axioms do much of the work in producing the fingerprint.

free parameters (2)
  • F(p) bounds 0.1 and 0.9 = not fitted; fixed constants
    Chosen to ensure positivity and low uncertainty; the proofs of covariance positivity depend on these bounds separating high and low probability regimes.
  • p in F(p) = sampled Uniform[0.2,0.5] in synthetic experiments; ranges (0.1,0.5] in theory
    p controls how strongly probabilities concentrate near 0.1/0.9; it is not fitted to real data but is a free shape parameter of the assumed generative model.
assumptions (7)
  • standard math Assumption 2.1: internal validity of RCT (ignorability of selection, ignorability of treatment, positivity).
    Standard RCT assumptions; used to identify CATErct via g_a(X) in Eq. (3).
  • domain assumption Assumption 4.1: exogeneity X⊥U|R, weak transportability Y^a⊥R|X,U, weak ignorability Y^a⊥S,A|X,U,R, positivity.
    Defines the unmeasured-confounder setting and is used in all Lemma 4.2 proofs.
  • ad hoc to paper Generative model Algorithm 1: for each x, p^T_{x,u} drawn independently from F(p) when U-bias is True, otherwise p^T_{x,u=1}=p^T_{x,u=0}.
    This independent-draw per x creates the contextual independence that makes variance track bias; it is the foundation of Section 4.2 and Theorem 4.4.
  • ad hoc to paper F(p) = Uniform([0.1,p] union [1-p,0.9]) for p in (0.1,0.5], the low-uncertainty distribution.
    Ensures conditional probabilities are away from 0.5, which is needed for the monotone relationship between |bias| and conditional variance.
  • domain assumption Assumption 4.5 for type 2 selection: transportability Y^a⊥R|X and ignorability Y^a⊥A|X,R.
    Isolates collider bias in Figure 1d; used in Theorem 4.6.
  • domain assumption Single binary unmeasured covariate U.
    Simplifies theory; authors argue empirically that continuous U gives similar results.
  • standard math Consistent nuisance estimators in Eqs. (15)-(18).
    Needed for Theorem 4.8 consistency of covariance estimator.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Uncovering Bias Mechanisms in Observational Studies." pith.science (2026). https://pith.science/paper/7VWAG5EG

@misc{pith2026250601191,
  author       = {Pith},
  title        = {Pith review of: Uncovering Bias Mechanisms in Observational Studies},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7VWAG5EG}},
  note         = {Machine review of arXiv:2506.01191}
}
read the original abstract

Observational studies are a key resource for causal inference but are often affected by systematic biases. Prior work has focused mainly on detecting these biases, via sensitivity analyses and comparisons with randomized controlled trials, or mitigating them through debiasing techniques. However, there remains a lack of methodology for uncovering the underlying mechanisms driving these biases, e.g., whether due to hidden confounding or selection of participants. In this work, we show that the relationship between bias magnitude and the predictive performance of nuisance function estimators (in the observational study) can help distinguish among common sources of causal bias. We validate our methodology through extensive synthetic experiments and a real-world case study, demonstrating its effectiveness in revealing the mechanisms behind observed biases. Our framework offers a new lens for understanding and characterizing bias in observational studies, with practical implications for improving causal inference.

Figures

Figures reproduced from arXiv: 2506.01191 by the authors.

Figure 1
Figure 1. Graphs of common biases in observational studies. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Evaluation of treatment assignment probability in three different subgroups in the WHI data [59]. As the conditioning set is expanded and the patient context is better specified, uncertainty in the decisions tend to decrease. Conditioning induces low uncertainty. The vignette above reveals another key prop￾erty of clinical decision-making: as more information about the patient is gathered, the uncertainty in downstr… view at source ↗
Figure 3
Figure 3. Pearson’s R (normalized covariance) in syn￾thetic experiments for the bias mechanisms in [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Left: Pearson’s R between bias functions and downstream variables for CHD and stroke in WHI experiments. Right: Average correlation signals under different synthetic settings: 1) type 2 selection bias (with selection probabilities matching WHI setting) and transportabi…
Figure 5
Figure 5. Figure 5: Different values of the correlation between the absolute bias function and the variance of [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Four different specifications of the selection model. See Eq. [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Common equivalent variants of Selection Bias Type 2. [PITH_FULL_IMAGE:figures/full_fig_p029_7.png]
Figure 8
Figure 8. Figure 8: Covariance signals for synthetic experiments with different covariate dimensionalities [PITH_FULL_IMAGE:figures/full_fig_p029_8.png]
Figure 9
Figure 9. Figure 9: Covariance signals for synthetic experiments with different selection mechanisms and [PITH_FULL_IMAGE:figures/full_fig_p030_9.png]
Figure 10
Figure 10. Figure 10: Covariance signals for synthetic settings involving different combinations of biases. [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]
Figure 11
Figure 11. Figure 11: Covariance signals for synthetic setting with continuous [PITH_FULL_IMAGE:figures/full_fig_p032_11.png]
Figure 12
Figure 12. Figure 12: Left: Covariance signals under different synthetic settings: top two have type 2 selection [PITH_FULL_IMAGE:figures/full_fig_p032_12.png]
Figure 13
Figure 13. Figure 13: Probability of giving HRT in different sequences of subgroups. On the [PITH_FULL_IMAGE:figures/full_fig_p033_13.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

66 extracted references · 60 canonical work pages

  1. [1]

    Akbari, E

    S. Akbari, E. Mokhtarian, A. Ghassami, and N. Kiyavash. Recursive causal structure learning in the presence of latent variables and selection bias.Advances in Neural Information Processing Systems, 34:10119–10130, 2021

  2. [2]

    Anandkumar, K

    A. Anandkumar, K. Chaudhuri, D. J. Hsu, S. M. Kakade, L. Song, and T. Zhang. Spectral meth- ods for learning multivariate latent tree structure.Advances in neural information processing systems, 24, 2011

  3. [3]

    Anandkumar, R

    A. Anandkumar, R. Ge, D. J. Hsu, S. M. Kakade, M. Telgarsky, et al. Tensor decompositions for learning latent variable models.J. Mach. Learn. Res., 15(1):2773–2832, 2014

  4. [4]

    Bareinboim and J

    E. Bareinboim and J. Pearl. Causal inference and the data-fusion problem.Proceedings of the National Academy of Sciences, 113(27):7345–7352, 2016

  5. [5]

    Barrett-Connor and D

    E. Barrett-Connor and D. Grady. Hormone replacement therapy, heart disease, and other considerations.Annual review of public health, 19(1):55–72, 1998

  6. [6]

    Boughdiri, C

    A. Boughdiri, C. Berenfeld, J. Josse, and E. Scornet. A unified framework for the transportability of population-level causal measures.arXiv preprint arXiv:2505.13104, 2025

  7. [7]

    Buitinck, G

    L. Buitinck, G. Louppe, M. Blondel, F. Pedregosa, A. Mueller, O. Grisel, V . Niculae, P. Pretten- hofer, A. Gramfort, J. Grobler, R. Layton, J. VanderPlas, A. Joly, B. Holt, and G. Varoquaux. API design for machine learning software: experiences from the scikit-learn project. InECML PKDD Workshop: Languages for Data Mining and Machine Learning, pages 108–...

  8. [8]

    Cadei, I

    R. Cadei, I. Demirel, P. De Bartolomeis, L. Lindorfer, S. Cremer, C. Schmid, and F. Locatello. Causal lifting of neural representations: Zero-shot generalization for causal inferences.arXiv preprint arXiv:2502.06343, 2025

Show all 66 references
  1. [9]

    T. T. Cai, H. Namkoong, and S. Yadlowsky. Diagnosing model performance under distribution shift.arXiv preprint arXiv:2303.02011, 2023

  2. [10]

    Chobtham and A

    K. Chobtham and A. C. Constantinou. Bayesian network structure learning with causal effects in the presence of latent variables. InInternational Conference on Probabilistic Graphical Models, pages 101–112. PMLR, 2020

  3. [11]

    D. Choo, K. Shiragur, and A. Bhattacharyya. Verification and search algorithms for causal dags. Advances in Neural Information Processing Systems, 35:12787–12799, 2022

  4. [12]

    Colnet, I

    B. Colnet, I. Mayer, G. Chen, A. Dieng, R. Li, G. Varoquaux, J.-P. Vert, J. Josse, and S. Yang. Causal inference methods for combining randomized trials and observational studies: a review. Statistical Science, 39(1):165–191, 2024

  5. [13]

    Colombo, M

    D. Colombo, M. H. Maathuis, et al. Order-independent constraint-based causal structure learning.J. Mach. Learn. Res., 15(1):3741–3782, 2014

  6. [14]

    Dagan, N

    N. Dagan, N. Barda, E. Kepten, O. Miron, S. Perchik, M. A. Katz, M. A. Hernán, M. Lipsitch, B. Reis, and R. D. Balicer. Bnt162b2 mrna covid-19 vaccine in a nationwide mass vaccination setting.New England Journal of Medicine, 384(15):1412–1423, 2021

  7. [15]

    I. J. Dahabreh, J. M. Robins, and M. A. Hernán. Benchmarking observational methods by comparing randomized trials and their emulations.Epidemiology, 31(5):614–619, 2020

  8. [16]

    De Bartolomeis, J

    P. De Bartolomeis, J. Abad, K. Donhauser, and F. Yang. Detecting critical treatment effect bias in small subgroups.arXiv preprint arXiv:2404.18905, 2024

  9. [17]

    De Bartolomeis, J

    P. De Bartolomeis, J. A. Martinez, K. Donhauser, and F. Yang. Hidden yet quantifiable: A lower bound for confounding strength using randomized trials. InInternational Conference on Artificial Intelligence and Statistics, pages 1045–1053. PMLR, 2024

  10. [18]

    Demirel, A

    I. Demirel, A. Alaa, A. Philippakis, and D. Sontag. Prediction-powered generalization of causal inferences. InInternational Conference on Machine Learning, 2024. 10

  11. [19]

    Demirel, E

    I. Demirel, E. De Brouwer, Z. M. Hussain, M. Oberst, A. A. Philippakis, and D. Sontag. Bench- marking observational studies with experimental data under right-censoring. InInternational Conference on Artificial Intelligence and Statistics, pages 4285–4293. PMLR, 2024

  12. [20]

    Fawkes, M

    J. Fawkes, M. O’Riordan, A. Vlontzos, O. Corcoll, and C. M. Gilligan-Lee. The hardness of validating observational studies with experimental data.arXiv preprint arXiv:2503.14795, 2025

  13. [21]

    S. P. Forbes and I. J. Dahabreh. Benchmarking observational analyses against randomized trials: a review of studies assessing propensity score methods.Journal of general internal medicine, 35:1396–1404, 2020

  14. [22]

    J. M. Franklin, E. Patorno, R. J. Desai, R. J. Glynn, D. Martin, K. Quinto, A. Pawar, L. G. Bessette, H. Lee, E. M. Garry, et al. Emulating randomized clinical trials with nonrandomized real-world evidence studies: first results from the rct duplicate initiative.Circulation, 1...

  15. [23]

    Glymour, K

    C. Glymour, K. Zhang, and P. Spirtes. Review of causal discovery methods based on graphical models.Frontiers in genetics, 10:524, 2019

  16. [24]

    Optimizing the use of real world evidence to inform regulatory decision- making, 2019

    Government of Canada. Optimizing the use of real world evidence to inform regulatory decision- making, 2019. URL https://www.canada.ca/en/health-canada/services/drugs-h ealth-products/drug-products/announcements/optimizing-real-world-evide nce-regulatory-decisions.html

  17. [25]

    W. Guo, S. L. Wang, P. Ding, Y . Wang, and M. Jordan. Multi-source causal inference using control variates under outcome selection bias.Transactions on Machine Learning Research,

  18. [26]

    Hartman, R

    E. Hartman, R. Grieve, R. Ramsahai, and J. S. Sekhon. From sample average treatment effect to population average treatment effect on the treated: combining experimental with observational studies to estimate population treatment effects.Journal of the Royal Statistical Society...

  19. [27]

    T. Hatt, J. Berrevoets, A. Curth, S. Feuerriegel, and M. van der Schaar. Combining obser- vational and randomized data for estimating heterogeneous treatment effects.arXiv preprint arXiv:2202.12891, 2022

  20. [28]

    Heinze-Deml, M

    C. Heinze-Deml, M. H. Maathuis, and N. Meinshausen. Causal structure learning.Annual Review of Statistics and Its Application, 5(1):371–391, 2018

  21. [29]

    M. A. Hernán and J. M. Robins. Using big data to emulate a target trial when a randomized trial is not available.American journal of epidemiology, 183(8):758–764, 2016

  22. [30]

    M. A. Hernán, S. Hernández-Díaz, and J. M. Robins. A structural approach to selection bias. Epidemiology, 15(5):615–625, 2004

  23. [31]

    M. A. Hernán, A. Alonso, R. Logan, F. Grodstein, K. B. Michels, W. C. Willett, J. E. Manson, and J. M. Robins. Observational studies analyzed like randomized experiments: an application to postmenopausal hormone therapy and coronary heart disease.Epidemiology, 19(6):766–779, 2008

  24. [32]

    M. A. Hernán, W. Wang, and D. E. Leaf. Target trial emulation: a framework for causal inference from observational data.Jama, 328(24):2446–2447, 2022

  25. [33]

    M. J. Holmberg and L. W. Andersen. Collider bias.Jama, 327(13):1282–1283, 2022

  26. [34]

    Hussain, M.-C

    Z. Hussain, M.-C. Shih, M. Oberst, I. Demirel, and D. Sontag. Falsification of internal and external validity in observational studies via conditional moment restrictions. InInternational Conference on Artificial Intelligence and Statistics, pages 5869–5898. PMLR, 2023

  27. [35]

    Z. M. Hussain, M. Oberst, M.-C. Shih, and D. Sontag. Falsification before extrapolation in causal effect estimation.Advances in Neural Information Processing Systems, 35:6161–6174, 2022. 11

  28. [36]

    G. W. Imbens and D. B. Rubin.Causal inference in statistics, social, and biomedical sciences. Cambridge University Press, 2015

  29. [37]

    Y . Jin, K. Guo, and D. Rothenhäusler. Diagnosing the role of observable distribution shift in scientific replications.arXiv preprint arXiv:2309.01056, 2023

  30. [38]

    Kallus, A

    N. Kallus, A. M. Puli, and U. Shalit. Removing hidden confounding by experimental grounding. Advances in neural information processing systems, 31, 2018

  31. [39]

    Karlsson and J

    R. Karlsson and J. Krijthe. Detecting hidden confounding in observational data using multiple environments.Advances in Neural Information Processing Systems, 36, 2024

  32. [40]

    R. K. Karlsson and J. H. Krijthe. Falsification of unconfoundedness by testing independence of causal mechanisms.arXiv preprint arXiv:2502.06231, 2025

  33. [41]

    Kaul and G

    S. Kaul and G. Gordon. Meta-analysis with untrusted data. InProceedings of the 4th Machine Learning for Health Symposium, volume 259 ofProceedings of Machine Learning Research, pages 563–593. PMLR, 15–16 Dec 2025

  34. [42]

    A. M. Lipsky and S. Greenland. Causal directed acyclic graphs.JAMA, 327(11):1083–1084, 2022

  35. [43]

    S. Lodi, A. Phillips, J. Lundgren, R. Logan, S. Sharma, S. R. Cole, A. Babiker, M. Law, H. Chu, D. Byrne, et al. Effect estimates in randomized trials and observational studies: comparing apples with apples.American journal of epidemiology, 188(8):1569–1577, 2019

  36. [44]

    J. E. Manson, J. Hsia, K. C. Johnson, J. E. Rossouw, A. R. Assaf, N. L. Lasser, M. Trevisan, H. R. Black, S. R. Heckbert, R. Detrano, et al. Estrogen plus progestin and the risk of coronary heart disease.New England Journal of Medicine, 349(6):523–534, 2003

  37. [45]

    I. Ng, X. Dong, H. Dai, B. Huang, P. Spirtes, and K. Zhang. Score-based causal discovery of latent variable causal models. InInternational Conference on Machine Learning, 2022

  38. [46]

    Nice real-world evidence framework, 2022

    NICE. Nice real-world evidence framework, 2022. URL https://www.nice.org.uk/corp orate/ecd9/chapter/overview

  39. [47]

    Oberst, A

    M. Oberst, A. D’Amour, M. Chen, Y . Wang, D. Sontag, and S. Yadlowsky. Understanding the risks and rewards of combining unbiased and possibly biased estimators, with applications to causal inference.arXiv preprint arXiv:2205.10467, 2023

  40. [48]

    Pearl and D

    J. Pearl and D. Mackenzie.The Book of Why: The New Science of Cause and Effect. Basic books, 2018

  41. [49]

    R. L. Prentice, R. Langer, M. L. Stefanick, B. V . Howard, M. Pettinger, G. Anderson, D. Barad, J. D. Curb, J. Kotchen, L. Kuller, et al. Combined postmenopausal hormone therapy and cardiovascular disease: toward resolving the discrepancy between observational studies and the ...

  42. [50]

    A. Ring, N. M. L. Battisti, M. W. Reed, E. Herbert, J. L. Morgan, M. Bradburn, S. J. Walters, K. A. Collins, S. E. Ward, G. R. Holmes, et al. Bridging the age gap: observational cohort study of effects of chemotherapy and trastuzumab on recurrence, survival and quality of life...

  43. [51]

    E. T. Rosenman. Methods for combining observational and experimental causal estimates: A review.Wiley Interdisciplinary Reviews: Computational Statistics, 17(2):e70027, 2025

  44. [52]

    Ruffini, M

    M. Ruffini, M. Casanellas, and R. Gavaldà. A new method of moments for latent variable models.Machine Learning, 107(8):1431–1455, 2018

  45. [53]

    Sadeghi and T

    K. Sadeghi and T. Soo. Conditions and assumptions for constraint-based causal structure learning.Journal of Machine Learning Research, 23(109):1–34, 2022. 12

  46. [54]

    Schuler, D

    A. Schuler, D. Walsh, D. Hall, J. Walsh, C. Fisher, C. P. for Alzheimer’s Disease, A. D. N. Initiative, and A. D. C. Study. Increasing the efficiency of randomized trial estimates via linear adjustment for a prognostic score.The International Journal of Biostatistics, 18(2):32...

  47. [55]

    Shanmugam, M

    K. Shanmugam, M. Kocaoglu, A. G. Dimakis, and S. Vishwanath. Learning causal graphs with small interventions.Advances in Neural Information Processing Systems, 28, 2015

  48. [56]

    Shimizu, P

    S. Shimizu, P. O. Hoyer, A. Hyvärinen, A. Kerminen, and M. Jordan. A linear non-gaussian acyclic model for causal discovery.Journal of Machine Learning Research, 7(10), 2006

  49. [57]

    Squires and C

    C. Squires and C. Uhler. Causal structure learning: A combinatorial perspective.Foundations of Computational Mathematics, 23(5):1781–1815, 2023

  50. [58]

    A. A. Tsiatis.Semiparametric theory and missing data, volume 4. Springer, 2006

  51. [59]

    Design of the women’s health initiative clinical trial and observational study.Controlled clinical trials, 19(1):61–109, 1998

    TWHI. Design of the women’s health initiative clinical trial and observational study.Controlled clinical trials, 19(1):61–109, 1998

  52. [60]

    M. J. V owels, N. C. Camgoz, and R. Bowden. D’ya like dags? a survey on structure learning and causal discovery.ACM Computing Surveys, 55(4):1–36, 2022

  53. [61]

    S. V . Wang, S. Schneeweiss, J. M. Franklin, R. J. Desai, W. Feldman, E. M. Garry, R. J. Glynn, K. J. Lin, J. Paik, E. Patorno, et al. Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials.Jama, 329(16):1376–1385, 2023

  54. [62]

    Y . Wang, M. Schröder, D. Frauen, J. Schweisthal, K. Hess, and S. Feuerriegel. Constructing confidence intervals for average treatment effects from multiple datasets. InInternational Conference on Learning Representations, 2025

  55. [63]

    Y . Xiao, H. Li, Y . Tang, and W. Zhang. Addressing hidden confounding with heterogeneous observational datasets for recommendation. InAdvances in Neural Information Processing Systems, 2024. 13 A Appendix A.1 Proofs Detailed proofs for the theoretical results in the main pape...

  56. [64]

    We have, b1(X) = (p Y u=1 −p Y u=0)(pA u=1 −p A u=0)/2(pA u=1 +p A u=0).(13)

    Confounding Bias– Consider Figure 1b where S⊥ ⊥A, U|X, R= 0 and assume that P(U= 1|R= 1) =P(U= 1|R= 0) = 1/2in the RCT and OS. We have, b1(X) = (p Y u=1 −p Y u=0)(pA u=1 −p A u=0)/2(pA u=1 +p A u=0).(13)

  57. [65]

    1 n nX i=1 |bb1(Xi)| ·(Yi −bηY (Xi))2 # | {z } (1) − EX,Y

    Selection Bias, Type 1– Consider Figure 1c where A⊥ ⊥S, U|X, R= 0and assume that P(U= 1|R= 1) =P(U= 1|R= 0) = 1/2in the RCT and OS. We have, b1(X) = (p Y u=1 −p Y u=0)(pS u=1 −p S u=0)/2(pS u=1 +p S u=0).(14) Proof of the Transportability Bias. We consider Figure 1a where S⊥ ⊥...

  58. [66]

    sequences

    Therefore: P(| ˆθn −θ| ≥ε)≤P |ˆθn −E[ ˆθn]| ≥ε 2 ≤ 4·Var( ˆθn) ε2 Taking the limit asn→ ∞: lim n→∞ P(| ˆθn −θ| ≥ε)≤lim n→∞ 4·Var( ˆθn) ε2 = 0 which proves ˆθn is consistent and we are done. A.2 Additional DAGs for Selection Bias Type 2 See Figure 7 for additional DAGs reflecti...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.