REVIEW 2 major objections 1 minor 20 references
Confounder Detection via Treatment Intent: A New Observational Study Design
T0 review · 2 major / 1 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read A new observational design elicits unobserved confounders by querying experts on why treatment decisions differ for matched pairs.
desk verdict The paper puts forward a matching-plus-expert-query design to surface unobserved confounders from treatment intent, backed by conditions and an ICU proof-of-concept, but the test substitutes NLP for actual human elicitation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The confounder detection via treatment intent design, which matches units on observed covariates and queries the decision-maker on the intent behind differing treatment assignments to surface hidden variables.
What would settle it
In the semi-synthetic ICU environment with known ground-truth confounders, run the matching-plus-query procedure with the text-note proxy and check whether the elicited variables recover the planted confounders; failure to recover them would falsify the central claim.
Extended reading notes
Core claim
Under the stated conditions the procedure elicits unobserved confounders; empirical evidence indicates EHRs collected in ICUs are subject to unobserved confounding, and a proof-of-concept using clinical text notes as proxy for physician knowledge succeeds in a semi-synthetic environment with known ground truth.
Load-bearing premise
The queried human expert can accurately identify and articulate the unobserved variables responsible for differing treatment decisions on the matched pairs, and the principled matching strategy has already controlled for all observed confounders.
Editorial extensions
If this is right
- Electronic health records collected in ICUs are subject to unobserved confounding.
- Clinical text notes can function as a usable proxy for physician knowledge when detecting hidden confounders.
- The design supplies a practical route to identify variables that must be measured before causal claims can be drawn from observational ICU data.
Reading between the lines
- If the elicited confounders can be measured in new data, standard adjustment methods could then be applied to reduce bias in treatment-effect estimates.
- The same matching-and-query pattern could be tried in other expert-decision domains such as loan approval or policy implementation where unobserved factors shape choices.
- Repeated application across many matched pairs might yield a catalog of recurring hidden variables that future data-collection protocols should record.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a new observational study design called 'confounder detection via treatment intent.' It uses a principled matching strategy to select pairs of units and queries a human expert on their treatment decisions to elicit unobserved confounders that explain residual treatment differences. The manuscript provides theoretical conditions under which the procedure can identify such confounders, presents empirical evidence that ICU electronic health records are subject to unobserved confounding, and includes a semi-synthetic proof-of-concept that substitutes NLP on clinical text notes for physician knowledge in a setting with known ground truth.
Significance. If the theoretical conditions are satisfied and expert elicitation reliably recovers the true unobserved factors, the design could provide a practical tool for diagnosing unobserved confounding in observational data, particularly in domains with accessible expert decision-makers such as clinical research. The combination of matching with targeted expert input is a novel contribution to causal inference methodology.
major comments (2)
- [Proof-of-concept experiment] The semi-synthetic proof-of-concept replaces human expert elicitation with an NLP model on clinical notes. This tests a proxy mechanism rather than the human-elicitation procedure that is central to the design and does not address whether experts can accurately name the specific unobserved variables driving treatment differences on matched pairs (as required by the weakest assumption in the stress-test note).
- [Empirical evidence section] The identifiability result depends on the matching having already controlled for all observed confounders at the level of individual pairs, yet the manuscript provides no quantitative balance diagnostics (e.g., standardized differences or pair-level overlap metrics) for the matched samples used in the ICU application.
minor comments (1)
- [Abstract] The abstract summarizes the theoretical conditions and empirical findings but does not state the key assumptions or report any quantitative results, which reduces immediate accessibility.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. We address each major comment below.
read point-by-point responses
-
Referee: [Proof-of-concept experiment] The semi-synthetic proof-of-concept replaces human expert elicitation with an NLP model on clinical notes. This tests a proxy mechanism rather than the human-elicitation procedure that is central to the design and does not address whether experts can accurately name the specific unobserved variables driving treatment differences on matched pairs (as required by the weakest assumption in the stress-test note).
Authors: We agree that the proof-of-concept substitutes an NLP model on clinical notes for direct human expert elicitation. The manuscript describes this explicitly as a proxy for physician knowledge to enable evaluation in a semi-synthetic setting with known ground truth. While this demonstrates the core matching and elicitation logic under controlled conditions, it does not constitute a direct test of human experts naming unobserved variables. A full human-expert validation study lies outside the scope of the present work due to resource and logistical constraints. We will add explicit discussion of this limitation and its implications for the weakest assumption. revision: partial
-
Referee: [Empirical evidence section] The identifiability result depends on the matching having already controlled for all observed confounders at the level of individual pairs, yet the manuscript provides no quantitative balance diagnostics (e.g., standardized differences or pair-level overlap metrics) for the matched samples used in the ICU application.
Authors: The referee correctly notes the absence of quantitative balance diagnostics for the matched pairs. We will include standardized mean differences, pair-level overlap metrics, and related diagnostics in the revised manuscript to verify that observed confounders are balanced at the individual-pair level as required by the identifiability result. revision: yes
Circularity Check
No circularity; derivation relies on external expert elicitation and matching
full rationale
The paper introduces a new observational study design that queries human experts on matched pairs to elicit unobserved confounders, supported by stated theoretical conditions for identifiability. No equations or claims reduce a target quantity to a fitted parameter by construction, nor does the argument rest on self-citations whose content is unverified within the paper. The semi-synthetic proof-of-concept substitutes NLP for the expert but does not claim to derive the causal effect from the same data used to fit any model. The central result is therefore self-contained against external benchmarks and does not exhibit any of the enumerated circularity patterns.
Assumptions & free parameters
assumptions (1)
- domain assumption Conditions under which querying experts on matched pairs elicits unobserved confounders
Cite this review
Pith. "Pith review of Confounder Detection via Treatment Intent: A New Observational Study Design." pith.science (2026). https://pith.science/paper/HD2L6IWD
@misc{pith2026260526413,
author = {Pith},
title = {Pith review of: Confounder Detection via Treatment Intent: A New Observational Study Design},
year = {2026},
howpublished = {\url{https://pith.science/paper/HD2L6IWD}},
note = {Machine review of arXiv:2605.26413}
}
read the original abstract
Understanding the effects of interventions is central to scientific progress, with randomized controlled trials (RCTs) regarded as the gold standard for causal inference in many applied fields. However, RCTs are costly, time-consuming, and often constrained by ethical or practical limitations, motivating the need for causal methods able to draw conclusions from observational data. While such data is collected at ever larger scale, making its use for causal inference is often hindered by the fact that not all variables affecting treatment allocation and the outcome are observed: an issue known as unobserved confounding. In this paper, we introduce a new study design called confounder detection via treatment intent. The idea is to query a human expert who makes treatment decisions, and ask them to compare pairs of units proposed by a principled matching strategy, with the goal of eliciting unobserved variables that explain why treatment decisions differ. We provide a theoretical basis for such a procedure, ascertaining conditions under which such a study design may elicit unobserved confounders. Building on this newly established foundations, we study treatment effects of interventions in the intensive care unit (ICU). First, we show empirical evidence strongly indicating that electronic health records (EHRs) collected in ICUs are subject to unobserved confounding. By using clinical text notes as a proxy for physicians' knowledge and leveraging natural language processing, we provide a proof of concept for our methodology in a semi-synthetic environment with a known ground truth.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Athey and S
S. Athey and S. Wager. Estimating treatment effects with causal forests: An application. Observational studies, 5(2):37–51, 2019
2019
- [2]
-
[3]
Bareinboim.Causal Artificial Intelligence: A Roadmap for Building Causally Intelligent Systems
E. Bareinboim.Causal Artificial Intelligence: A Roadmap for Building Causally Intelligent Systems. Online, 2025. URLhttps://causalai-book.net/. Draft version
2025
-
[4]
Bareinboim, A
E. Bareinboim, A. Forney, and J. Pearl. Bandits with unobserved confounders: A causal approach.Advances in neural information processing systems, 28, 2015
2015
-
[5]
Bareinboim, J
E. Bareinboim, J. D. Correa, D. Ibeling, and T. Icard. On pearl’s hierarchy and the foundations of causal inference. InProbabilistic and Causal Inference: The Works of Judea Pearl, page 507–556. Association for Computing Machinery, New York, NY , USA, 1st edition, 2022
2022
-
[6]
R. C. Bone, R. A. Balk, F. B. Cerra, R. P. Dellinger, A. M. Fein, W. A. Knaus, R. M. Schein, and W. J. Sibbald. Definitions for sepsis and organ failure and guidelines for the use of innovative therapies in sepsis.Chest, 101(6):1644–1655, 1992
1992
-
[7]
M. Casey. Use of electronic health records in us hospitals.New England Journal of Medicine, 368(16):1469–1470, 2013
2013
-
[8]
Chen and C
T. Chen and C. Guestrin. Xgboost: A scalable tree boosting system. InProceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pages 785–794, 2016
2016
Show all 20 references
-
[9]
Esteban, A
A. Esteban, A. Anzueto, I. Alia, F. Gordo, C. Apezteguia, F. Palizas, D. Cide, R. Goldwaser, L. Soto, G. Bugedo, et al. How is mechanical ventilation employed in the intensive care unit? an international utilization review.American journal of respiratory and critical care medi...
2000
-
[10]
M. O. Harhay, J. Wagner, S. J. Ratcliffe, S. J. Hsieh, I. S. Douglas, and M. P. Kerlin. Random- ized controlled trials.Journal of thoracic disease, 11(7):E79, 2019
2019
-
[11]
W. R. Hersh, M. G. Weiner, P. J. Embi, J. R. Logan, T. H. Payne, E. V . Bernstam, H. P. Lehmann, G. Hripcsak, T. H. Hartzog, J. J. Cimino, et al. Caveats for the use of operational electronic health record data in comparative effectiveness research.Medical care, pages S30– S37, 2013
2013
-
[12]
A. E. Johnson, T. J. Pollard, L. Shen, L.-W. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. A. Celi, and R. G. Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3:160035, 2016. doi: 10.1038/sdata.2016.35. URLhttps://www.nature. com/arti...
2016 doi
-
[13]
Karlin and Y
S. Karlin and Y . Rinott. Classes of orderings of measures and related correlation inequalities. i. multivariate totally positive distributions.Journal of Multivariate Analysis, 10(4):467–498, 1980
1980
-
[14]
S. Lee, J. D. Correa, and E. Bareinboim. General identifiability with arbitrary surrogate exper- iments. InUncertainty in artificial intelligence, pages 389–398. PMLR, 2020
2020
-
[15]
Neumann, D
M. Neumann, D. King, I. Beltagy, and W. Ammar. Scispacy: fast and robust models for biomedical natural language processing. InProceedings of the 18th BioNLP workshop and shared task, pages 319–327, 2019
2019
-
[16]
Pearl.Causality: Models, Reasoning, and Inference
J. Pearl.Causality: Models, Reasoning, and Inference. Cambridge University Press, New York, 2000. 2nd edition, 2009
2000
-
[17]
Rodemund, B
N. Rodemund, B. Wernly, C. Jung, et al. The salzburg intensive care database (sicdb): an openly available critical care dataset.Intensive Care Medicine, 49:700–702, 2023. doi: 10.1007/s00134-023-07046-3. URLhttps://link.springer.com/article/10.1007/ s00134-023-07046-3. 11
2023 doi
-
[18]
Rosenman et al
R. Rosenman et al. Combining observational and randomized data to study heterogeneous treatment effects.arXiv preprint arXiv:2010.12791, 2020
2010
-
[19]
K. F. Schulz, D. G. Altman, and D. Moher. Randomised trials, human nature, and research ethics.The Lancet, 381(9878):1030–1031, 2018
2018
-
[20]
P. J. Thoral, J. M. Peppink, R. H. Driessen, et al. Sharing icu patient data responsibly under the society of critical care medicine/european society of intensive care medicine joint data science collaboration: the amsterdam university medical centers database (amsterdamumcdb)...
2021 doi
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.