{"id":"474b4be2-3bf0-4617-bd00-fd1b001c06c4","arxiv_id":"2501.12803","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Instrumental causal forests show that Indonesia's PKH cash transfer has heterogeneous effects on maternal health care use, with supply-side readiness and household poverty shaping who benefits and who does not.","lead":"This study uses machine learning to see whether Indonesia's cash transfer program helps some mothers more than others, depending on village health services and household poverty. It finds that the program's effects on maternal health care vary widely, with supply-side factors and poverty levels shaping who benefits.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The CLATEs rely on an unstated monotonicity assumption; without it, the headline negative 2013 post-natal effects can be an artifact of non-positive IV weights rather than genuine complier harm.","rationale":"The paper has real strengths: a genuinely randomized subdistrict-level assignment, an instrument that is strongly relevant with plausible exclusion, out-of-bag honest estimation, cluster-robust inference, and LATEs that line up with earlier PKH evaluations. The causal-forest machinery is applied carefully, and the continuous-outcome appendix is a useful robustness check. My read of the central claim is that it stands or falls on whether the estimated CLATEs are actually complier effects. The reader's weakest assumption already named IV validity, and within that package the most under-defended component is monotonicity. The paper never states it, and the non-compliance pattern plus post-2007 programme expansion make it a genuine threat. My proposed test is feasible with the paper's own data: because Z is randomized, the conditional ITT is identified without monotonicity, and under monotonicity it must be the compliance score times the CLATE. If the sign patterns disagree, the headline negative post-natal CLATEs cannot be interpreted causally. This does not change the reader's CONDITIONAL verdict; it sharpens the condition that must be met before the strongest claim can be accepted.","tokens_in":16804,"tokens_out":9564,"duration_ms":108453,"concrete_test":"Using the same data and covariate vector, estimate a causal forest for the conditional intention-to-treat effect τ_ITT(x)=E[Y|Z=1,X=x]−E[Y|Z=0,X=x] on the 2013 post-natal outcome. Under monotonicity and exclusion, τ_ITT(x)=P(complier|X=x)·τ_CLATE(x), so the sign pattern of the conditional ITT should match the sign pattern of the CLATE wherever the compliance score is positive. Compare the proportion of mothers with negative τ_ITT with the proportion with negative CLATE, and check the estimated compliance scores in the negative-CLATE regions. If negative conditional ITT is rare or absent, the negative CLATEs are likely an artifact of non-monotonic weights; if the negative pattern persists in the ITT and in the compliance-score-weighted comparison, the headline heterogeneity finding survives this threat.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that PKH has heterogeneous, supply-side-predictable effects on maternal care—is identified by instrumental forests through Eq. (1), the ratio Cov(Y,Z|X)/Cov(D,Z|X). For a binary instrument and binary treatment, this ratio is a conditional local average treatment effect only if the instrument satisfies relevance, independence, exclusion, and monotonicity (no defiers). Section 3.1 states relevance and exclusion but never discusses monotonicity. The design gives reason for concern: assignment is at the 2007 subdistrict level, yet Table 1 shows 10% of control-arm mothers enrolled in 2009 and 14% in 2013, while 51–52% of treated-arm mothers did not enroll. Because the programme expanded after 2007 and transfer values fell, a household assigned to control could plausibly enrol later while a household assigned to treatment could drop out or never enrol, so the assumption D_i(1)≥D_i(0) for every i is not guaranteed. If there are defiers, Eq. (1) is a weighted average of complier effects with possibly negative weights; the negative CLATEs in the 2013 post-natal histogram (Figure 2), which drive the abstract's strongest finding, could then arise from unstable weights rather than from women who were genuinely made less likely to attend check-ups. The paper does not report any diagnostic or sensitivity analysis for this assumption, so the causal interpretation of the heterogeneity results is not yet secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper estimates heterogeneous effects of Indonesia's conditional cash transfer programme (PKH) on four maternal health care utilisation outcomes, using random assignment to PKH subdistricts as an instrument for actual enrolment. The authors apply instrumental causal forests to obtain conditional local average treatment effects (CLATEs) for the 2009 and 2013 waves, and then summarise heterogeneity with best linear predictors (BLP), classification analysis (CLAN), and depth-two policy trees. They report positive average effects on good assisted delivery in both years, and on pre-natal and post-natal visit thresholds in 2009 only. The CLATE histograms show wide dispersion, especially for post-natal visits in 2013, where some estimated effects are negative. Supply-side variables and household poverty indicators are claimed to predict the heterogeneity. The appendix provides continuous-outcome robustness checks.","tokens_in":17076,"tokens_out":6045,"duration_ms":68256,"significance":"If the estimates are valid, the paper makes a useful empirical contribution to the CCT literature by moving beyond average effects and by demonstrating an instrumental-forest workflow for a large-scale randomised policy evaluation. The use of a randomised instrument, out-of-bag prediction, cluster-robust standard errors at the subdistrict level, and three complementary characterisations of heterogeneity are clear strengths. The policy trees and CLAN results also connect the statistical findings to actionable targeting questions. However, the causal interpretation of the CLATEs rests on the instrumental-variables assumptions, and the manuscript is currently silent on one of them (monotonicity). Since the headline finding of negative post-natal effects in 2013 is driven by CLATEs whose interpretation requires that assumption, the central claim is not yet fully secured. The paper does not provide code or data, so computational reproducibility cannot be verified, but the methods are standard and externally established.","major_comments":[{"comment":"The identification statement for the CLATE is incomplete. Equation (1) identifies a conditional local average treatment effect only if the instrument satisfies relevance, independence, exclusion, and monotonicity (no defiers). The paper explicitly states relevance and exclusion but never discusses monotonicity. This is not a minor omission: Table 1 shows substantial two-sided noncompliance, with 10% of control-arm mothers enrolled in 2009 and 14% in 2013, while 51-52% of treated-arm mothers were not enrolled. Given the programme expansion after 2007 and the decline in transfer value, the assumption that D_i(1) >= D_i(0) for every mother is not guaranteed. Without monotonicity, Eq. (1) is a weighted average of per-group effects with possibly negative weights, so the negative CLATEs in Figure 2 for 2013 post-natal visits could be an artefact of unstable IV weights rather than evidence that some complier mothers were made less likely to attend post-natal check-ups. The authors should either defend monotonicity in this setting or provide sensitivity analyses, such as bounds, alternative principal-strata assumptions, or a discussion of the direction of possible bias.","section":"Section 3.1, Eq. (1)"},{"comment":"The treatment of supply-side variables in the nuisance functions is unclear and potentially consequential. The text in the analytical steps says that the propensity score e(x) is estimated without supply-side variables, while Table 2 labels its last column as 'Used in m(x)' and marks all supply-side variables as 'No'. It is therefore not clear which nuisance functions (m, e, g) include which covariates. If supply-side variables are excluded from e(x) and g(x), and if they are correlated with the instrument or with enrolment conditional on covariates, the residualised treatment and instrument may still be confounded, which would bias the CLATE estimates. Because supply-side variables are central to the paper's heterogeneity claims, this issue is load-bearing. The authors should clarify the exact specification of each nuisance function and, ideally, show robustness to including the full covariate vector in all nuisance functions.","section":"Section 3.1 and Table 2"},{"comment":"The BLP and CLAN results are based on many individual significance tests: roughly 30 covariates across four outcomes and two years, with 95% confidence intervals, and no adjustment for multiple testing. Without such adjustment, some of the reported 'significant' predictors of heterogeneity are likely to be false positives. The qualitative agreement across BLP, CLAN, and policy trees is reassuring, but it is not a formal correction. The authors should either report multiple-testing-corrected confidence sets, pre-specify a smaller set of effect modifiers, or explicitly frame the individual coefficient tests as exploratory and highlight only the patterns that replicate across the complementary analyses.","section":"Section 4.2 and Figure 3"}],"minor_comments":[{"comment":"The notation Cov[Y, Z | X_i = x] is informal; each term should be written as a conditional covariance given X_i = x, not a covariance of conditional objects.","section":"Section 3.1, Eq. (1)"},{"comment":"The caption refers to 'ATE point estimates' from the AIPTW estimator, but the estimates are LATEs because the analysis uses an instrument. Please use the correct estimand label.","section":"Figure 2 caption"},{"comment":"The coding of terciles is described inconsistently: Table 2 says 'q1 = highest quantity' while Figure 3 says 'q1 = largest quantity', and Figure B.7 uses q1-q4 with q4 as the reference. Please define the direction of the tercile indicators once and use consistent wording.","section":"Table 2 and Figure 3 notes"},{"comment":"Please clarify whether the supply-side covariates are measured at baseline (2007) or at the follow-up waves (2009/2013). If they are measured after treatment, they could be affected by PKH and would be bad controls rather than pre-treatment effect modifiers.","section":"Section 2.3"},{"comment":"The policy trees are interpreted qualitatively and no uncertainty quantification is attached to the tree structure. Please state explicitly that the trees are descriptive summaries and do not carry confidence statements, or add an appropriate inferential procedure.","section":"Section 4.4"},{"comment":"There are several typographical and grammatical errors, for example 'hetreatment effects', 'facilites', and 'painting of pictures of enrolled mothers being typically of worth socioeconomic status'. A careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for an economics and global health audience, and the empirical strategy is sensible. The main risk is that the causal interpretation of the heterogeneity results depends on an unstated and potentially implausible monotonicity assumption. This is fixable in revision through explicit discussion and sensitivity analysis. I do not see a novelty or attribution problem: the prior work by Kreif et al. (2022) is cited, and the present application to PKH is distinct. If the authors address the monotonicity concern and clarify the nuisance-function specification, I would expect the paper to be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe key thing to know: this is a competent applied use of instrumental causal forests on a well-known CCT experiment, and the qualitative finding of heterogeneity in PKH's maternal-health effects is probably true. But the causal interpretation of the headline CLATEs is not yet secured, because the paper never states or defends the monotonicity assumption required for Eq. (1) to be a complier effect.\n\nWhat's new: they apply Athey et al.'s instrumental forest to Indonesia's PKH, with a randomized instrument for enrollment, and they combine BLP, CLAN, and policy trees. That is a useful demonstration, and the complementary analyses are sensible. The LATE estimates line up with earlier work (Cahyadi et al.), which gives confidence in the implementation.\n\nWhere it gets soft: the monotonicity issue is real. With 10–14% of controls enrolled and 51–52% of treated not enrolled, defiers are not implausible. Under non-monotonicity, the negative 2013 post-natal CLATEs could be weighted averages with negative weights, not evidence that some women were made less likely to attend. The paper states relevance and exclusion but never monotonicity; no sensitivity check is reported. That is a load-bearing gap for the strongest claim.\n\nThere is also a direct contradiction between the abstract and the results: the abstract says mothers in areas with more doctors, nurses, and delivery assistants were more likely to benefit, but the CLAN and BLP sections report that the lowest tercile of supply had the highest effects in 2013. The paper needs to fix that before anything else. And the complete-case analysis is done without missingness diagnostics; multiple testing is not addressed; no replication code/data.\n\nThat said, the central heterogeneity claim is supported by the histograms and several complementary analyses; the soft spots are fixable. This is not a desk-reject. A serious referee should ask for a monotonicity discussion (and preferably a test or bounding exercise), consistency between abstract and body, and a missing-data check. The citation pattern looks fine, including the self-citation to Kreif et al., which is appropriate given the earlier application of causal forests to Indonesia.\n\nWho it's for: applied health economists and anyone working on causal-forest IV applications. It's a useful teaching/application paper once the IV assumptions are handled honestly.\n\nMy recommendation: engage with it; send it to peer review with the expectation of major revision.","headline":"Useful applied IV-forest study of PKH, but the headline negative 2013 post-natal findings rest on an unstated monotonicity assumption; needs a major revision before I'd trust the causal claim.","tokens_in":17611,"tokens_out":2402,"would_cite":false,"duration_ms":24087,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Instrumental causal forests show that Indonesia's PKH cash-transfer effects on maternal care vary by village health supply and household poverty, with the largest swings in 2013 post-natal visits.","keywords":["conditional cash transfers","Indonesia","PKH","maternal health care utilisation","treatment effect heterogeneity","instrumental causal forests","conditional local average treatment effects","randomized experiment"],"falsifier":"Check whether the randomisation predicts outcomes among mothers who never enrolled (never-takers); if it does, the exclusion restriction fails. Also compute the compliance difference P[enrolled|treatment subdistrict, X] − P[enrolled|control subdistrict, X] within finely defined covariate cells; a negative difference in any cell would reveal defiers and invalidate the CLATE interpretation.","tokens_in":16601,"feed_emoji":"🤰","tokens_out":10537,"duration_ms":93570,"temperature":0.7,"pith_summary":"The paper asks whether the effect of Indonesia's flagship conditional cash transfer programme, PKH, on maternal health care is the same for every mother, and answers that it is not. Treating the programme's 2007 random assignment to subdistricts as an instrument for actual enrolment, the authors estimate conditional local average treatment effects with instrumental causal forests for four outcomes: assisted delivery, facility delivery, four or more pre-natal visits, and two or more post-natal visits. They find that supply-side conditions—the number of doctors, nurses, midwives, and traditional birth attendants in the village—and household poverty indicators predict who benefits, and that benefit patterns shift between the 2009 and 2013 follow-up surveys. The most striking heterogeneity is in 2013 post-natal visits, where some complier mothers are predicted to be less likely to attend two check-ups after receiving the transfer. If these estimates are right, average programme effects conceal meaningful winners and losers, and targeting could be improved by accounting for local health-system capacity.","feed_headline":"PKH's maternal-care gains hinge on village health supply","feed_subtitle":"A causal-forest analysis shows average effects hide winners and losers, especially for post-natal visits in 2013.","key_machinery":"The central mechanism is the instrumental causal forest, a tree-based estimator from the generalized random-forest family that recursively partitions the covariate space into leaves where the conditional 2SLS estimand $\\mathrm{Cov}[Y,Z\\mid X]/\\mathrm{Cov}[D,Z\\mid X]$ is approximately constant, using honest sample splitting and out-of-bag prediction. Here $Y$ is the utilisation outcome, $D$ is actual PKH enrolment, and $Z$ is the randomised offer to live in a treatment subdistrict. The forest returns CLATE estimates for every mother; doubly robust scores aggregate them into an overall LATE, and supplementary analyses (best linear predictors, classification analyses, policy trees) turn the forest output into interpretable statements about which covariates drive heterogeneity: the supply of health workers, household amenities, and survey wave.","core_discovery":"On the paper's own terms, the central discovery is that the effect of PKH enrolment on maternal health-care utilisation is a conditional local average treatment effect (CLATE) that varies substantially with observable covariates, not a single average. Using the randomised subdistrict assignment as an instrument, the authors report that for complier mothers PKH raises the probability of good assisted delivery by about 15–16 percentage points in both 2009 and 2013, but has no significant average effect on facility delivery in either year, and increases the probability of meeting the pre-natal and post-natal visit thresholds only in 2009. Beyond averages, the estimated CLATE distributions span negative and positive values for every outcome, and the 2013 post-natal-visit distribution is the widest, ranging from roughly −0.5 to 0.7. The drivers of this heterogeneity include the per-capita supply of health workers, household poverty (lack of electricity, clean water, septic tank), and survey wave; best-linear-predictor and classification analyses point to supply-side readiness as a key moderator, and policy trees choose health-worker supply variables as splitting criteria in three of four 2013 outcomes.","pith_inferences":["A natural extension is to estimate separate CLATEs by the type of birth attendant available, since the paper's supply variables lump doctors, nurses, midwives, and traditional birth attendants together; splitting them could reveal whether the negative 2013 post-natal effects come from substitution toward traditional attendants.","The results imply that cost-effectiveness calculations should use the joint distribution of CLATEs, not the LATE, because targeting the most-affected quartile could multiply health benefits per rupiah transferred; the paper stops short of drawing this implication.","If supply-side readiness is the true mechanism, a replication in a region with uniformly low supply should find near-zero average effects on assisted delivery; that is a concrete test of the paper's interpretation.","The 2013 policy trees' reliance on health-worker supply suggests demand-side cash transfers and supply-side investment are complements, so a combined intervention bundling PKH with a midwife-incentive payment could be evaluated against PKH alone."],"forward_implications":["If the estimates are correct, the average LATE hides a wide distribution: in 2013 some complier mothers are predicted to lower their post-natal-visit attendance, so describing the programme as 'positive on average' would be misleading for those women.","Supply-side readiness is a first-order moderator: expanding PKH in villages with richer health-worker supply would raise average assisted-delivery gains, while in low-supply villages the cash alone is unlikely to convert into better maternal care.","The timing of measurement matters: effects on pre- and post-natal visits present in 2009 disappear by 2013, consistent with the transfer shrinking from 14% to 7% of household consumption, so any evaluation window must be stated along with the estimate.","Policy trees suggest that simple, interpretable rules—such as enrolling only households in areas with above-median health-worker supply—could capture much of the benefit, and these allocation rules are directly testable."],"supporting_citations":[{"why":"Supplies the PKH experimental design, the instrument of randomised subdistrict assignment, and the IV strategy that this paper extends to heterogeneity analysis.","marker":"Cahyadi et al. (2020)"},{"why":"Introduces generalised random forests and the instrumental forest estimator used to compute the conditional local average treatment effects.","marker":"Athey et al. (2019)"},{"why":"Provides the causal-forest application and honest-splitting guidance that underpins the estimation and out-of-bag prediction.","marker":"Athey and Wager (2019)"},{"why":"Gives the policy-tree learning algorithm used to derive optimal treatment allocation rules.","marker":"Athey and Wager (2021)"},{"why":"Provides the best-linear-predictor and classification-analysis tools for inferring which covariates drive heterogeneity.","marker":"Chernozhukov et al. (2018b)"},{"why":"Supplies debiased machine-learning inference for conditional average treatment effects used in the BLP analysis.","marker":"Semenova and Chernozhukov (2021)"},{"why":"First PKH impact evaluation that this paper extends and compares against for heterogeneity patterns.","marker":"Alatas (2011)"},{"why":"Earlier PKH evaluation whose average-effect estimates and risk-stratified analyses motivate the outcome thresholds used here.","marker":"Kusuma et al. (2016)"}],"fun_headline_variants":["Causal forest shows PKH's maternal-care boost depends on local health workers","Indonesia's cash transfer: maternal care gains vary by village health staff","PKH's effect on maternal care hinges on local doctor and midwife supply","Instrumental causal forests reveal PKH winners and losers in maternal care"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The estimates are complier causal effects only if the randomisation offer changes health-care use solely through actual enrolment, no mother does the opposite of her assignment, and the 2007 randomisation remains valid through 2013 despite programme expansion and a halved transfer value; the paper asserts relevance and exclusion but does not test monotonicity or spillovers.","fun_headline_variants_meta":{"raw":{"variants":["Causal forest shows PKH's maternal-care boost depends on local health workers","Indonesia's cash transfer: maternal care gains vary by village health staff","PKH's effect on maternal care hinges on local doctor and midwife supply","Instrumental causal forests reveal PKH winners and losers in maternal care"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2877,"prompt_tokens":972,"completion_tokens":1905,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":1826}},"tokens_in":588,"tokens_out":1905,"duration_ms":13613,"temperature":1.0,"reasoning_tokens":1826,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:46:00.907821+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check whether the randomisation predicts outcomes among mothers who never enrolled (never-takers); if it does, the exclusion restriction fails. Also compute the compliance difference P[enrolled|treatment subdistrict, X] − P[enrolled|control subdistrict, X] within finely defined covariate cells; a negative difference in any cell would reveal defiers and invalidate the CLATE interpretation.","supporting_citations":[{"cited_title":"A., Prima, R","cited_arxiv_id":null,"evidence_quote":"Supplies the PKH experimental design, the instrument of randomised subdistrict assignment, and the IV strategy that this paper extends to heterogeneity analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"First PKH impact evaluation that this paper extends and compares against for heterogeneity patterns."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier PKH evaluation whose average-effect estimates and risk-stratified analyses motivate the outcome thresholds used here."}],"review_version":1}