{"id":"a2532fd6-073d-4275-8756-8bf8be128592","arxiv_id":"2507.20755","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"AI-chosen live calls improved maternal health message listenership, but the claimed downstream health behavior gains rest on a few uncorrected tests and a self-selected survey sample.","lead":"AI-scheduled phone call interventions in a maternal health program improved how often women listened to automated health messages. The same interventions produced mixed evidence of changes in health behaviors like taking supplements after childbirth.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central outcome claim rests on a per-protocol comparison of women who answered intervention calls against controls who were never offered one; unobserved drivers of call pickup can fully explain the reported effects, and the analysis does not test or bound this selection.","rationale":"The reader's weakest assumption correctly identifies the per-protocol selection problem as the load-bearing issue. The study randomizes arms, but the outcome analysis compares a selected subset of the intervention arm (ID′) with a control group (IC) that was never offered a live intervention call. The Whittle index is a balancing score for the allocation decision, not for the beneficiary's decision to answer a call, so unobserved health motivation and availability can jointly drive call pickup and survey responses. This is a genuine internal-validity threat to the paper's central claim that AI-targeted interventions improve health behaviors, independent of the listenership results. The reader also correctly notes the statistical fragility: only two of 23 outcomes reach p<0.05 without multiplicity correction, and the abstract's mention of iron as significant is contradicted by Table 1's p=0.098. The paper does provide real-world RCT infrastructure, blinded interviewers, and a credible listenership result, so the non-finding should not be read as dismissing the engagement contribution. However, the headline behavioral claim requires either an intention-to-treat analysis with pre-specified endpoints or formal bounds that rule out selection-driven effects. I agree with the reader's REJECT verdict and would not change it.","tokens_in":862,"tokens_out":904,"duration_ms":41369,"concrete_test":"Compute Lee (2009) bounds for the three Table 1 outcomes, treating selection into ID′ (answering the intervention call and later the survey) as endogenous relative to the IC survey responders, with randomized arm assignment as the instrument. If the 95% bounds for the calcium-intake difference and birth-weight knowledge include zero, the per-protocol effect is not identified. As a complement, regress ID′ pickup on pre-intervention covariates (age, parity, gestational age, prior listenership trajectories, phone activity) within the intervention arm; any significant predictors beyond those balanced by design invalidate the claim that Whittle-index-matched groups are exchangeable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The decisive step is in Sections 3.4.2 and 4.1: from the randomized intervention list ID, only those who actually answered the live intervention call (ID′) are surveyed, and these are matched to control-arm women IC who never had the opportunity to answer such a call. The Whittle-index match is designed to make the model's selection decision ignorable; it cannot make the decision to pick up a live call ignorable. If women who answer an unscheduled human call differ in health motivation, phone access, literacy, or availability—all plausibly correlated with taking postnatal supplements or recalling birth weight—the comparison is confounded even with perfect Whittle-index balance. This is not addressed by blinding interviewers or by matching on the index. In addition, Table 1 reports iron intake with p=0.098 and calcium with p=0.041 among 23 outcomes, so with any multiple-comparison correction neither endpoint survives; the abstract's phrase 'iron or calcium' overstates the evidence. The listening outcome in Figure 1 is a valid descriptive check of the intervention's engagement effect, but the leap from listenership to behavior change inherits the selection problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a field randomized controlled trial with 34,453 beneficiaries of the mMitra maternal-health voice-call program, assigned to an intervention arm that received weekly AI-scheduled live service calls chosen by a restless-bandit/DFL Whittle-index policy, versus a control arm that received only automated calls. For evaluation, the authors surveyed 23 knowledge and behavior questions in cohorts 1 and 2, comparing intervention-arm women who answered the live intervention call (ID') to control-arm women selected by a simulated Whittle-index list (IC). They report statistically significant improvements in three outcomes (postnatal iron, postnatal calcium, and birth-weight knowledge), alongside a listenership gain. Cohort 3 shows no significant behavior effect, which the authors attribute to higher baseline listenership.","tokens_in":11992,"tokens_out":4674,"duration_ms":53452,"significance":"The question addressed is important: demonstrating that AI-scheduled engagement interventions change health behaviors would extend prior listenership results and matter for mHealth programs. Strengths include the large deployed sample, the attempt to run a real-world trial with blinded interviewers, and the explicit construction of a counterfactual via Whittle-index matching; the listenership plots provide plausible descriptive evidence of engagement effects. However, the causal claims for behavior change rest on a per-protocol comparison with known selection and on uncorrected multiple testing. As presented, the evidence does not support the abstract's conclusions. If the selection and multiple-testing issues could be addressed with pre-specified intention-to-treat analyses and sensitivity bounds, the study would be valuable.","major_comments":[{"comment":"The central comparison is not between randomized arms: the intervention group analyzed is ID', the subset of the randomized intervention list who picked up the live intervention call, while the control group IC consists of women who were never offered such a call but were selected by a simulated Whittle-index run. Whittle-index matching can make the algorithm's assignment decision ignorable, but it cannot make the beneficiary's decision to pick up a live call ignorable. Factors such as health motivation, phone access, availability, or literacy that predict answering an unscheduled call plausibly also predict postnatal supplement use or knowledge of birth weight; if so, the estimated differences are confounded even with perfect balance on the index. The paper does not report a sensitivity analysis, an instrumental-variable analysis, or bounds that exploit the randomized ID list, so the behavior-change effect is not identified from the design as described.","section":"Section 3.4.2 and 4.1"},{"comment":"The paper tests 23 survey outcomes and reports only two p-values below 0.05 (calcium after delivery p=0.0413 and birth-weight knowledge p=0.0080) and one at p=0.0981 (iron after delivery). Under any standard multiple-comparison correction, such as Bonferroni with 23 tests at 0.05/23 approximately 0.00217, none of these outcomes survives. The abstract's phrasing 'iron or calcium supplements' and 'statistically significant improvements' is therefore not supported by the reported analyses. The paper should specify a pre-registered primary outcome and report adjusted p-values or false-discovery-rate controls.","section":"Table 1 and Tables 4-5"},{"comment":"Survey nonresponse is large and differs by arm: only 701 of 4,495 intervention-arm selected women and 850 of 4,495 control-arm selected women completed the survey. The authors acknowledge that survey response is nonrandom, but then re-match respondents on the Whittle index. Matching on a score that predicts the model's assignment cannot correct for differential selection into the survey or into answering the intervention call, because these selection processes can depend on unobserved post-treatment factors. The paper should report intention-to-treat comparisons on the full randomized ID list, or provide explicit bounds under stated assumptions about selection, before the effect estimate can be considered causal.","section":"Section 3.4.2 and Table 3"},{"comment":"The key results exclude Cohort 3, where no significant behavioral difference was found, and the combined all-cohort analysis is relegated to the appendix rather than reported in the main text. The explanation that Cohort 3 had higher baseline listenership is plausible but not tested; if the intervention effect on behavior operates through listenership, the paper should model this explicitly, for example with an interaction or mediation analysis. A pre-specified analysis plan covering all cohorts and the pooling rule is needed before the headline claim can be accepted.","section":"Section 4.3 and Appendix"}],"minor_comments":[{"comment":"The question count is inconsistent: the text says 13 single-choice questions and 8 multi-answer questions, which sums to 21, not 23; please reconcile the counts.","section":"Section 3.4.1"},{"comment":"The y-axis label 'Cumulative Diff. in Listenership (in sec)' and the caption 'Drop in Intervention Group is significantly lower' conflate cumulative difference with weekly drop; please clarify the definition and add error bars or confidence intervals.","section":"Section 4.2.1 and Figure 1"},{"comment":"The wording should distinguish the survey population (IC plus ID') from the matched analysis sample, since the matching described in Section 4.1 uses only survey responders with similar Whittle indices.","section":"Section 3.4.2"},{"comment":"The table reports percentage improvements without sample sizes, baseline rates, or confidence intervals; please add these for each row.","section":"Table 1"},{"comment":"The use of sklearn's train_test_split to stratify beneficiaries should be described with enough detail to establish the randomization sequence and allocation concealment.","section":"Section 3.2"},{"comment":"Please state explicitly when the Whittle index used for matching is computed and whether it is the final DFL index at the time of intervention selection.","section":"Section 4.1"}],"recommendation":"reject","confidential_remarks":"The manuscript addresses an important question and reports a large field deployment, but the per-protocol identification problem and the multiple-testing issue are central to the claimed behavioral results. I do not see how these can be fixed within a minor revision; a redesign with pre-specified intention-to-treat analyses, multiple-comparison control, and sensitivity analyses would be needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is the first from the ARMMAN/Google line of work to claim that AI-scheduled RMAB/DFL interventions improve downstream health behaviors, not just listenership. That is the right question to ask, and the listenership result is the strong part: a large randomized trial (34k women) with a clear engagement gain that reproduces the group's earlier findings. The Whittle-index matching scheme used to construct the control comparison is a sensible adaptation of Boehmer et al. and is worth reading on its own.\n\nThe behavioral claim is not supported by the analysis as it stands. The comparison is per-protocol: only intervention-arm women who actually answered the live call (ID′) are surveyed, then matched to control-arm women who were never offered that call. Matching on the Whittle index handles the model's selection decision, but it cannot handle the woman's decision to pick up the phone. Health motivation, phone access, or simple availability plausibly drive both call pickup and supplement adherence or birth-weight recall, so the reported effects could be selection, not intervention. An intention-to-treat analysis, or at least a sensitivity analysis that bounds the selection effect, would be needed to close that gap.\n\nThe statistical case is also thinner than the abstract implies. Twenty-three outcomes are compared, only two cross p<0.05, and the iron result that feeds the abstract's 'iron or calcium' phrasing is p=0.098. No multiple-comparison correction is reported. The cohort 3 null result is explained away post hoc via baseline listenership differences, which adds to the sense that the headline is being stretched.\n\nTo be clear, I don't think this is a cynical paper. The authors openly discuss the survey non-response problem and try to re-balance; they improved the questionnaire relative to their prior preprint. But the fix does not address the structural selection problem, because the control group never had the option to comply. The listenership results and the evaluation methodology are genuinely useful; the causal behavior-change claim needs more work.\n\nThis is a paper worth sending to a serious referee: the data are real, the problem matters, and the evaluation question is hard. A good referee should demand an ITT analysis or a clear per-protocol framing with explicit caveats, and should push back on the abstract's wording. Not a desk reject, but not a publish-as-is either.","headline":"Strong listenership results and a useful matching method, but the causal behavior-change claim is undercut by a per-protocol comparison and multiple-testing issues.","tokens_in":12605,"tokens_out":3582,"would_cite":false,"duration_ms":40773,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI-selected service calls improved postnatal calcium use and birth-weight knowledge in a large maternal health trial.","keywords":["restless multi-armed bandits","Whittle index","maternal health","mHealth","health behavior change","counterfactual matching","randomized controlled trial","AI for social impact"],"falsifier":"Compare the matched intervention and control groups on a survey question about a health fact that the automated voice messages never covered; if the intervention group scores higher on that placebo item, the Whittle-index matching has not removed selection into call pickup, and the behavioral gains would be suspect.","tokens_in":11596,"feed_emoji":"📞","tokens_out":8262,"duration_ms":92358,"temperature":0.7,"pith_summary":"The paper asks whether AI-targeted service calls, already known to slow listener dropout in an automated maternal-health voice-message program, also change what mothers know and do. To answer it, the authors ran a two-arm randomized trial with 34,453 enrolled beneficiaries across three cohorts, where a restless-bandit (Whittle-index) model chose which intervention-arm women received a live health-worker call each week. Compared with a 'dummy' control group who would have been selected by the same model but were never called, women who actually received the calls answered survey questions better: calcium intake after delivery improved by 28% (p=0.041) and knowing the baby's birth weight improved by 8.97% (p=0.008). A question on continued iron supplementation moved in the same direction (21.7%) but did not reach conventional significance. If the comparison is accepted, it is the first demonstration in this program that AI-scheduled engagement gains translate into measurable health behaviors and knowledge.","feed_headline":"AI phone calls lifted calcium use and birth-weight knowledge","feed_subtitle":"In a 34,000-mother trial, AI-targeted calls improved key health behaviors and knowledge.","key_machinery":"The load-bearing object is the Whittle index, a scalar assigned to each beneficiary by the restless-bandit scheduling model: it represents the priority or expected marginal benefit of intervening on that arm, and the model uses it to decide each week which intervention-arm women receive a live call. The paper uses the same index twice. First, it is the scheduling rule that generates the intervention list in the treatment arm. Second, it is run 'in dummy mode' on the control arm to identify the women who would have been called, and the index value then serves as the matching variable: each intervention woman who answered her call is paired with a control woman of similar Whittle index in the same cohort. This turns the index from a decision rule into a balancing score for counterfactual comparison, an evaluation strategy the paper adopts from earlier work on index-based treatment allocation.","core_discovery":"The central claim is that listenership improvements caused by AI-scheduled live intervention calls carry through to health behavior change. The paper reports statistically significant improvements in two of the main outcomes in cohorts 1 and 2: the share of mothers still taking calcium pills after delivery rose 28% (p=0.0413) and the share correctly reporting the baby's birth weight rose 8.97% (p=0.0080); iron-pill continuation improved 21.74% but at p=0.0981, a positive trend the authors do not present as conclusive. The causal reading rests on a matched counterfactual: rather than comparing all assigned intervention women with controls, the survey is limited to women who actually picked up the intervention call, matched to control-arm women with nearly equal Whittle indices in the same cohort, the index being the same quantity the scheduling algorithm uses to rank who most needs a call. The authors do not claim the same effect in cohort 3, where baseline listenership was already high and no statistically significant differences appeared.","pith_inferences":["The paper's per-protocol comparison implicitly assumes the Whittle index is a sufficient balancing score; a sensitivity analysis with an unobserved-confounder bound would test how large a hidden selection effect would have to be to erase the calcium result.","If the result replicates, one could optimize the scheduling model directly for behavioral endpoints rather than listenership, since the present study treats listenership as the intermediate outcome.","The same matched-index evaluation could be transferred to vaccination reminders, chronic-disease follow-up, and nutrition programs where call pickup is self-selected and dropout is nonrandom.","The weaker iron result (p=0.098) may indicate which behaviors are most responsive to call encouragement; distinguishing knowledge effects from supply-side barriers would require cost or availability data the paper does not include."],"forward_implications":["Postnatal micronutrient adherence, a behavior that health systems struggle to support after delivery, can be improved by phone-based AI-scheduled encouragement.","Knowledge endpoints such as birth-weight recall serve as measurable downstream proxies for engagement gains in mHealth programs.","The Whittle-index-matched dummy-control design can be applied to other sequential intervention trials where who actually receives the treatment is nonrandom.","Targeting may matter most where baseline engagement is low; the null result for the high-listenership cohort suggests diminishing returns when awareness is already good."],"supporting_citations":[{"why":"Foundational paper defining the restless-bandit index policy used here for scheduling and matching.","marker":"[27]"},{"why":"Supplies the decision-focused learning algorithm that produces the Whittle indices used for scheduling and matching.","marker":"[26]"},{"why":"Prior field deployment evidence that RMAB-selected calls improve listenership, the intermediate outcome this paper extends.","marker":"[16]"},{"why":"Previous preliminary study whose inconclusive results motivate the survey and counterfactual matching changes made here.","marker":"[9]"},{"why":"Provides the statistical rationale for comparing index-selected intervention recipients with dummy-index control beneficiaries.","marker":"[6]"},{"why":"Evidence that the automated voice-message service itself improves maternal knowledge and practices, motivating the behavior endpoints.","marker":"[17]"},{"why":"Links regular listenership in the program to improved health literacy, another background support for the chosen endpoints.","marker":"[11]"}],"fun_headline_variants":["AI-scheduled calls boost maternal calcium use and birth-weight recall","Targeted AI calls raise postnatal calcium adherence in 34k trial","AI call targeting improves maternal health behaviors, not just listenership","Beyond listenership: AI calls drive real behavior change in mothers","AI intervention calls increase calcium uptake and birth-weight knowledge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That two mothers with the same Whittle index are interchangeable in every way that affects the survey answers, even though one picked up her intervention call and the other, an identical-index control, was never given the chance.","fun_headline_variants_meta":{"raw":{"variants":["AI-scheduled calls boost maternal calcium use and birth-weight recall","Targeted AI calls raise postnatal calcium adherence in 34k trial","AI call targeting improves maternal health behaviors, not just listenership","Beyond listenership: AI calls drive real behavior change in mothers","AI intervention calls increase calcium uptake and birth-weight knowledge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1283,"prompt_tokens":940,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":258}},"tokens_in":556,"tokens_out":343,"duration_ms":4288,"temperature":1.0,"reasoning_tokens":258,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T13:16:46.318705+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the matched intervention and control groups on a survey question about a health fact that the automated voice messages never covered; if the intervention group scores higher on that placebo item, the Whittle-index matching has not removed selection into call pickup, and the behavioral gains would be suspect.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Foundational paper defining the restless-bandit index policy used here for scheduling and matching."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the decision-focused learning algorithm that produces the Whittle indices used for scheduling and matching."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior field deployment evidence that RMAB-selected calls improve listenership, the intermediate outcome this paper extends."},{"cited_title":"Preliminary Study of the Impact of AI-Based Interventions on Health and Behavioral Outcomes in Maternal Health Programs","cited_arxiv_id":"2407.11973","evidence_quote":"Previous preliminary study whose inconclusive results motivate the survey and counterfactual matching changes made here."},{"cited_title":"Evaluating the Effectiveness of Index-Based Treatment Allocation","cited_arxiv_id":"2402.11771","evidence_quote":"Provides the statistical rationale for comparing index-selected intervention recipients with dummy-index control beneficiaries."},{"cited_title":"Murthy, S","cited_arxiv_id":null,"evidence_quote":"Evidence that the automated voice-message service itself improves maternal knowledge and practices, motivating the behavior endpoints."},{"cited_title":"Hegde and R","cited_arxiv_id":null,"evidence_quote":"Links regular listenership in the program to improved health literacy, another background support for the chosen endpoints."}],"review_version":1}