{"id":"7b20dfed-f3f0-4cad-9114-6a4ab696b4d6","arxiv_id":"2607.08706","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Patients randomized to access a pre-visit AI chatbot received 4.6 pp fewer prescriptions and 2.7 pp more diagnostic tests, reflecting the chatbot's encoded caution against medications and clean recommendations for testing.","lead":"A large field experiment at a Chinese hospital found that patients given access to an AI chatbot before their doctor visit received fewer prescriptions and more diagnostic tests, mirroring the chatbot's liability-driven caution against medications. This matters because it shows AI design choices—made by developers to limit legal risk—can propagate into real clinical decisions at scale.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The observed pattern (reduced prescribing, increased testing) is consistent with general patient-information effects, not uniquely with AI directionality; the paper's own conversation logs could distinguish these but are not used for that purpose.","rationale":"The reader correctly identifies the weakest assumption: the causal channel between AI advice content and clinical outcomes is inferred from pattern matching rather than directly observed. I agree this is the most load-bearing concern. The paper's design is strong for establishing ITT effects and documenting AI advice patterns, but the specific claim about directionality propagating requires distinguishing it from general information effects that predict the same outcome pattern. This distinction matters because the paper's policy implications—that developer guardrail choices propagate into real decisions at scale—depend on directionality being the operative mechanism, not just one of several consistent explanations. The concern is addressable rather than fatal: the conversation logs already contain the information needed to test whether outcomes vary with advice content. The 17% take-up rate and differential survey response rates are secondary concerns that the paper handles reasonably through IV estimation and appropriate caveats. The unidentified LLM model limits reproducibility but does not undermine the internal validity of the experiment. CONDITIONAL remains the appropriate verdict: the core empirical findings are solid, but the central interpretive claim about directionality specifically would be substantially strengthened by the proposed subgroup analysis using existing data.","tokens_in":39889,"tokens_out":3584,"duration_ms":186685,"concrete_test":"Among the 998 chatbot users, split by whether their conversation logs mention medications (58.9% of turns do, per Table 2) versus those whose conversations did not discuss medications at all. Reestimate the prescribing effect (Table 3, col 1) separately for these two subgroups using the same within-physician specification. If the prescribing reduction is concentrated among users whose conversations included medication cautions, this supports the directionality channel. If prescribing reductions are similar regardless of whether medications were discussed, the pattern likely reflects general preparation rather than directional advice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that AI guardrail *directionality*—cautioning against medications while recommending testing—propagates into clinical decisions. The key evidence is that effect directions mirror advice directions (§4.2). But this pattern is also predicted by mechanisms the paper itself acknowledges (§3.1): better-informed patients constrain overtreatment (reducing prescriptions) and improved symptom articulation leads to more appropriate test ordering (increasing testing). In a Chinese hospital with documented overprescription of TCM and antibiotics, virtually any pre-visit information intervention would likely reduce unnecessary prescriptions. The heterogeneity analyses (Tables 7–9) show effects concentrated among receptive physicians and intensive prescribers, which is consistent with patients bringing specific requests shaped by AI advice but also with general information effects disciplining expert behavior in a credence-goods setting. The cross-model comparison (Figures A1–A2) documents variation in caution levels across models but cannot link that variation to outcome variation within the experiment. Critically, the paper possesses conversation logs for 998 users—recording which topics each user discussed and what stance the AI took—but does not examine whether clinical effects vary with the specific content of those conversations. Without this, the paper establishes that chatbot access changes outcomes and that the changes are *consistent with* directional advice, but cannot confirm that the directionality of the advice is the operative mechanism rather than a general information or preparation effect that happens to produce the same pattern.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper reports a preregistered field experiment (AEARCTR-0015851) conducted at a large Chinese public hospital, randomizing 11,666 outpatient visits to patient access to an LLM-based chatbot prior to consultation. The authors document that the chatbot's advice is systematically directional: it cautions against medications (especially TCM and antibiotics) while issuing clean recommendations for diagnostic testing, a pattern they attribute to liability-driven guardrails in AI training. They then show that chatbot access reduces prescription rates by 4.6 pp (ITT) and increases diagnostic testing by 2.7 pp, with effects concentrated among physicians receptive to patient input and those with higher baseline prescribing intensity. Survey evidence reveals lower patient satisfaction and intended compliance, while physicians report clearer symptom descriptions. The paper contributes to the economics of AI, credence goods, and patient information literatures by examining how design choices encoded in AI systems propagate into real-world expert-client decisions.","tokens_in":40251,"tokens_out":1442,"duration_ms":293603,"significance":"The paper addresses a timely and important question: whether the defensive guardrails encoded in generative AI systems propagate into real-world decisions at scale. The experimental design is strong—preregistered, two-layer randomization with physician fixed effects, balance confirmed (Table 1), and ITT analysis preserving randomization integrity. The IV estimates with F-statistics exceeding 1,100 (Tables 5–6) support the TOT analysis. The conversation-log analysis using an independent GPT-4o classification pipeline (Appendix B) and the cross-model comparison across six LLMs (Figures A1–A2) provide external validation that the directional pattern is not specific to one model. The finding that effects do not spill over to untreated patients (Figure A3) and do not persist in physician practice post-experiment (Figure 3) is economically informative. The paper is well-positioned for a general economics journal given its scale, policy relevance, and methodological rigor.","major_comments":[{"comment":"The central causal claim—that AI guardrail *directionality* propagates into clinical decisions—rests on the observation that effect directions mirror advice directions (§4.2). However, the paper acknowledges alternative mechanisms (§3.1): better-informed patients constrain overtreatment (reducing prescriptions) and improved symptom articulation leads to more appropriate test ordering (increasing testing). In a setting with documented overprescription of TCM and antibiotics, virtually any pre-visit information intervention could produce this pattern. The paper possesses conversation logs for 998 users—recording which topics each user discussed and what stance the AI took—but does not examine whether clinical effects vary with the specific content of those conversations. A natural test would be to split chatbot users by whether their conversation included medication cautions versus testing","section":null},{"comment":"The TOT estimates (Table 5) are approximately five times the ITT estimates, which the authors attribute to the 17.1% take-up rate. However, because take-up is self-selected (Appendix Table A2 shows younger, male, employed patients are more likely to use the chatbot), the LATE may reflect selection on unobservables correlated with both chatbot use and clinical outcomes. The paper should discuss whether compliers differ systematically from the full treated population in ways that affect the external validity of the TOT estimates, and whether the IV assumptions (exclusion restriction, monotonicity) are plausible given that treatment assignment is access to a tool that patients choose to use.","section":null},{"comment":"The welfare discussion (§7) is appropriately cautious but could be strengthened. The paper finds reduced prescribing (potentially welfare-improving if overtreatment is the margin) but also reduced patient compliance (potentially welfare-reducing if beneficial care is declined). The two-week revisit reduction (Table 3, column 6) is marginally significant and the authors note it may reflect care-seeking behavior rather than health improvements. Given that the paper shifts healthcare utilization patterns at scale based on AI developer priorities, a more structured welfare framework—even a simple conceptual one distinguishing between the AI developer's objective function and the health system's—would sharpen the policy implications.","section":null}],"minor_comments":[{"comment":"§4.1: The cross-model comparison (Figures A1–A2) uses only first-turn responses for all models, including the experiment chatbot, so the shares differ from Table 2. This is noted but could be more prominently flagged to avoid confusion when comparing Table 2 and the figures.","section":null},{"comment":"Table 2: The mention rate for antibiotics in patient messages is 0.1%, which seems extremely low. A brief note on whether this reflects patients rarely asking about antibiotics by name or a keyword-matching limitation would help interpretation.","section":null},{"comment":"§5.3: The patient-level persistence analysis tracks patients 3–6 months post-experiment. The chatbot link expires after the visit, but patients may have consulted other AI tools independently. The paper acknowledges this possibility but cannot distinguish between persistent preference shifts and continued AI use. A brief discussion of what survey or administrative data could help separate these channels would be useful.","section":null},{"comment":"Figure 5, Panels (b) and (c): The medication non-purchase and test non-completion rates show no significant differences, but the confidence intervals appear wide. Given the 13% survey response rate and potential selection, a note on whether these null results should be interpreted as evidence of no effect or as underpowered would be helpful.","section":null},{"comment":"§2.2: The paper states that drug expenditures represented 'more than 60 percent of inpatient spending and roughly 70 percent of outpatient spending' in Chinese public hospitals. Given ongoing payment reforms and the zero-markup policy mentioned later (footnote 26), clarifying whether these figures reflect the study period or an earlier benchmark would improve accuracy.","section":null},{"comment":"Appendix B.3: The stance classification treats 'the doctor may prescribe...' as a clean recommendation (footnote 21). This is a defensible coding choice but could inflate recommendation rates for models that use deferral phrasing. A sensitivity check excluding deferral-only recommendations would strengthen the cross-model comparison.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a strong field experiment on a highly relevant topic. The main substantive gap—the inability to distinguish directional advice effects from general information effects—is the most important concern, but it is addressable with the existing conversation-log data. If the authors can show that clinical effects vary with the specific content of patient-AI conversations (e.g., medication cautions vs. testing recommendations), the central claim would be substantially strengthened. Without that, the paper still makes a valuable contribution but the causal mechanism remains suggestive rather than demonstrated. The paper is well-suited for a top general economics journal."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The paper to know about: a preregistered RCT at a Chinese hospital where 10,000+ outpatient visits were randomized to give patients AI chatbot access before seeing a doctor. Prescriptions dropped 4.6pp, diagnostic testing rose 2.7pp, and the direction of each effect matches the chatbot's advice stance—heavy caution on medications, clean recommendations for testing. That mapping between AI guardrail directionality and clinical outcomes is the real contribution here, and it's genuinely new as the first large-scale randomized evidence on patient-side AI in healthcare. The two-layer physician randomization, the conversation log analysis, and the cross-model comparison (showing the caution pattern varies across developers, so it's design, not medical evidence) are all well done. The persistence analysis tracking patients months later is a nice touch. This deserves attention regardless of what happens in review. The soft spot is real but specific: the paper claims directionality is the operative mechanism, but it can't distinguish that from a general information effect. In a hospital system with documented overprescription of TCM and antibiotics, virtually any pre-visit information intervention would reduce unnecessary prescriptions. Better-informed patients constraining overtreatment is a standard credence-goods story that predicts the same pattern. The paper has conversation logs for 998 users—recording which topics each discussed and what stance the AI took—but never examines whether clinical effects vary with conversation content. That analysis would directly test whether patients who received medication cautions had different prescription outcomes than those who discussed other topics. Without it, the paper establishes that chatbot access changes outcomes and that the changes are consistent with directional advice, but not that directionality is the channel rather than general preparation. The 17% take-up rate and unspecified LLM model are secondary concerns—the IV estimates handle take-up, and the cross-model comparison partially addresses the model question. The differential survey response rates across groups warrant caution on the patient satisfaction findings but don't undermine the administrative outcomes. This is a well-designed experiment on an important question. The mechanism gap is addressable with data the authors already have, which is both the frustration and the opportunity. Recommend serious peer review—the experimental work earns it, and the conversation-log analysis could substantially strengthen the central claim if added.","headline":"Solid field experiment with a mechanism claim that the existing data could test but doesn't","tokens_in":40595,"tokens_out":1517,"would_cite":true,"duration_ms":63919,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"AI's Hidden Guardrails Reshape Real Doctors' Decisions","keywords":["generative AI","healthcare","field experiment","AI guardrails","credence goods","prescribing behavior","diagnostic testing","patient-physician relationship"],"falsifier":"If patients who used the chatbot but received only neutral, non-directional medical information (no systematic caution toward medications or encouragement of testing) showed the same reductions in prescribing and increases in testing, then the directional-advice propagation mechanism would be falsified—the effects would be attributable to general patient preparation rather than to the AI's encoded guardrails.","tokens_in":40158,"feed_emoji":"","tokens_out":1250,"duration_ms":165436,"temperature":0.7,"pith_summary":"When patients consult a generative AI chatbot before seeing a doctor, the advice they receive is not neutral. Liability-driven guardrails baked into the AI's training make it cautious about recommending medications—especially Traditional Chinese Medicine and antibiotics—while it freely and confidently recommends diagnostic tests. This paper shows, through a randomized field experiment at a Chinese hospital with over 10,000 outpatient visits, that this directional advice propagates into actual clinical decisions. Patients given chatbot access left with fewer prescriptions and more test orders, mirroring the AI's stance. The effects were concentrated among physicians receptive to patient input and those who prescribed most heavily at baseline. The AI did not make care neutral; it shifted it in the direction its developers' liability concerns pointed—away from drugs, toward testing. Patients also reported lower satisfaction and lower intended compliance with their doctors' recommendations, suggesting the AI shifted authority within the doctor-patient relationship. No spillovers to other patients and no lasting change in physician behavior were found; the effects traveled through patients' own use and partially persisted in their later visits.","feed_headline":"AI's Hidden Guardrails Reshape Real Doctors' Decisions","feed_subtitle":"Patients who consult AI before seeing a doctor leave with fewer prescriptions and more tests—mirroring the liability-driven biases bakedinto","key_machinery":"The mechanism has three links. First, AI developers encode liability-driven guardrails into their models, producing directional advice: heavy caution around medications (especially TCM and antibiotics) and clean encouragement of diagnostic testing. Second, patients carry this directional advice into their brief outpatient consultations, where they can request or question treatments. Third, physicians—particularly those open to patient input and those with intensive baseline prescribing—adjust their decisions in the direction the AI pointed. The two-layer randomization (physicians exposed vs. unexposed; within exposed physicians, patients treated vs. control) isolates the direct effect of AI-","core_discovery":"The central discovery is a propagation mechanism: defensive guardrails encoded in AI training—designed to limit developer liability by cautioning against medication recommendations while freely recommending diagnostic tests—travel through patient consultations into real clinical decisions at scale. In a randomized experiment, offering patients pre-visit AI chatbot access reduced prescription rates by 4.6 percentage points and increased diagnostic testing by 2.7 percentage points, with the direction of each effect matching the AI's advice stance. The chatbot cautioned against medications in 70-91% of mentions but issued clean test recommendations 94.5% of the time. Treatment-on-the-treated IV","pith_inferences":["If the directionality of AI advice reflects developer liability concerns rather than clinical evidence, then different developers' models would produce systematically different clinical outcomes for the same patients—a testable prediction partially supported by the cross-model variation in caution levels documented in the paper's appendix.","The finding that effects were largest among heavy prescribers suggests AI-assisted patients function as a check on physician-induced demand, but the paper cannot distinguish whether reduced prescriptions represent curtailed overtreatment or withheld beneficial care.","If the testing effect persists but the prescribing effect fades, this asymmetry in durability may reflect that diagnostic testing is easier for patients to request independently, while medication decisions remain more firmly physician-controlled—a distinction with implications for which AI-driven behavior changes are likely to endure.","The divergence between patient-reported worse communication and physician-reported better communication suggests the AI shifted expectations rather than communication quality itself, which if generalizable means AI tools may systematically lower patient satisfaction even when they improve consultation efficiency."],"forward_implications":["If AI guardrails propagate into clinical decisions, then the design choices of a small number of AI developers effectively become health policy, shifting prescribing and testing patterns for millions of patients without democratic deliberation or clinical oversight.","Unequal access to AI tools may widen health disparities: adopters were disproportionately younger, male, and employed, and the absence of spillovers means benefits do not diffuse to patients who lack access.","The same propagation mechanism likely operates in other credence-good markets—legal services, financial advice, skilled trades—wherever clients consult AI before meeting an expert, carrying the AI's directional priorities into the expert-client relationship.","If a single exposure to the chatbot durably shifted patient testing behavior months later, then even intermittent AI use could produce lasting changes in healthcare utilization patterns that outlast the tool itself.","Regulatory frameworks for medical AI may need to address not only what AI says directly to patients but also how it reshapes downstream expert decisions, since the clinical effects are mediated through the patient-physician interaction rather than through autonomous AI action."],"fun_headline_variants":["AI Chatbot Guardrails Propagate Into Real Clinical Decisions","Pre-Visit AI Access Cuts Prescriptions, Raises Diagnostic Tests","Patients Carrying AI Advice Shift Doctor Prescribing Behavior","Liability-Driven AI Bias Travels Through Patients to Doctors","AI Caution Against Medication Reaches Real Clinical Practice"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The paper assumes that the observed clinical changes were caused by patients carrying the AI's specific medication cautions and testing recommendations into their consultations, but it does not directly observe what patients actually said during the visits or whether they explicitly referenced the AI's advice. The effects could in principle arise from general patient preparation, improved symptom articulation, or other channels unrelated to the AI's directional stance.","fun_headline_variants_meta":{"raw":{"variants":["AI Chatbot Guardrails Propagate Into Real Clinical Decisions","Pre-Visit AI Access Cuts Prescriptions, Raises Diagnostic Tests","Patients Carrying AI Advice Shift Doctor Prescribing Behavior","Liability-Driven AI Bias Travels Through Patients to Doctors","AI Caution Against Medication Reaches Real Clinical Practice"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":621,"prompt_tokens":538,"completion_tokens":83,"prompt_tokens_details":null},"tokens_in":538,"tokens_out":83,"duration_ms":17396,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T02:37:54.857278+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If patients who used the chatbot but received only neutral, non-directional medical information (no systematic caution toward medications or encouragement of testing) showed the same reductions in prescribing and increases in testing, then the directional-advice propagation mechanism would be falsified—the effects would be attributable to general patient preparation rather than to the AI's encoded guardrails.","supporting_citations":[],"review_version":1}