Pith. sign in

REVIEW 3 major objections 6 minor 24 references

Directional AI Advice: Experimental Evidence from Healthcare

T0 review · 3 major / 6 minor · reviewed 2026-07-10 · glm-5.2

Pith's one-line read AI's Hidden Guardrails Reshape Real Doctors' Decisions

desk verdict Solid field experiment with a mechanism claim that the existing data could test but doesn't read the letter →

arxiv 2607.08706 v1 pith:XLCKUC2I submitted 2026-07-09 econ.GN q-fin.EC

classification econ.GNq-fin.EC
keywords generativeAIhealthcarefieldexperimentguardrailscredencegoodsprescribingbehaviordiagnostictestingpatient-physicianrelationship
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

When patients consult a generative AI chatbot before seeing a doctor, the advice they receive is not neutral. Liability-driven guardrails baked into the AI's training make it cautious about recommending medications—especially Traditional Chinese Medicine and antibiotics—while it freely and confidently recommends diagnostic tests. This paper shows, through a randomized field experiment at a Chinese hospital with over 10,000 outpatient visits, that this directional advice propagates into actual clinical decisions. Patients given chatbot access left with fewer prescriptions and more test orders, mirroring the AI's stance. The effects were concentrated among physicians receptive to patient input and those who prescribed most heavily at baseline. The AI did not make care neutral; it shifted it in the direction its developers' liability concerns pointed—away from drugs, toward testing. Patients also reported lower satisfaction and lower intended compliance with their doctors' recommendations, suggesting the AI shifted authority within the doctor-patient relationship. No spillovers to other patients and no lasting change in physician behavior were found; the effects traveled through patients' own use and partially persisted in their later visits.

What carries the argument

The mechanism has three links. First, AI developers encode liability-driven guardrails into their models, producing directional advice: heavy caution around medications (especially TCM and antibiotics) and clean encouragement of diagnostic testing. Second, patients carry this directional advice into their brief outpatient consultations, where they can request or question treatments. Third, physicians—particularly those open to patient input and those with intensive baseline prescribing—adjust their decisions in the direction the AI pointed. The two-layer randomization (physicians exposed vs. unexposed; within exposed physicians, patients treated vs. control) isolates the direct effect of AI-

What would settle it

If patients who used the chatbot but received only neutral, non-directional medical information (no systematic caution toward medications or encouragement of testing) showed the same reductions in prescribing and increases in testing, then the directional-advice propagation mechanism would be falsified—the effects would be attributable to general patient preparation rather than to the AI's encoded guardrails.

Watch

Extended reading notes

Core claim

The central discovery is a propagation mechanism: defensive guardrails encoded in AI training—designed to limit developer liability by cautioning against medication recommendations while freely recommending diagnostic tests—travel through patient consultations into real clinical decisions at scale. In a randomized experiment, offering patients pre-visit AI chatbot access reduced prescription rates by 4.6 percentage points and increased diagnostic testing by 2.7 percentage points, with the direction of each effect matching the AI's advice stance. The chatbot cautioned against medications in 70-91% of mentions but issued clean test recommendations 94.5% of the time. Treatment-on-the-treated IV

Load-bearing premise

The paper assumes that the observed clinical changes were caused by patients carrying the AI's specific medication cautions and testing recommendations into their consultations, but it does not directly observe what patients actually said during the visits or whether they explicitly referenced the AI's advice. The effects could in principle arise from general patient preparation, improved symptom articulation, or other channels unrelated to the AI's directional stance.

Editorial extensions

If this is right

  • If AI guardrails propagate into clinical decisions, then the design choices of a small number of AI developers effectively become health policy, shifting prescribing and testing patterns for millions of patients without democratic deliberation or clinical oversight.
  • Unequal access to AI tools may widen health disparities: adopters were disproportionately younger, male, and employed, and the absence of spillovers means benefits do not diffuse to patients who lack access.
  • The same propagation mechanism likely operates in other credence-good markets—legal services, financial advice, skilled trades—wherever clients consult AI before meeting an expert, carrying the AI's directional priorities into the expert-client relationship.
  • If a single exposure to the chatbot durably shifted patient testing behavior months later, then even intermittent AI use could produce lasting changes in healthcare utilization patterns that outlast the tool itself.
  • Regulatory frameworks for medical AI may need to address not only what AI says directly to patients but also how it reshapes downstream expert decisions, since the clinical effects are mediated through the patient-physician interaction rather than through autonomous AI action.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the directionality of AI advice reflects developer liability concerns rather than clinical evidence, then different developers' models would produce systematically different clinical outcomes for the same patients—a testable prediction partially supported by the cross-model variation in caution levels documented in the paper's appendix.
  • The finding that effects were largest among heavy prescribers suggests AI-assisted patients function as a check on physician-induced demand, but the paper cannot distinguish whether reduced prescriptions represent curtailed overtreatment or withheld beneficial care.
  • If the testing effect persists but the prescribing effect fades, this asymmetry in durability may reflect that diagnostic testing is easier for patients to request independently, while medication decisions remain more firmly physician-controlled—a distinction with implications for which AI-driven behavior changes are likely to endure.
  • The divergence between patient-reported worse communication and physician-reported better communication suggests the AI shifted expectations rather than communication quality itself, which if generalizable means AI tools may systematically lower patient satisfaction even when they improve consultation efficiency.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper reports a preregistered field experiment (AEARCTR-0015851) conducted at a large Chinese public hospital, randomizing 11,666 outpatient visits to patient access to an LLM-based chatbot prior to consultation. The authors document that the chatbot's advice is systematically directional: it cautions against medications (especially TCM and antibiotics) while issuing clean recommendations for diagnostic testing, a pattern they attribute to liability-driven guardrails in AI training. They then show that chatbot access reduces prescription rates by 4.6 pp (ITT) and increases diagnostic testing by 2.7 pp, with effects concentrated among physicians receptive to patient input and those with higher baseline prescribing intensity. Survey evidence reveals lower patient satisfaction and intended compliance, while physicians report clearer symptom descriptions. The paper contributes to the economics of AI, credence goods, and patient information literatures by examining how design choices encoded in AI systems propagate into real-world expert-client decisions.

Significance. The paper addresses a timely and important question: whether the defensive guardrails encoded in generative AI systems propagate into real-world decisions at scale. The experimental design is strong—preregistered, two-layer randomization with physician fixed effects, balance confirmed (Table 1), and ITT analysis preserving randomization integrity. The IV estimates with F-statistics exceeding 1,100 (Tables 5–6) support the TOT analysis. The conversation-log analysis using an independent GPT-4o classification pipeline (Appendix B) and the cross-model comparison across six LLMs (Figures A1–A2) provide external validation that the directional pattern is not specific to one model. The finding that effects do not spill over to untreated patients (Figure A3) and do not persist in physician practice post-experiment (Figure 3) is economically informative. The paper is well-positioned for a general economics journal given its scale, policy relevance, and methodological rigor.

major comments (3)
  1. The central causal claim—that AI guardrail *directionality* propagates into clinical decisions—rests on the observation that effect directions mirror advice directions (§4.2). However, the paper acknowledges alternative mechanisms (§3.1): better-informed patients constrain overtreatment (reducing prescriptions) and improved symptom articulation leads to more appropriate test ordering (increasing testing). In a setting with documented overprescription of TCM and antibiotics, virtually any pre-visit information intervention could produce this pattern. The paper possesses conversation logs for 998 users—recording which topics each user discussed and what stance the AI took—but does not examine whether clinical effects vary with the specific content of those conversations. A natural test would be to split chatbot users by whether their conversation included medication cautions versus testing
  2. The TOT estimates (Table 5) are approximately five times the ITT estimates, which the authors attribute to the 17.1% take-up rate. However, because take-up is self-selected (Appendix Table A2 shows younger, male, employed patients are more likely to use the chatbot), the LATE may reflect selection on unobservables correlated with both chatbot use and clinical outcomes. The paper should discuss whether compliers differ systematically from the full treated population in ways that affect the external validity of the TOT estimates, and whether the IV assumptions (exclusion restriction, monotonicity) are plausible given that treatment assignment is access to a tool that patients choose to use.
  3. The welfare discussion (§7) is appropriately cautious but could be strengthened. The paper finds reduced prescribing (potentially welfare-improving if overtreatment is the margin) but also reduced patient compliance (potentially welfare-reducing if beneficial care is declined). The two-week revisit reduction (Table 3, column 6) is marginally significant and the authors note it may reflect care-seeking behavior rather than health improvements. Given that the paper shifts healthcare utilization patterns at scale based on AI developer priorities, a more structured welfare framework—even a simple conceptual one distinguishing between the AI developer's objective function and the health system's—would sharpen the policy implications.
minor comments (6)
  1. §4.1: The cross-model comparison (Figures A1–A2) uses only first-turn responses for all models, including the experiment chatbot, so the shares differ from Table 2. This is noted but could be more prominently flagged to avoid confusion when comparing Table 2 and the figures.
  2. Table 2: The mention rate for antibiotics in patient messages is 0.1%, which seems extremely low. A brief note on whether this reflects patients rarely asking about antibiotics by name or a keyword-matching limitation would help interpretation.
  3. §5.3: The patient-level persistence analysis tracks patients 3–6 months post-experiment. The chatbot link expires after the visit, but patients may have consulted other AI tools independently. The paper acknowledges this possibility but cannot distinguish between persistent preference shifts and continued AI use. A brief discussion of what survey or administrative data could help separate these channels would be useful.
  4. Figure 5, Panels (b) and (c): The medication non-purchase and test non-completion rates show no significant differences, but the confidence intervals appear wide. Given the 13% survey response rate and potential selection, a note on whether these null results should be interpreted as evidence of no effect or as underpowered would be helpful.
  5. §2.2: The paper states that drug expenditures represented 'more than 60 percent of inpatient spending and roughly 70 percent of outpatient spending' in Chinese public hospitals. Given ongoing payment reforms and the zero-markup policy mentioned later (footnote 26), clarifying whether these figures reflect the study period or an earlier benchmark would improve accuracy.
  6. Appendix B.3: The stance classification treats 'the doctor may prescribe...' as a clean recommendation (footnote 21). This is a defensible coding choice but could inflate recommendation rates for models that use deferral phrasing. A sensitivity check excluding deferral-only recommendations would strengthen the cross-model comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the paper is an RCT with independently measured inputs and outcomes

full rationale

The paper's central claim—that AI guardrail directionality propagates into clinical decisions—rests on two independently measured data sources: (1) conversation logs classified by an external pipeline (GPT-4o stance classification applied to observed transcripts, Table 2), and (2) clinical outcomes from hospital administrative records estimated via randomized assignment (Tables 3–6). No variable is defined in terms of another in a circular fashion. The ITT and IV estimates use random assignment as an instrument, with first-stage F-statistics exceeding 1,100, which is standard and not self-referential. The cross-model comparison (Figures A1–A2) uses six external LLMs to validate that the directional pattern is not specific to the deployed model. No load-bearing self-citations are invoked; the theoretical framing draws on external work (Arrow 1963, Darby and Karni 1973, Finkelstein et al. 2022, Currie et al. 2011, etc.). The connection between conversation-log directionality and clinical-outcome direction is interpretive (the paper states 'the direction of each treatment effect mirrors the direction of the AI's advice'), not a definitional identity. Whether this interpretive link is fully identified versus alternative mechanisms (general information effects) is a correctness/identification concern, not a circularity concern—the paper's derivation chain does not reduce to its inputs by construction.

Assumptions & free parameters 1 free parameters · 4 assumptions · 1 invented entities

The paper relies on standard experimental economics methodology. The key axioms are domain assumptions about the mechanism linking AI advice to clinical outcomes, which are supported by the data but not independently verified through direct observation of consultation content.

free parameters (1)
  • None fitted to central claim
    The experimental design uses randomization rather than parameter fitting. Regression coefficients are estimated, not fitted to produce the claim. No free parameters are introduced to make the directional-propagation argument work.
assumptions (4)
  • domain assumption The chatbot's directional stance (cautioning against medications, recommending testing) reflects liability-driven guardrails encoded by AI developers rather than neutral medical evidence.
    Stated in §1 and §2.1. Supported by cross-model comparison (§4.1, Figures A1-A2) showing variation across models, and by cited developer guidelines (OpenAI, Anthropic). The causal attribution to liability motives is inferred, not directly tested.
  • domain assumption Patients who used the chatbot carried its advice into their consultations in a way that influenced physician decisions.
    Invoked in §4.2: 'the direction of each treatment effect mirrors the direction of the AI's advice.' The paper does not directly observe consultation content but infers the channel from the correspondence between advice direction and outcome direction.
  • domain assumption The 17% take-up rate does not introduce selection bias that invalidates the ITT estimates.
    Discussed in §3.4 and §4.3. Take-up is self-selected (younger, male, employed patients more likely to use), but ITT estimates preserve randomization. The TOT estimates are acknowledged as local average treatment effects for compliers.
  • domain assumption Standard errors need not be clustered at the physician level for the main results.
    Stated in §4.3 footnote 24, citing Abadie et al. (2023). Robustness checks with clustered SEs (Appendix Tables A3-A6) show core results on prescribing and testing remain significant, though the revisit effect loses marginal significance.
invented entities (1)
  • None independent evidence
    purpose: No new entities, particles, forces, or constructs are postulated.
    The paper is an empirical study using existing technologies and standard econometric methods.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Directional AI Advice: Experimental Evidence from Healthcare." pith.science (2026). https://pith.science/paper/XLCKUC2I

@misc{pith2026260708706,
  author       = {Pith},
  title        = {Pith review of: Directional AI Advice: Experimental Evidence from Healthcare},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XLCKUC2I}},
  note         = {Machine review of arXiv:2607.08706}
}
read the original abstract

Generative AI is fast becoming the first place people turn for expert advice. The advice it provides can be directional rather than neutral, shaped in part by the choices of its designers and regulators. When clients consult AI before meeting an expert, they carry this directional advice into a relationship that once rested on the expert's judgment alone. We study its consequences in healthcare through a large-scale preregistered field experiment at a Chinese hospital, where we randomize patients' access to an AI chatbot before their outpatient visit. Examination of the conversation logs shows that the chatbot routinely cautions against the use of medications, especially Traditional Chinese Medicine and antibiotics, while issuing clean recommendations for diagnostic testing, consistent with the liability-driven guardrails encoded in AI training. This directionality propagates into clinical practice. Prescription rates decline among treated patients while diagnostic testing increases, and these effects are more pronounced among physicians who are receptive to patient input and those with more intensive prescribing styles. Beyond shifting healthcare utilization, survey results show that AI access reduces patient compliance and satisfaction, shifting the balance of authority between patients and physicians.

Figures

Figures reproduced from arXiv: 2607.08706 by the authors.

Figure 1
Figure 1. [PITH_FULL_IMAGE:figures/full_fig_p040_1.png] view at source ↗
Figure 2
Figure 2. Prescriptions and Clinical Practices by Treatment Status Notes: This figure reports mean clinical outcomes for patients in the Treatment, Control, and Unexposed groups, with 95% confidence intervals (dashed lines). 41 [PITH_FULL_IMAGE:figures/full_fig_p041_2.png] view at source ↗
Figure 3
Figure 3. Dynamic Effects of Exposure to Treated Patients on Prescriptions and Clinical Practices 42 [PITH_FULL_IMAGE:figures/full_fig_p042_3.png] view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: Dynamic Effects of Exposure to Treated Patients on Prescriptions and Clinical Practices (Continued) Notes: This figure presents the point estimates with 95% confidence intervals from equation (2). Data are aggregated to the physician-week level using all patient visits…
Figure 4
Figure 4. Figure 4: Persistent Effects of Treatment Assignment on Outpatient Utilization and Clinical Practice in the Post-Experiment Period (October–December 2025) Notes: This figure examines whether assignment to chatbot access during the experiment continues to affect outpatient utiliz…
Figure 5
Figure 5. Figure 5: Patient Behavior and Perceptions by Treatment Status Notes: Panels (a)–(c) report outcomes constructed from administrative medical records and capture patient attrition and realized post-consultation behavior. Panels (d)–(f) report outcomes from a post-consultation pat…
Figure 6
Figure 6. Figure 6: Physicians’ Perceptions of Patient Behaviors during Consultations Notes: This figure reports physicians’ survey responses collected after the conclusion of the experiment, comparing physicians randomized to the Exposed and Unexposed groups. Outcomes reflect physicians’…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

24 extracted references · 24 canonical work pages

  1. [1]

    When Should You Adjust Standard Errors for Clustering?,

    Abadie, Alberto, Susan Athey, Guido W Imbens, and Jeffrey M Wooldridge, “When Should You Adjust Standard Errors for Clustering?,”The Quarterly Journal of Economics, 2023,138(1), 1–35. Abaluck, Jason, Robert Pless, Nirmal Ravi, Anja Sautmann, and Aaron Schwartz, “Does LLM Assistance Improve Healthcare Delivery? An Evaluation Using On-site Physicians and La...

  2. [2]

    The Value of Rating Systems in Credence Goods Markets,

    Angerer, Silvia, Daniela Glätzle-Rützler, Wanda Mimra, Thomas Rittmannsberger, and Chris- tian Waibel, “The Value of Rating Systems in Credence Goods Markets,”The Economic Journal, 2026, p. ueag011. Arrow, Kenneth J, “Uncertainty and the Welfare Economics of Medical Care,”The American Economic Review, 1963,53(5), 941–973. Artmann, Elisabeth, Hessel Ooster...

  3. [3]

    Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum,

    Ayers, John W, Adam Poliak, Mark Dredze, Eric C Leas, Zechariah Zhu, Jessica B Kelley, Dennis J Faix, Aaron M Goodman, Christopher A Longhurst, Michael Hogarth et al., “Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum,”JAMA internal medicine, 2023,183(6), 589–596. Badinski, Ivan, ...

  4. [4]

    What Drives Taxi Drivers? A Field Experiment on Fraud in a Market for Credence Goods,

    Balafoutas, Loukas, Adrian Beck, Rudolf Kerschbamer, and Matthias Sutter, “What Drives Taxi Drivers? A Field Experiment on Fraud in a Market for Credence Goods,”Review of Economic Studies, 2013,80(3), 876–891. , Rudolf Kerschbamer, and Matthias Sutter, “Second-Degree Moral Hazard in a Real-World Credence Goods Market,”The Economic Journal, 2017,127(599), ...

  5. [5]

    Do Pharmacists Buy Bayer? Informed Shoppers and the Brand Premium,

    Bronnenberg, Bart J, Jean-Pierre Dubé, Matthew Gentzkow, and Jesse M Shapiro, “Do Pharmacists Buy Bayer? Informed Shoppers and the Brand Premium,”The Quarterly Journal of Economics, 2015,130(4), 1669–1726. Brynjolfsson, Erik, Danielle Li, and Lindsey Raymond, “Generative AI at Work,”The Quarterly Journal of Economics, 2025,140(2), 889–942. Busse, Meghan R...

  6. [6]

    Physicians Treating Physicians: Relational and Informational Advantages in Treatment and Survival,

    Chen, Stacey H, Jennjou Chen, Hongwei Chuang, and Tzu-Hsin Lin, “Physicians Treating Physicians: Relational and Informational Advantages in Treatment and Survival,”Journal of Labor Economics, 2025, 43(1), 15–46. Chen, Yiqun, Petra Persson, and Maria Polyakova, “The Roots of Health Inequality and the Value of Intrafamily Expertise,”American Economic Journa...

  7. [7]

    DoctorDecision-MakingandPatientOutcomes,

    Currie, Janet, WBentleyMacLeod, andKateMusen, “DoctorDecision-MakingandPatientOutcomes,” Journal of Economic Literature, 2026,64(1), 141–194. , Wanchuan Lin, and Juanjuan Meng, “Social Networks and Externalities from Gift Exchange: Evidence from a Field Experiment,”Journal of Public Economics, 2013,107, 19–30. , , and , “Addressing Antibiotic Abuse in Chi...

  8. [8]

    Is More Information Better? The Effects of “Report Cards

    34 Dranove, David, Daniel Kessler, Mark McClellan, and Mark Satterthwaite, “Is More Information Better? The Effects of “Report Cards” on Health Care Providers,”Journal of Political Economy, 2003,111 (3), 555–588. Dulleck, Uwe and Rudolf Kerschbamer, “On Doctors, Mechanics, and Computer Specialists: The Economics of Credence Goods,”Journal of Economic lite...

Show all 24 references
  1. [9]

    Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?,

    Filippas, Apostolos, John J Horton, and Benjamin S Manning, “Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?,” in “Proceedings of the 25th ACM Conference on Economics and Computation” 2024, pp. 614–615. Finkelstein, Amy, Matthew Gentzko...

  2. [10]

    Health Services as Credence Goods: A Field Experiment,

    Gottschalk, Felix, Wanda Mimra, and Christian Waibel, “Health Services as Credence Goods: A Field Experiment,”The Economic Journal, 2020,130(629), 1346–1383. Gruber, Jonathan and Maria Owings, “Physician Financial Incentives and Cesarean Section Delivery,” RAND Journal of Econ...

  3. [11]

    Credence Goods Markets, Online Information and Repair Prices: A Natural Field Experiment,

    Kerschbamer, Rudolf, Daniel Neururer, and Matthias Sutter, “Credence Goods Markets, Online Information and Repair Prices: A Natural Field Experiment,”Journal of Public Economics, 2023,222, 104891. Kessler, Daniel and Mark McClellan, “Do Doctors Practice Defensive Medicine?,”Th...

  4. [12]

    Overprescribing in China, Driven by Financial Incentives, Results in Very High Use of Antibiotics, Injections, and Corticosteroids,

    Li, Yongbin, Jing Xu, Fang Wang, Bin Wang, Liqun Liu, Wanli Hou, Hong Fan, Yeqing Tong, Juan Zhang, and Zuxun Lu, “Overprescribing in China, Driven by Financial Incentives, Results in Very High Use of Antibiotics, Injections, and Corticosteroids,”Health affairs, 2012,31(5), 10...

  5. [13]

    Does Patient Demand Contribute to the Overuse of Prescription Drugs?,

    , , and Simone Schaner, “Does Patient Demand Contribute to the Overuse of Prescription Drugs?,” American Economic Journal: Applied Economics, 2022,14(1), 225–260. Lu, Fangwen, “Insurance Coverage and Agency Problems in Doctor Prescriptions: Evidence from a Field Experiment in ...

  6. [14]

    Physician Agency,

    McGuire, Thomas G, “Physician Agency,”Handbook of Health Economics, 2000,1, 461–536. Mimra, Wanda, Alexander Rasch, and Christian Waibel, “Second Opinions in Markets for Expert Services: Experimental Evidence,”Journal of Economic Behavior & Organization, 2016,131, 106–125. Mul...

  7. [15]

    Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,

    Noy, Shakked and Whitney Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,”Science, 2023,381(6654), 187–192. Peng, Sida, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer, “The Impact of AI on Developer Productivity: Evidence fro...

  8. [16]

    What Drives Fraud in a Credence Goods Market?–Evidence from a Field Study,

    Rasch, Alexander and Christian Waibel, “What Drives Fraud in a Credence Goods Market?–Evidence from a Field Study,”Oxford Bulletin of Economics and Statistics, 2018,80(3), 605–624. Reimers, Imke and Joel Waldfogel, “AI and the Quantity and Quality of Creative Products: Have LL...

  9. [17]

    Disclaimers and Referral Patterns for Medical Advice across Urgency Levels: Large Language Model Evaluation Study,

    Reis, Florian, Louis Agha-Mir-Salim, Richard Hickstein, Moritz Reis, Sophie K Piper, Felix Balzer, and Sebastian Daniel Boie, “Disclaimers and Referral Patterns for Medical Advice across Urgency Levels: Large Language Model Evaluation Study,”Journal of Medical Internet Researc...

  10. [18]

    A Longitudinal Analysis of Declining Medical Safety Messaging in Generative AI Models,

    Working paper. Sharma, Sonali, Ahmed M Alaa, and Roxana Daneshjou, “A Longitudinal Analysis of Declining Medical Safety Messaging in Generative AI Models,”npj Digital Medicine, 2025,8(1),

  11. [19]

    Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment,

    Spatharioti, Sofia Eleni, David M Rothschild, Daniel G Goldstein, and Jake M Hofman, “Comparing Traditional and LLM-based Search for Consumer Choice: A Randomized Experiment,”arXiv preprint arXiv:2307.03744,

  12. [20]

    Decision Aids for People Facing Health Treatment or Screening Decisions,

    Stacey, Dawn, France Légaré, Krystina Lewis, Michael J Barry, Carol L Bennett, Karen B Eden, Margaret Holmes-Rovner, Hilary Llewellyn-Thomas, Anne Lyddiatt, Richard Thomson et al., “Decision Aids for People Facing Health Treatment or Screening Decisions,”Cochrane Database of S...

  13. [21]

    Because a patient may appear in more than one experimental-period visit with differing assignments, we group each patient by their first such visit

    We track the same patients from October to December 2025 by matching their subsequent outpatient visits through unique patient identifiers. Because a patient may appear in more than one experimental-period visit with differing assignments, we group each patient by their first ...

  14. [22]

    All regressions control for patient characteristics (age, gender, occupation, and an indicator for first visit) and include physician fixed effects

    High TCM (High Antibiotic, High Opioid) is an indicator equal to one if the physician’s pre-experiment mean prescribing rate for TCM (antibiotics, opioids) exceeds the department median, and zero otherwise. All regressions control for patient characteristics (age, gender, occu...

  15. [23]

    Any visit (dummy) Number of visits (1) (2) (3) (4) Treated−0.011−0.009−0.022−0.012 (0.011) (0.011) (0.047) (0.046) Patient Char. N Y N Y Mean 0.46 0.46 1.24 1.24 Observations 8,166 8,166 8,166 8,166 Notes:This table reports the persistent effects of treatment assignment on pat...

  16. [24]

    very satisfied

    Prescribed medication Diagnostic tests (1) (2) (3) (4) Treated−0.002−0.003 0.016 ∗ 0.016∗ (0.008) (0.008) (0.010) (0.009) Patient Char. N Y N Y Mean 0.83 0.83 0.21 0.21 Observations 10,149 10,149 10,149 10,149 Notes:This table reports persistent effects of treatment assignment...

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.