Pith. sign in

REVIEW 2 major objections 1 minor 5 references

Analysis of 15,000 user reviews identifies three recurring breakdowns in AI healthcare chatbots and links privacy concerns to the worst experiences.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.3

2026-06-26 02:07 UTC pith:P4CIQZAK

load-bearing objection The paper maps three breakdown categories from 15k chatbot reviews but skips key validation and sampling details. the 2 major comments →

arxiv 2606.27302 v1 pith:P4CIQZAK submitted 2026-06-25 cs.HC cs.AI

AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns

classification cs.HC cs.AI
keywords AI healthcare chatbotsuser reviewsinformation infrastructurebreakdownsprivacy concernstopic modelingdigital health systemsuser experience
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper studies how AI healthcare chatbots perform when people seek health information or manage their own care. It draws on more than 15,000 reviews across 59 apps and applies topic modeling plus interpretive reading to surface patterns in what goes wrong. Three categories of problems appear repeatedly: barriers to access and unreliable service, weak user experience and interaction quality, and problems with billing plus customer support. Privacy and security worries stand out as tied to the most negative user reactions. Treating the chatbots as information infrastructures makes visible how these failures shape everyday use of digital health tools.

Core claim

By analyzing over 15,000 user reviews from 59 AI healthcare chatbot apps through topic modeling and interpretive analysis, the study identifies three recurring breakdowns: access barriers and service unreliability, user experience and interaction quality, and billing and customer support issues. Privacy and security concerns are associated with the most negative experiences. Framing these systems as information infrastructures shows how failures in access, usability, and trust affect users who rely on them for health information and self-management.

What carries the argument

Topic modeling combined with interpretive analysis of user reviews, applied to AI healthcare chatbots viewed as information infrastructures.

Load-bearing premise

The 15,000 reviews from the 59 apps are treated as representative enough of real-world experiences that the identified breakdowns reflect actual patterns rather than sampling or interpretation artifacts.

What would settle it

A new collection of several thousand recent reviews from a comparable set of apps that yields either no dominant breakdown categories or shows privacy concerns uncorrelated with negative sentiment would undermine the central findings.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Designers can target the three breakdown areas to reduce user friction in health chatbots.
  • Policymakers receive evidence that privacy protections matter for maintaining user trust in these tools.
  • Information professionals gain concrete categories for evaluating and improving digital health systems.
  • Users seeking health information may encounter repeated access, interaction, or billing problems that limit the chatbots' usefulness.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Similar large-scale review analysis could be extended to other consumer-facing AI health tools to check whether the same breakdown types appear.
  • If the identified issues remain unaddressed, reliance on these chatbots for self-management could decline over time as users seek alternatives.
  • The infrastructure framing suggests that isolated fixes in one area, such as privacy, may not resolve problems rooted in access or interaction quality.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper analyzes over 15,000 user reviews from 59 AI healthcare chatbot apps. Using topic modeling and interpretive analysis, it identifies three recurring breakdowns (access barriers and service unreliability, user experience and interaction quality, billing and customer support issues) and reports that privacy and security concerns are associated with the most negative experiences. The work frames these chatbots as information infrastructures to draw implications for designers, policymakers, and information professionals.

Significance. If the analytic pipeline is shown to be reliable, the study supplies a large-scale empirical view of real-world user breakdowns in AI healthcare chatbots and links specific concerns (especially privacy) to sentiment. This could usefully inform infrastructure-oriented design and policy work in digital health.

major comments (2)
  1. [Abstract/Methods] Abstract and Methods: the central claims rest on topic modeling plus interpretive analysis, yet the manuscript supplies no validation steps, inter-rater reliability metrics, topic-model hyperparameters, or controls for review-platform bias. These omissions are load-bearing for the identification of the three breakdowns and the privacy-negative-sentiment association.
  2. [Data Collection] Data section: the 59 apps are described as the study corpus, but no sampling frame, inclusion criteria, or fraction of total available reviews is reported. This directly affects the representativeness claim for self-selected app-store data.
minor comments (1)
  1. [Abstract] Abstract: the sentence on privacy could be tightened to state the exact association measure used (e.g., sentiment-score comparison) rather than the current phrasing.

Simulated Author's Rebuttal

2 responses · 0 unresolved

Thank you for the constructive feedback on our manuscript. We address each major comment below, agreeing where revisions are needed to improve transparency and rigor.

read point-by-point responses
  1. Referee: [Abstract/Methods] Abstract and Methods: the central claims rest on topic modeling plus interpretive analysis, yet the manuscript supplies no validation steps, inter-rater reliability metrics, topic-model hyperparameters, or controls for review-platform bias. These omissions are load-bearing for the identification of the three breakdowns and the privacy-negative-sentiment association.

    Authors: We agree these details are essential. In the revised manuscript we will expand the Methods section to specify the topic modeling algorithm, number of topics, hyperparameters, and software used, along with model validation metrics such as coherence scores. We will also describe the interpretive coding process in detail and report inter-rater reliability statistics calculated on a coded subset. A new limitations paragraph will discuss review-platform bias and any mitigation steps (e.g., use of verified reviews). These additions will directly underpin the reported breakdowns and privacy-sentiment association. revision: yes

  2. Referee: [Data Collection] Data section: the 59 apps are described as the study corpus, but no sampling frame, inclusion criteria, or fraction of total available reviews is reported. This directly affects the representativeness claim for self-selected app-store data.

    Authors: We acknowledge the reporting gap. The revision will add a dedicated Data Collection subsection detailing the sampling frame, search terms and inclusion criteria used to identify the 59 AI healthcare chatbot apps, the collection time window, and the total reviews available versus those analyzed (noting platform API limits). We will also explicitly frame the self-selected nature of app-store data as a limitation and discuss implications for representativeness. revision: yes

Circularity Check

0 steps flagged

No circularity: empirical observational study grounded in external user data

full rationale

This is a qualitative empirical paper that applies topic modeling and interpretive analysis to 15,000 external user reviews. No equations, fitted parameters, predictions, or derivations exist that could reduce to inputs by construction. No self-citation chains, uniqueness theorems, or ansatzes are invoked as load-bearing for the central claims. The analysis is self-contained against external benchmarks (app-store reviews), satisfying the default expectation of no significant circularity.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Empirical HCI study; no mathematical free parameters, no domain axioms beyond standard assumptions of topic modeling validity, and no invented entities. Relies on prior literature for topic modeling and interpretive methods.

pith-pipeline@v0.9.1-grok · 5659 in / 1128 out tokens · 26348 ms · 2026-06-26T02:07:57.409731+00:00 · methodology

0 comments
read the original abstract

AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on users remains to be studied. This study examines over 15,000 user reviews from 59 AI healthcare chatbot apps to explore how these systems function in everyday informational and emotional contexts. Topic modeling and interpretive analysis identify three recurring breakdowns: access barriers and service unreliability, user experience and interaction quality, and billing and customer support issues. Privacy and security concerns are associated with the most negative experiences. By framing AI healthcare chatbots as information infrastructures, our findings highlight how failures in access, usability, and trust affect users, offering actionable insights for designers, policymakers, and information professionals aiming to improve digital health systems.

Figures

Figures reproduced from arXiv: 2606.27302 by Ece Gumusel, Masooda Bashir, Muhammad Hassan, Ramazan Yener.

Figure 1
Figure 1. Figure 1: Percentage distribution of the three concern categories by platform [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Rating distributions across the topics Category 3: Billing, Customer Support, and Trust This category captures topics related to billing practices, customer support, trust and users’ billing concerns. It covers two topics: Customer Support and Billing (n = 689; M = 1.53) and Charges and Refund Concerns (n = 2,086; M = 1.55). Although it is the smallest category by volume (2,775 reviews; 18.4% of the corpus… view at source ↗
Figure 3
Figure 3. Figure 3: Rating distributions for SPR vs. non-flagged reviews [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

5 extracted references · 1 canonical work pages

  1. [1]

    Those things are written by lawyers, and programmers are reading that

    Andalibi, N. (2020). Disclosure, privacy, and stigma on social media: Examining non-disclosure of distressing experiences. ACM transactions on computer-human interaction (TOCHI), 27(3), 1-43. Andalibi, N., Haimson, O. L., Choudhury, M. D., & Forte, A. (2018). Social support, reciprocity, and anonymity in responses to sexual abuse disclosures on social med...

  2. [2]

    G., Tong, H

    Laranjo, L., Dunn, A. G., Tong, H. L., Kocaballi, A. B., Chen, J., Bashir, R., ... & Coiera, E. (2018). Conversational agents in healthcare: a systematic review. Journal of the American Medical Informatics Association, 25(9), 1248-1258. Li, H., Chen, Y., Luo, J., Wang, J., Peng, H., Kang, Y., ... & Song, Y. (2023). Privacy in large language models: Attack...

  3. [3]

    umbrella concepts

    Ocepek, M. G. (2018). Bringing out the everyday in everyday information behavior. Journal of Documentation, 74(2), 398-411. Pagano, D., & Maalej, W. (2013, July). User feedback in the appstore: An empirical study. In 2013 21st IEEE international requirements engineering conference (RE) (pp. 125-134). IEEE. Reddy, S., Allan, S., Coghlan, S., & Cooper, P. (...

  4. [4]

    N., Wisniewski, H., Halamka, J

    Vaidyam, A. N., Wisniewski, H., Halamka, J. D., Kashavan, M. S., & Torous, J. B. (2019). Chatbots and conversational agents in mental health: a review of the psychiatric landscape. The Canadian Journal of Psychiatry, 64(7), 456-464. Van Dijck, J., Poell, T., & De Waal, M. (2018). The platform society: Public values in a connective world. Oxford university...

  5. [5]

    S., & Davidson, E

    Winter, J. S., & Davidson, E. (2019). Big data governance of personal health information and challenges to contextual integrity. The Information Society, 35(1), 36-51. Yener, R., Chen, G. H., Gumusel, E., & Bashir, M. (2025). Can I Trust This Chatbot? Assessing User Privacy in AI‐ Healthcare Chatbot Applications. Proceedings of the Association for Informa...