REVIEW 2 major objections 1 minor 5 references
Analysis of 15,000 user reviews identifies three recurring breakdowns in AI healthcare chatbots and links privacy concerns to the worst experiences.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-26 02:07 UTC pith:P4CIQZAK
load-bearing objection The paper maps three breakdown categories from 15k chatbot reviews but skips key validation and sampling details. the 2 major comments →
AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By analyzing over 15,000 user reviews from 59 AI healthcare chatbot apps through topic modeling and interpretive analysis, the study identifies three recurring breakdowns: access barriers and service unreliability, user experience and interaction quality, and billing and customer support issues. Privacy and security concerns are associated with the most negative experiences. Framing these systems as information infrastructures shows how failures in access, usability, and trust affect users who rely on them for health information and self-management.
What carries the argument
Topic modeling combined with interpretive analysis of user reviews, applied to AI healthcare chatbots viewed as information infrastructures.
Load-bearing premise
The 15,000 reviews from the 59 apps are treated as representative enough of real-world experiences that the identified breakdowns reflect actual patterns rather than sampling or interpretation artifacts.
What would settle it
A new collection of several thousand recent reviews from a comparable set of apps that yields either no dominant breakdown categories or shows privacy concerns uncorrelated with negative sentiment would undermine the central findings.
If this is right
- Designers can target the three breakdown areas to reduce user friction in health chatbots.
- Policymakers receive evidence that privacy protections matter for maintaining user trust in these tools.
- Information professionals gain concrete categories for evaluating and improving digital health systems.
- Users seeking health information may encounter repeated access, interaction, or billing problems that limit the chatbots' usefulness.
Where Pith is reading between the lines
- Similar large-scale review analysis could be extended to other consumer-facing AI health tools to check whether the same breakdown types appear.
- If the identified issues remain unaddressed, reliance on these chatbots for self-management could decline over time as users seek alternatives.
- The infrastructure framing suggests that isolated fixes in one area, such as privacy, may not resolve problems rooted in access or interaction quality.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes over 15,000 user reviews from 59 AI healthcare chatbot apps. Using topic modeling and interpretive analysis, it identifies three recurring breakdowns (access barriers and service unreliability, user experience and interaction quality, billing and customer support issues) and reports that privacy and security concerns are associated with the most negative experiences. The work frames these chatbots as information infrastructures to draw implications for designers, policymakers, and information professionals.
Significance. If the analytic pipeline is shown to be reliable, the study supplies a large-scale empirical view of real-world user breakdowns in AI healthcare chatbots and links specific concerns (especially privacy) to sentiment. This could usefully inform infrastructure-oriented design and policy work in digital health.
major comments (2)
- [Abstract/Methods] Abstract and Methods: the central claims rest on topic modeling plus interpretive analysis, yet the manuscript supplies no validation steps, inter-rater reliability metrics, topic-model hyperparameters, or controls for review-platform bias. These omissions are load-bearing for the identification of the three breakdowns and the privacy-negative-sentiment association.
- [Data Collection] Data section: the 59 apps are described as the study corpus, but no sampling frame, inclusion criteria, or fraction of total available reviews is reported. This directly affects the representativeness claim for self-selected app-store data.
minor comments (1)
- [Abstract] Abstract: the sentence on privacy could be tightened to state the exact association measure used (e.g., sentiment-score comparison) rather than the current phrasing.
Simulated Author's Rebuttal
Thank you for the constructive feedback on our manuscript. We address each major comment below, agreeing where revisions are needed to improve transparency and rigor.
read point-by-point responses
-
Referee: [Abstract/Methods] Abstract and Methods: the central claims rest on topic modeling plus interpretive analysis, yet the manuscript supplies no validation steps, inter-rater reliability metrics, topic-model hyperparameters, or controls for review-platform bias. These omissions are load-bearing for the identification of the three breakdowns and the privacy-negative-sentiment association.
Authors: We agree these details are essential. In the revised manuscript we will expand the Methods section to specify the topic modeling algorithm, number of topics, hyperparameters, and software used, along with model validation metrics such as coherence scores. We will also describe the interpretive coding process in detail and report inter-rater reliability statistics calculated on a coded subset. A new limitations paragraph will discuss review-platform bias and any mitigation steps (e.g., use of verified reviews). These additions will directly underpin the reported breakdowns and privacy-sentiment association. revision: yes
-
Referee: [Data Collection] Data section: the 59 apps are described as the study corpus, but no sampling frame, inclusion criteria, or fraction of total available reviews is reported. This directly affects the representativeness claim for self-selected app-store data.
Authors: We acknowledge the reporting gap. The revision will add a dedicated Data Collection subsection detailing the sampling frame, search terms and inclusion criteria used to identify the 59 AI healthcare chatbot apps, the collection time window, and the total reviews available versus those analyzed (noting platform API limits). We will also explicitly frame the self-selected nature of app-store data as a limitation and discuss implications for representativeness. revision: yes
Circularity Check
No circularity: empirical observational study grounded in external user data
full rationale
This is a qualitative empirical paper that applies topic modeling and interpretive analysis to 15,000 external user reviews. No equations, fitted parameters, predictions, or derivations exist that could reduce to inputs by construction. No self-citation chains, uniqueness theorems, or ansatzes are invoked as load-bearing for the central claims. The analysis is self-contained against external benchmarks (app-store reviews), satisfying the default expectation of no significant circularity.
Axiom & Free-Parameter Ledger
read the original abstract
AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on users remains to be studied. This study examines over 15,000 user reviews from 59 AI healthcare chatbot apps to explore how these systems function in everyday informational and emotional contexts. Topic modeling and interpretive analysis identify three recurring breakdowns: access barriers and service unreliability, user experience and interaction quality, and billing and customer support issues. Privacy and security concerns are associated with the most negative experiences. By framing AI healthcare chatbots as information infrastructures, our findings highlight how failures in access, usability, and trust affect users, offering actionable insights for designers, policymakers, and information professionals aiming to improve digital health systems.
Figures
Reference graph
Works this paper leans on
-
[1]
Those things are written by lawyers, and programmers are reading that
Andalibi, N. (2020). Disclosure, privacy, and stigma on social media: Examining non-disclosure of distressing experiences. ACM transactions on computer-human interaction (TOCHI), 27(3), 1-43. Andalibi, N., Haimson, O. L., Choudhury, M. D., & Forte, A. (2018). Social support, reciprocity, and anonymity in responses to sexual abuse disclosures on social med...
2020
-
[2]
Laranjo, L., Dunn, A. G., Tong, H. L., Kocaballi, A. B., Chen, J., Bashir, R., ... & Coiera, E. (2018). Conversational agents in healthcare: a systematic review. Journal of the American Medical Informatics Association, 25(9), 1248-1258. Li, H., Chen, Y., Luo, J., Wang, J., Peng, H., Kang, Y., ... & Song, Y. (2023). Privacy in large language models: Attack...
-
[3]
umbrella concepts
Ocepek, M. G. (2018). Bringing out the everyday in everyday information behavior. Journal of Documentation, 74(2), 398-411. Pagano, D., & Maalej, W. (2013, July). User feedback in the appstore: An empirical study. In 2013 21st IEEE international requirements engineering conference (RE) (pp. 125-134). IEEE. Reddy, S., Allan, S., Coghlan, S., & Cooper, P. (...
2018
-
[4]
N., Wisniewski, H., Halamka, J
Vaidyam, A. N., Wisniewski, H., Halamka, J. D., Kashavan, M. S., & Torous, J. B. (2019). Chatbots and conversational agents in mental health: a review of the psychiatric landscape. The Canadian Journal of Psychiatry, 64(7), 456-464. Van Dijck, J., Poell, T., & De Waal, M. (2018). The platform society: Public values in a connective world. Oxford university...
2019
-
[5]
S., & Davidson, E
Winter, J. S., & Davidson, E. (2019). Big data governance of personal health information and challenges to contextual integrity. The Information Society, 35(1), 36-51. Yener, R., Chen, G. H., Gumusel, E., & Bashir, M. (2025). Can I Trust This Chatbot? Assessing User Privacy in AI‐ Healthcare Chatbot Applications. Proceedings of the Association for Informa...
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.