REVIEW 4 major objections 6 minor 2 references
Performance of a large language model-Artificial Intelligence based chatbot for counseling patients with sexually transmitted infections and genital diseases
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A GPT-4-based chatbot built specifically for sexual health counseling can deliver accurate, empathetic, and factually correct STI information, this proof-of-concept study reports, with relevance as its main weakness.
desk verdict Single-arm expert role-play evaluation of a GPT-4 STI chatbot; narrow results plausible, broad 'accurate/correct' claim not supported by the 5.0/0-SD correctness scores. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a multi-agent system over GPT-4-0613 constrained by a Deterministic Finite Automaton-style conversational flow, plus four overlaid prompt modules: a general STI information module that reproduces a venereologist's stepwise diagnostic reasoning, an emotional recognition module, an acute stress disorder detection module, and a psychotherapy module, with a parallel question-suggestion agent. The DFA principle supplies the state machine that governs transitions between modules; the prompt engineering supplies the persona, the metacognitive instruction to reason through differentials, and the rule always to recommend professional medical evaluation. The evaluation instrument is the six-item Numerical Rating Scale used by the venereologist actors.
What would settle it
Re-run the same 30 patient prompts with the STI-specific modules stripped out, using only the underlying GPT-4 model, and compare the six rating criteria; if the generic model matches Otiz's scores, the claimed contribution of the DFA and module architecture collapses. Independent confirmation would come from a blinded trial where actual patients with confirmed STIs interact with Otiz and their transcripts are graded by venereologists against human-counselor transcripts.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a domain-specialized LLM conversational agent, Otiz, can reproduce the counseling function of a venereologist for common genital conditions: it suggested the correct diagnosis or differentials, gave factually correct information, wrote in language lay users could follow, and responded with emotional warmth. The evidence is a table of mean numerical rating scores: diagnostic accuracy 4.1–4.7 across the four STIs, overall accuracy 4.3–4.6, comprehensibility 4.2–4.4, empathy 4.5–4.8, and correctness of information 5.0 with zero standard deviation in every disease class. Non-STI diagnostic scores were lower (2.9–3.5), with the STI/non-STI difference statistically significant ($p=0.038$), and relevance was the weak dimension, which the authors attribute to redundant responses. Inter-observer agreement was strong, with only 12.7% of paired ratings differing by more than one point.
Load-bearing premise
The central claim rests on the assumption that scores given by venereologists role-playing patients on a 0–5 rating scale reflect how well the chatbot would counsel real patients in real settings.
Editorial extensions
If this is right
- If the scores generalize, Otiz-type chatbots could handle first-line STI counseling and post-visit follow-up in resource-limited clinics, freeing specialists for complex cases.
- The consistently perfect correctness score implies a tool that does not add misinformation, addressing a key safety concern for patient-facing medical AI.
- The lower relevance scores indicate that the next concrete engineering target is response concision, reducing redundancy while keeping coverage.
- The weaker non-STI diagnostic performance suggests separate modules for non-infectious genital diseases are needed before broader genital-health use.
- Because the authors position Otiz as supplemental rather than replacement, the practical deployment outcome is a triage and counseling aid, not autonomous diagnosis.
Reading between the lines
- A natural extension is a blinded randomized comparison of Otiz against a generic LLM without the STI-specific modules and against human counselors, using the same vignettes, to isolate the contribution of the DFA and prompt architecture from the base model's abilities.
- The uniform 5.0 correctness score invites a ceiling-effect check: future evaluations should include prompts with deliberately planted near-miss information to verify the grader scale can detect errors.
- The DFA-plus-LLM hybrid suggests a general pattern for medical chatbots: use deterministic state control for safety-critical conversational flow and LLM generation for language, which could transfer to other stigmatized or sensitive health domains.
- Real-world validity hinges on whether actual patients, who are less medically articulate than venereologist actors, phrase symptoms differently; prompts derived from real patient transcripts would be a stronger test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes Otiz, a GPT-4-0613-based multi-agent chatbot for STI counseling, and reports an evaluation in which 23 venereologists role-playing patients rated 60 interactions (30 prompts, each rated by two raters) across four STIs and two non-STIs on six criteria. Mean scores were high for diagnostic accuracy, overall accuracy, correctness of information, comprehensibility, and empathy, while relevance scores were lower (2.9-3.6). The authors conclude that chatbots like Otiz can provide accurate, correct, discrete, and empathetic STI information and could reduce burden on healthcare systems.
Significance. If the result holds, this is a useful proof-of-concept for a disease-specific conversational agent in sexual health, a domain with stigma and access barriers. Strengths of the study include independent double ratings of each prompt, inclusion of non-STI conditions as a contrast, and transparent disclosure of the developers' affiliation. However, the lack of any baseline comparator, the role-play setting without real patients, and the suspicious ceiling effect in the correctness-of-information scores mean the current evidence supports only a narrow claim about simulated interactions, not the broad conclusion stated in the abstract and discussion.
major comments (4)
- [Methods III and Table 1] The central pillar of the conclusion, 'correctness of information', is reported as 5.0 with zero standard deviation across all six disease classes and all 60 evaluations. Such unanimity is implausible for a genuine quality measure and strongly suggests a ceiling effect, an anchoring effect, or a rubric in which a score of 5 meant only 'no explicitly dangerous misinformation' rather than 'complete and appropriate information'. Without a discriminating rubric, raw transcripts, or a coded analysis of factual errors, this score cannot bear the weight of the claim that Otiz provides 'accurate, correct' information. Please re-analyze the interactions or revise the conclusion to reflect the actual level of evidence.
- [Methods IV and Table 1] The comparison between STI and non-STI diagnostic accuracy uses a Wilcoxon signed-rank test, described as paired non-parametric data. The study design pairs two raters per prompt, not STI versus non-STI conditions; there are 4 STI classes and 2 non-STI classes, so the pairing structure for a signed-rank test is unclear from the methods. Please specify the unit of analysis (per prompt, per evaluation, or per condition), justify the pairing, and report effect sizes and confidence intervals. The current p=0.038 may be based on an invalid or arbitrary pairing.
- [Discussion and Conclusion] The conclusion that 'AI conversational agents like Otiz can provide accurate, correct... information' overgeneralizes because no baseline or comparison arm is included. The high scores could reflect leniency by evaluators who knew they were rating a specific product, the relative simplicity of the initiating prompts, or the absence of challenging edge cases. The authors themselves acknowledge in the Limitations section that the role-play evaluation may not capture real-world diversity and that blinded trials against human counselors and non-specific chatbots are needed; this undermines the categorical tone of the abstract conclusion. The conclusion should be limited to something like 'in simulated interactions with venereologist actors, Otiz received high ratings on several quality dimensions, but relevance was lower and real-world accuracy remains unverified.'
- [Results and Table 1] Inter-observer agreement is reported as '19 out of 150 pairs (12.7%)', but the design has 30 prompts × 6 criteria = 180 paired ratings. The denominator of 150 is unexplained and should be reconciled (for example, if one criterion was excluded, say so). Also, the text refers to 'discharge/proctitis' while Table 1 lists 'Gonorrhea/Chlamydia/UTI'; the terminology should be consistent.
minor comments (6)
- [Methods III] Please state explicitly whether the two venereologists who designed the prompts are the same individuals as any of the 23 evaluators or are authors on the paper; this information is relevant for assessing potential bias in prompt selection and scoring.
- [Abstract] The abstract reports 'correctness of information (5.0)' without a standard deviation or confidence interval, which conceals the ceiling effect; please include a measure of dispersion or a note that all ratings were identical.
- [Discussion] The claim that this is the 'first proof-of-concept of an STI conversational agent in the world' is contradicted by the manuscript's own reference 13 (SHIHbot) and other sexual-health chatbots; please revise to a more precise statement about the specific design or evaluation approach.
- [Table 1] Please add the number of evaluations underlying each cell (n=10 per disease class, n=60 total) and clarify whether the reported values are means across all evaluations in that class.
- [Discussion] The statement that Otiz 'can alleviate the burden on healthcare systems' is speculative; no data on provider time, patient outcomes, or cost are presented. Please remove or explicitly label this as a hypothesis.
- [References] Reference 22 is a vendor blog about NHS-recommended apps; consider replacing it with a peer-reviewed source describing regulatory approvals and evidence for mental-health chatbots.
Circularity Check
No significant circularity: the evaluation is an external expert rating study, and self-citations are background only.
full rationale
The paper's central claim (Otiz can provide accurate, correct, comprehensible, and empathetic STI information) rests on scores assigned by 23 independently recruited venereologists who role-played patients and rated six criteria on a 0-5 scale. These scores are observations, not fitted parameters or quantities derived from the chatbot's construction; there is no equation in which a reported score reduces to a training target or to an input used to build Otiz. The three self-citations (refs 2, 8, 12) appear in the Introduction to support background statements about STI burden, AI in STI care, and computer-vision research; none is load-bearing for the evaluation or conclusion. The references to DFA principles are supported by external textbooks and reviews (refs 14-15), not by an author-derived uniqueness theorem. The Limitations section explicitly concedes that actor-based role-play 'may not fully capture the diversity of user needs and preferences in real-world settings,' and the 5.0/0-SD correctness score is plausibly a ceiling or leniency artifact; however, a ceiling effect is a measurement-validity concern, not circularity, because the score is an independent observation rather than a defined consequence of the rating rubric. No claim in the paper is equivalent by construction to its input, and no prediction is statistically forced by a fitted value. The analysis is therefore self-contained with respect to the circularity question, even though the clinical validity of the scores can be questioned for other reasons.
Assumptions & free parameters
assumptions (4)
- domain assumption Expert venereologist ratings on a 6-point NRS are a valid measure of chatbot counseling quality and medical accuracy.
- domain assumption The 30 initiating prompts written by two expert venereologists are representative of real patient presentations for the six conditions.
- domain assumption Venereologists acting as patients can reproduce the conversational behavior and needs of real patients.
- domain assumption GPT-4's base medical knowledge is reliable enough that any factual errors would be caught by the raters.
Cite this review
Pith. "Pith review of Performance of a large language model-Artificial Intelligence based chatbot for counseling patients with sexually transmitted infections and genital diseases." pith.science (2026). https://pith.science/paper/E6WZPZ67
@misc{pith2026241212166,
author = {Pith},
title = {Pith review of: Performance of a large language model-Artificial Intelligence based chatbot for counseling patients with sexually transmitted infections and genital diseases},
year = {2026},
howpublished = {\url{https://pith.science/paper/E6WZPZ67}},
note = {Machine review of arXiv:2412.12166}
}
read the original abstract
Introduction: Global burden of sexually transmitted infections (STIs) is rising out of proportion to specialists. Current chatbots like ChatGPT are not tailored for handling STI-related concerns out of the box. We developed Otiz, an Artificial Intelligence-based (AI-based) chatbot platform designed specifically for STI detection and counseling, and assessed its performance. Methods: Otiz employs a multi-agent system architecture based on GPT4-0613, leveraging large language model (LLM) and Deterministic Finite Automaton principles to provide contextually relevant, medically accurate, and empathetic responses. Its components include modules for general STI information, emotional recognition, Acute Stress Disorder detection, and psychotherapy. A question suggestion agent operates in parallel. Four STIs (anogenital warts, herpes, syphilis, urethritis/cervicitis) and 2 non-STIs (candidiasis, penile cancer) were evaluated using prompts mimicking patient language. Each prompt was independently graded by two venereologists conversing with Otiz as patient actors on 6 criteria using Numerical Rating Scale ranging from 0 (poor) to 5 (excellent). Results: Twenty-three venereologists did 60 evaluations of 30 prompts. Across STIs, Otiz scored highly on diagnostic accuracy (4.1-4.7), overall accuracy (4.3-4.6), correctness of information (5.0), comprehensibility (4.2-4.4), and empathy (4.5-4.8). However, relevance scores were lower (2.9-3.6), suggesting some redundancy. Diagnostic scores for non-STIs were lower (p=0.038). Inter-observer agreement was strong, with differences greater than 1 point occurring in only 12.7% of paired evaluations. Conclusions: AI conversational agents like Otiz can provide accurate, correct, discrete, non-judgmental, readily accessible and easily understandable STI-related information in an empathetic manner, and can alleviate the burden on healthcare systems.
Reference graph
Works this paper leans on
-
[2]
Emotional recognition module- activated post-diagnosis, it assesses emotional states through text analysis, identifies emotions ranging from anxiety to relief, and adapts subsequent responses for empathetic care. It utilizes sentiment analysis techniques to identify the user's mood and adapt the chatbot's responses accordingly.17 3. Acute Stress Disorder ...
work page 2010
-
[27]
Conversational agents in healthcare: a systematic review
Laranjo L, Dunn AG, Tong HL, et al. Conversational agents in healthcare: a systematic review. J Am Med Inform Assoc 2018;25(9):1248–58. 28. Kocaballi AB, Berkovsky S, Quiroz JC, et al. The Personalization of Conversational Agents in Health Care: Systematic Review. J Med Internet Res 2019;21(11):e15360. 29. Bickmore TW, Trinh H, Olafsson S, et al. Patient ...
arXiv 2018
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.