{"id":"582f734c-059e-4529-88a3-82678bda8a43","arxiv_id":"2507.22902","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"In a retrospective sample of 500 urgent-care telehealth visits, a proprietary AI doctor matched clinicians' top diagnosis 81% of the time and treatment plans 99.2% of the time, but the design cannot support claims of comparable clinical decision-making.","lead":"This preprint compares a proprietary multi-agent AI system, Doctronic, against board-certified clinicians on 500 urgent-care telehealth visits. It reports 81% diagnostic agreement and 99.2% treatment-plan alignment, but the study measured concordance rather than clinical outcomes and was run by the system's owners.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99.2% treatment-plan 'safety' figure rests on a compatibility rubric that counts omitted workups as consistent, so the safety conclusion does not follow.","rationale":"The reader's verdict already rejects the paper on the grounds that concordance is not accuracy and that clinician anchoring makes concordance suspect. I agree that those are real, but the more internal and decisive problem is that the headline safety metric is defined so permissively that it cannot support the safety claim even as concordance. The LLM-judge prompt in Appendix 2 explicitly labels as consistent any pair in which one plan is a superset of the other or one plan is a very brief version of the other. That means the 99.2% figure can be high even when the AI adds unnecessary testing or omits needed testing, as long as the core treatment 'approach' overlaps. The paper's own low-similarity Case 3 illustrates this: the expert judged the AI plan superior, but the LLM judge had called the treatment plans not concordant, so the metric is not a safety validator. The discordant-case human review cannot fix this because it only looked at cases where the top diagnosis differed; the overwhelming majority of encounters labeled safe were never checked for plan-level safety. The advertised conclusion 'autonomously and safely assess' requires a plan-level correctness standard, not a pairwise compatibility standard. This does not require assuming bad faith; the limitations section is candid that the study measured agreement, not accuracy. But the abstract and conclusion draw a stronger inference than the measurement supports. A blinded, guideline-based safety re-rating of a sample of the compatible pairs would settle the matter. If the compatible plans are almost all genuinely safe and appropriate, the treatment-plan finding could be rehabilitated; if not, the central safety claim fails. My reading therefore leaves the reader's REJECT verdict unchanged.","tokens_in":22723,"tokens_out":4543,"duration_ms":50067,"concrete_test":"Select a random sample of 100 encounter pairs from the 496 pairs the LLM judge labeled 'treatment plan compatible.' Have two independent board-certified clinicians, blinded to AI/human origin and to the compatibility label, score each AI plan against condition-specific guidelines (e.g., IDSA/CDC/AAFP) for safety and appropriateness, using a pre-registered binary outcome (safe vs unsafe/inappropriate). Compute the proportion of plans rated unsafe and require, say, at most 5% for the safety claim to stand. If the 99.2% figure is inflated by the permissive superset/terse-plan rubric, this sample will reveal a higher unsafe rate, directly testing whether treatment-plan 'compatibility' supports 'autonomously and safely.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing problem is in the definition of the primary safety endpoint. Appendix 2, Prompt 3 explicitly counts two treatment plans as 'clinically consistent' if one plan is a superset of the other ('one plan includes all key elements of the other plan plus additional elements') or if one plan is very terse but both 'provide similar treatment approaches' (criterion 6). It also counts a plan that skips confirmatory testing as consistent if the final treatment approach is the same (criterion 2). Under this rubric, an AI plan that orders extensive labs, imaging, and referrals will be 'compatible' with a clinician plan that orders only supportive care, and a clinician note that says 'watch for now' is 'compatible' with a broader AI workup. Thus the reported 99.2% treatment-plan alignment is not evidence that the AI's plan is safe, guideline-concordant, or even substantively similar; it is evidence only that the two plans are not overtly contradictory. The human expert review of discordant pairs covered only the 97 top-diagnosis discordances, not the 496 pairs classified as compatible, so compatible-but-unsafe plans were never examined. The conclusion that the system 'can autonomously and safely assess and provide an appropriately documented treatment plan' therefore rests on a metric that cannot detect the omission of indicated tests, red-flag workups, or contraindicated-but-not-contradictory management.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a retrospective comparison of an autonomous multi-agent LLM system (Doctronic) against board-certified clinicians on 500 consecutive urgent-care telehealth encounters. The primary endpoints are top-1 and top-4 diagnostic concordance, treatment-plan alignment, and safety metrics, adjudicated by an LLM judge and by human expert review. The authors report 81% top-1 concordance, 95.4% top-4 concordance, 99.2% treatment-plan alignment, and zero clinical hallucinations, and conclude that the AI system can autonomously and safely assess and document treatment plans with consistency matching board-certified clinicians.","tokens_in":22957,"tokens_out":3921,"duration_ms":39403,"significance":"If the claims were supported, this would be a landmark result: the first demonstration that an autonomous LLM-based system matches clinicians in real-world urgent-care decision-making. The paper has notable strengths: it uses 500 consecutive real-world encounters rather than vignettes, includes the verbatim adjudication prompts in an appendix, and provides reproducible similarity metrics. However, the central claim of comparable clinical decision-making is not supported by the data as presented. The clinician comparator was exposed to the AI note before the encounter, the treatment-plan rubric is permissive to the point of counting omitted workups as consistent, the human expert review was unblinded and single-reviewer, and the LLM judge was not quantitatively validated. These issues are load-bearing for the paper's conclusions, so the result, while potentially important, is not established by this study.","major_comments":[{"comment":"The study design states that \"Clinicians were given a copy of the AI-generated documentation before their telehealth visit.\" This creates a direct anchoring effect: the human comparator is not an independent assessment, because the clinician has already seen Doctronic's diagnosis and plan. The 81% and 99.2% concordance figures therefore measure agreement conditional on exposure to the AI output, not agreement between two independent decision-makers. The Limitations section acknowledges this anchoring effect, but the abstract and conclusion nevertheless assert a \"consistency matching that of board-certified clinicians.\" This is a load-bearing design issue that cannot be remedied by reanalysis of the present data.","section":"Methods, Clinicians"},{"comment":"The treatment-plan alignment endpoint is defined by rubrics that make \"clinically consistent\" extremely permissive. Criterion 2 counts as consistent a plan that omits a confirmatory test while the other does not, as long as the final treatment agrees; criterion 5 counts a plan that is a superset of the other (\"more comprehensive but includes the same core approach\"); criterion 6 counts a one-sentence plan as consistent if it provides a \"similar treatment approach.\" Under these criteria, a plan that omits diagnostic testing, red-flag workups, or indicated monitoring can be classified as aligned. The Results section states that the 99.2% alignment represents plans \"judged to be clinically compatible and guideline-concordant,\" but the rubric cannot detect the omission of indicated tests or non-contradictory but unsafe management. The primary safety endpoint therefore does not measure safety or guideline concordance.","section":"Appendix 2, Prompt 3"},{"comment":"The expert review of discordant pairs was performed by a single board-certified physician who was not blinded to the origin of the notes; the authors state that blinding \"proved impractical\" because AI notes were easily recognizable. No inter-rater reliability or second reviewer is reported. Moreover, this review covered only the 97 pairs in which the LLM judge found a top-1 diagnosis mismatch. The 496 pairs classified as treatment-aligned were never examined for omitted workups or other unsafe-but-compatible plans. Consequently, the claims that AI was superior in 36.1% of discordant cases and that no harmful errors occurred rest on an unblinded, single-reviewer assessment of a non-representative subset.","section":"Methods, Evaluation by Human Experts"},{"comment":"The LLM-judge protocol is described as \"developed and validated on a set of paired notes\" with human judges, but no validation statistics are reported. There is no human-LLM agreement rate, no sample size for the validation set, no inter-rater reliability for the human judges, and no analysis of the multi-run consensus beyond an unspecified \"consensus multi-run\" protocol. Since the LLM judge is the sole adjudicator for the primary concordance and safety endpoints across all 500 pairs, the absence of quantitative validation of this instrument is a load-bearing gap.","section":"Methods, Blinded LLM-Judge Prompts"},{"comment":"The Limitations section explicitly states that \"the findings show agreement, not accuracy\" and that \"Ground truth based on follow-up outcomes was not considered.\" However, the Conclusion asserts that Doctronic \"can autonomously and safely assess and provide an appropriately documented treatment plan\" and the abstract claims \"comparable clinical decision-making to human providers.\" Agreement without ground truth cannot establish safety, accuracy, or comparable decision-making; the conclusion overstates what the study can support. This internal inconsistency between the stated limitations and the stated conclusions needs to be resolved.","section":"Limitations and Conclusion"}],"minor_comments":[{"comment":"The abstract says the endpoints were \"assessed by blinded LLM-based adjudication and expert human review,\" but the human expert review was explicitly unblinded in the Methods.","section":"Abstract"},{"comment":"Table 1 lists \"Influenza 20\" and \"Acute bronchitis 17\" as the most common conditions, while Appendix 3 lists frequencies of 22 and 18 for the same categories; the source of this discrepancy should be clarified.","section":"Table 1 and Appendix 3"},{"comment":"The text has a typo, \"biomedical co rpora,\" which should read \"biomedical corpora.\"","section":"Methods, Similarity and Style Metrics"},{"comment":"The Jaccard index is reported as \"0.087 ± 0.0450.\" with a trailing period after a number; this is a formatting error.","section":"Table 3"},{"comment":"The phrase \"evaluated on the bias of treatment plan consistency\" should read \"evaluated on the basis of treatment plan consistency.\"","section":"Discussion"}],"recommendation":"reject","confidential_remarks":"The study is entirely authored by equity owners of the commercial system under evaluation, Doctronic. The conflict is disclosed, but the combination of a vendor-authored rubric, an anchored clinician comparator, an unblinded single-expert review, and an unvalidated LLM judge leaves the central claim without independent support. The transparent reporting of prompts and data is commendable, but the design flaws are such that a simple revision would not fix them; the conclusions would require a new study with clinicians who do not receive the AI note, a validated outcome-based ground truth, and blinded adjudication with inter-rater reliability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first published evaluation of an end-to-end autonomous multi-agent LLM system running actual telehealth urgent-care encounters rather than vignettes, and the 500-case real-world dataset is worth something. The authors also do a few things right: they publish the full judge prompts, they show paired SOAP notes in the appendix, and they state plainly in the limitations that the findings show agreement, not accuracy, and that giving clinicians the AI note first creates an anchoring bias. All disclosures are up front, since every author is a Doctronic equity holder. Any serious discussion of clinical AI evaluation should be able to point to this paper as the current state of the art in vendor self-evaluation, for better and for worse.\n\nThe soft spots are not small. The primary safety endpoint, treatment-plan consistency at 99.2%, is defined by Prompt 3 in Appendix 2, and the rubric explicitly counts a plan that omits confirmatory testing as compatible if the final treatment approach is the same, and a very terse 'watch for now' as consistent with a broad workup so long as the core approach overlaps. That means the 99.2% figure is really a measure of 'not overtly contradictory,' not a measure of safety or guideline-concordant management. An AI plan that skips a needed test, a red-flag workup, or a contraindicated but not contradictory medication would sail through as compatible. The human expert review only looked at the 97 top-diagnosis discordant pairs, so the 496 treatment-plan-compatible pairs were never screened for these misses. On top of that, the expert adjudication was single-reviewer, not blinded (the authors admit the AI notes were recognizable), and had no inter-rater check. The clinician comparator was anchored by reading the AI's note before the visit, so the 81% top-diagnosis concordance is arguably inflated by the study design itself.\n\nNone of this makes the paper worthless. As a benchmark of what an autonomous LLM agent does when given a real clinical conversation, the transcript-level data and the full prompt set are reusable. But the conclusion that the system 'can autonomously and safely assess... with a consistency matching that of board-certified clinicians' does not follow from the evidence. The evidence supports a much narrower claim: high agreement with clinicians under conditions designed to maximize agreement, with no detected hallucinations in this sample.\n\nWho gets value from this? Methodologists working on LLM-as-judge validation, and anyone teaching pitfalls in clinical AI benchmarking. It deserves peer review because the question is important and the dataset is unique; the reviewers should push for an unanchored clinician control, a treatment-plan rubric that penalizes omitted workups, and independent blinded adjudication with inter-rater reliability. My own read is that the central claim is not established, but the paper is credible enough as a serious evaluation attempt to take through review.","headline":"First real-world autonomous LLM telehealth evaluation, but the treatment-plan compatibility rubric makes the 'safety' headline unsupported.","tokens_in":23558,"tokens_out":3670,"would_cite":false,"duration_ms":38221,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A multi-agent LLM-based system autonomously ran 500 real urgent-care telehealth encounters, matching board-certified clinicians on the top diagnosis in 81% of cases and on treatment plans in 99.2%, with zero clinical hallucinations.","keywords":["autonomous AI doctor","large language models","multi-agent system","urgent care telehealth","diagnostic concordance","LLM-as-a-judge","clinical safety","hallucination"],"falsifier":"A prospective randomized study where clinicians evaluate the same patients without seeing the AI note, and where an independent panel adjudicates against actual follow-up outcomes, would settle the claim. If top-diagnosis concordance falls substantially once anchoring is removed, or if the AI's plans are associated with more return visits or adverse events than clinicians' plans, the conclusion that the AI matches board-certified clinicians in safety and accuracy would be falsified.","tokens_in":22502,"feed_emoji":"🩺","tokens_out":6196,"duration_ms":53331,"temperature":0.7,"pith_summary":"This paper reports a retrospective evaluation of Doctronic, a proprietary multi-agent LLM-based system that runs an entire urgent-care telehealth encounter—history taking, reasoning, diagnosis, and SOAP documentation—without a human in the loop. Comparing 500 consecutive real encounters against board-certified clinicians, the authors find the AI's top diagnosis matched the clinician's in 81% of cases, the treatment plans were judged compatible in 99.2%, and no hallucinated diagnoses or treatments were found. The authors' central claim is that an autonomous AI doctor can safely assess and document urgent-care patients with consistency comparable to human clinicians, a result that would matter because projected clinician shortages are large. The paper frames this as the first large-scale validation of an end-to-end autonomous LLM-based clinical system in a real-world setting.","feed_headline":"Autonomous AI doctor matches clinicians on 81% of diagnoses","feed_subtitle":"In 500 real telehealth visits, the AI also matched 99.2% of treatment plans, with zero hallucinations.","key_machinery":"The load-bearing object is Doctronic itself, a cloud-native system of more than 100 LLM-powered agents that mimic a care team: agents take a full history, synthesize the conversation, generate a differential diagnosis with at least four entries, and produce a SOAP note; clinicians could use that note during their own visit. The evaluation machinery is a blinded LLM-as-judge protocol built on four pre-tested prompts—top-4 concordance, top-1 concordance, treatment-plan compatibility, and a 0-10 Comparative Summary Score (CSS)—plus surface and embedding-based similarity metrics and a single board-certified physician's review of all discordant pairs. This machinery is what turns the raw encounter pairs into the reported concordance and safety numbers.","core_discovery":"On the paper's own terms, the discovery is that a multi-agent LLM architecture can carry out a complete urgent-care visit autonomously and produce decisions that align with board-certified clinicians: 81.0% top-diagnosis concordance (405/500), 95.4% top-4 concordance, 99.2% treatment-plan compatibility (496/500), and zero clinical hallucinations. In an expert review of the 97 discordant pairs, the AI was judged superior in 36.1%, the clinician superior in 9.3%, and the rest equivalent, ambiguous, or the same diagnosis documented with low specificity. The authors conclude that the system can autonomously and safely assess patients and provide an appropriately documented treatment plan, and they offer it as a potential answer to healthcare workforce shortages.","pith_inferences":["We infer that the 81% concordance is likely an upper bound: clinicians were given the AI note before their visit, so anchoring bias may have inflated agreement; a blinded replication would probably show lower raw concordance.","The 99.2% treatment-plan compatibility figure should be read against the paper's own compatibility criteria, which count different NSAIDs or broader work-ups as 'compatible'; outcome-based safety comparisons would be a stricter test.","If the zero-hallucination result is taken at face value, it is a property of this specific multi-agent architecture plus grounded conversation, not of LLMs generally; stress-testing with rare or contradictory presentations would clarify how much the architecture, rather than the evaluation, is responsible.","The study's text-only, English-language telehealth setting leaves open whether the same autonomy transfers to video examinations, non-English patients, or in-person care, which would need separate validation."],"forward_implications":["If the central claim holds, fully autonomous AI systems can perform end-to-end urgent care encounters, including history taking, reasoning, and documentation, without direct human-in-the-loop intervention.","Such systems could be deployed as first-line triage in after-hours, rural, or resource-constrained settings, shortening wait times and expanding access.","The high treatment-plan alignment suggests AI can produce guideline-concordant management plans in the large majority of urgent-care presentations.","The finding that many apparent diagnostic disagreements stem from low-specificity clinician documentation implies that benchmarking methods must use semantic adjudication rather than exact wording matching.","The authors' claim that major clinical errors can be reduced to near zero with appropriate grounding implies a design target for future clinical AI: architecture and grounding, not just model scale, determine safety."],"supporting_citations":[{"why":"Establishes the methodological standard for blinded expert-rated evaluation of LLM diagnostic accuracy that this study builds on.","marker":"[11]"},{"why":"Provides the prior real-world comparison of AI intake recommendations versus physician recommendations in virtual urgent care, the direct benchmark for this study's concordance rates.","marker":"[12]"},{"why":"Supplies the LLM-as-judge methodology that the blinded adjudication prompts rely on.","marker":"[13]"},{"why":"Defines assessment of AI-generated clinical notes and hallucination that the documentation-quality and safety metrics extend.","marker":"[14]"}],"fun_headline_variants":["Autonomous AI doctor matches MDs in 500-visit real-world trial","AI doctor hits 81% diagnosis match, zero hallucinations in real care","AI clinician outperforms human doctors in 36% of disagreements","First autonomous AI doctor validated in urgent care telehealth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that the AI is as safe and accurate as clinicians rests on treating agreement with those clinicians—who had read the AI's note before their own visit—as a proxy for clinical correctness, a premise the paper itself states as 'agreement, not accuracy.'","fun_headline_variants_meta":{"raw":{"variants":["Autonomous AI doctor matches MDs in 500-visit real-world trial","AI doctor hits 81% diagnosis match, zero hallucinations in real care","AI clinician outperforms human doctors in 36% of disagreements","First autonomous AI doctor validated in urgent care telehealth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000717,"raw_usage":{"total_tokens":3256,"prompt_tokens":1012,"completion_tokens":2244,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":2171}},"tokens_in":628,"tokens_out":2244,"duration_ms":22557,"temperature":1.0,"reasoning_tokens":2171,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:03:50.113977+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective randomized study where clinicians evaluate the same patients without seeing the AI note, and where an independent panel adjudicates against actual follow-up outcomes, would settle the claim. If top-diagnosis concordance falls substantially once anchoring is removed, or if the AI's plans are associated with more return visits or adverse events than clinicians' plans, the conclusion that the AI matches board-certified clinicians in safety and accuracy would be falsified.","supporting_citations":[{"cited_title":"Ann Intern Med 178:498-506, 2025","cited_arxiv_id":null,"evidence_quote":"Provides the prior real-world comparison of AI intake recommendations versus physician recommendations in virtual urgent care, the direct benchmark for this study's concordance rates."}],"review_version":1}