{"id":"ffe425c3-bed0-41e5-ac8a-7d783289c706","arxiv_id":"2412.12538","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"The authors report a scalable AI-patient-actor benchmark and claim their own chatbot August reaches 81.8% top-one diagnostic accuracy on 400 vignettes, but the evaluation is self-run and the comparison baselines are undocumented.","lead":"This paper introduces a benchmark that uses AI-simulated patients to test how well a health chatbot, called August, diagnoses illnesses in conversation. It reports August at 81.8% first-choice diagnostic accuracy on 400 clinical vignettes, outperforming symptom checkers and several family doctors, though the evaluation was run by August's own developers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'significantly outperforming' claim rests on comparator scores in Section 5.1 with no stated measurement protocol; until Avey, Ada, and the clinicians are re-run under this benchmark's exact conditions, the headline accuracy comparison is unsupported.","rationale":"I read the manuscript in good faith and confirm its internal arithmetic: 327/400 is 81.75%, 340/400 is 85.0%, and 383/400 is 95.75%. The Appendix interaction is illustrative, and Section 7 is candid about the limits of vignettes and simulated patients. Those are real strengths. The load-bearing weakness, however, is exactly where the reader placed it: the comparator scores are asserted in Section 5.1 with no stated measurement protocol. The abstract's 'significantly outperforming traditional symptom checkers' and the Section 5.1 claims against clinicians and question counts all depend on comparability that the manuscript does not establish. Because Section 3.1 states only that the 400 vignettes were selected from Hammoud et al., the natural reading is that the baseline numbers were inherited from that prior work rather than regenerated under this paper's protocol. That matters because the protocol differs in at least three consequential ways: conversational versus form-based input, the AI patient actor's communication style, and the modified matching rubric. Each of those can change scores substantially, so the comparison is not valid on the evidence presented. The closed-loop design, with patient actors styled on August's own conversations and a vendor-modified rubric, makes the absence of an independent measurement more acute rather than less. The paper also reports no confidence intervals or significance tests, and the Dermatology row of 13 cases shows how noisy small subsamples are. None of this requires the 81.8% figure to be wrong; it requires the comparative claims to be treated as unverified. A same-protocol re-measurement of the comparators, or a published per-case mapping plus identical rubric and judge agreement, would be the decisive test. If such information were added, the verdict could move toward CONDITIONAL; as submitted, the reader's REJECT is the appropriate posture.","tokens_in":15503,"tokens_out":3007,"duration_ms":29321,"concrete_test":"Re-run Avey and Ada Health and the three family-medicine clinicians on the same 400-vignette subset, with the same AI patient-actor prompts, conversation interface, human judges, and Gilbert-based matching rubric used for August, and report per-system scores and inter-rater agreement. If comparators shift by more than 5 percentage points or the symptom-checker question count changes from 29, the headline comparisons in Section 5.1 and the Abstract are not valid. If the authors intend to rely on published scores from Hammoud et al., they must provide a per-vignette mapping showing exact overlap and rubric equivalence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 reports Avey at 67.5%, Ada at 54.2%, three clinicians at 49.7%/61.3%/72.5%, and Table 5 gives symptom checkers a mean of 29 questions, but the paper never states how these comparators were measured. Section 3.1 says the 400 vignettes came from Hammoud et al.; it does not say the comparators were run on those same 400 vignettes, through the same AI patient actors, against the same judges and the same modified Gilbert matching rubric. If these numbers are carried over from Hammoud/Gilbert rather than re-measured, then 'significantly outperforming' compares August's conversational top-1 score on this protocol to structured symptom-checker scores from a different vignette set and a different scoring convention. That breaks the central superiority claim regardless of whether 327/400 is correctly computed. The risk is heightened by the closed-loop elements: patient-actor communication style derived from August's own anonymized conversations (Section 3.2) and a vendor-modified rubric (Section 3.3). The paper's own Section 7 concedes patient actors do not fully replicate real patients, so an externally sourced baseline cannot silently inherit the conversational benchmark's assumptions. No confidence intervals or significance tests appear anywhere, and small-count specialties such as Dermatology (13 cases, 100%) show the noise floor.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a conversational benchmarking framework for health AI that combines 400 clinical vignettes from Hammoud et al. with LLM-based patient actors and human judges. It applies this framework to August, a health AI developed by the authors' own company, and reports a top-one diagnostic accuracy of 327/400 (81.8%), top-two accuracy of 340/400 (85.0%), specialist referral accuracy of 383/400 (95.8%), and a mean of 16 questions per consultation. The authors claim that August significantly outperforms two symptom checkers (Avey 67.5%, Ada 54.2%) and three family medicine physicians (49.7%, 61.3%, 72.5%). The paper also describes the benchmark design, outcome measures, strengths and limitations, and future plans for real-world validation.","tokens_in":15588,"tokens_out":6326,"duration_ms":57916,"significance":"If the methodology were independent and the comparator measurements were valid, the framework would address a genuine gap: there is no widely accepted, scalable method for evaluating conversational diagnostic AI. The use of 400 validated clinical vignettes, explicit matching criteria, and internally consistent tabulations are useful starting points. The paper also takes some care to report limitations, including the fidelity of AI patient actors and the absence of physical examination data. However, the empirical contribution as presented does not support the headline comparative claim: the comparator scores lack a described measurement protocol, no confidence intervals or significance tests are provided, and the benchmark has closed-loop features that favor the evaluated system. The framework itself is potentially valuable but is not yet independently reproducible from the manuscript.","major_comments":[{"comment":"The central claim that August 'significantly outperforms' Avey, Ada, and the three clinicians is unsupported because the manuscript never states how the comparator scores were measured. It is not reported whether Avey, Ada, and the clinicians were run on the same 400 vignettes, through the same AI patient actors, with the same conversational input format, the same judge instructions, or the same modified Gilbert matching rubric; the number of clinicians and their testing conditions are also absent. If these figures were taken from Hammoud et al. or Gilbert et al. rather than re-measured under this benchmark's protocol, then the comparison is invalid. The authors must either report the full comparator protocol, re-measure the comparators under the benchmark's exact conditions, or remove the comparative superiority claim.","section":"Section 5.1"},{"comment":"No confidence intervals or significance tests appear anywhere in the paper, so the word 'significantly' is not backed by statistical inference. With 400 cases, a binomial confidence interval should accompany the overall 81.8% and 85.0% estimates, and pairwise comparisons to each comparator should be tested. This is especially important because several specialty rows are based on very small counts (e.g., Dermatology 13/13, Hematology 16/23), where the noise floor is high and the reported percentages cannot support strong conclusions.","section":"Section 5.1 and Tables 1-2"},{"comment":"The benchmark is closed-loop in a way that threatens the comparative claim. The patient actors' communication style was explicitly derived from 'patterns from anonymized conversations between a health AI and actual users,' and the health AI in question appears to be August itself. This means the test distribution was shaped by August's own interaction patterns, potentially giving August an advantage over comparators that were not developed or tuned under the same conditions. The authors should construct patient actors independently of the evaluated system, or at minimum report a sensitivity analysis with patient actors not informed by August's conversations.","section":"Section 3.2"},{"comment":"The scoring rubric is a vendor-modified version of the Gilbert et al. criteria, but the paper does not report inter-rater reliability among the human judges, whether the judges were blinded to the identity of the system, or how the modified rubric affects comparability with previously published symptom-checker scores. Because August's outputs were judged under this modified rubric while external comparator scores may come from a different rubric, observed differences can be an artifact of the scoring criteria. The authors should provide inter-rater reliability statistics, describe judge blinding, and demonstrate that the modified rubric does not change the rank ordering of systems.","section":"Section 3.3"},{"comment":"The paper claims a 'standardized and scalable framework' and states an intention to make benchmarking accessible to others, but it does not provide the patient-actor prompts, the judge instructions, the vignette selection criteria, or the code used to run the benchmark. The data availability statement only directs readers to a contact email. Without these artifacts, the framework is not reproducible and the scalability claim cannot be independently assessed. The authors should release the benchmark artifacts or provide a detailed specification sufficient for independent reimplementation.","section":"Sections 3 and 9"}],"minor_comments":[{"comment":"References [50] and [63] are the same work (Hammoud et al., JMIR AI 2024) and should be merged into a single citation.","section":"References"},{"comment":"The sentence about symptom checkers relying on 'generative adversarial networks' is not accurate for the systems under discussion and is not supported by the cited literature; this should be corrected or removed.","section":"Section 1.1"},{"comment":"The 'Correct Specialty Identified' outcome lacks an explicit definition of what counts as the correct specialty and how the referral was judged; Table 4 also does not define 'Common' versus 'Less Common' incidence or report case counts for those subgroups.","section":"Section 4.2 and Table 4"},{"comment":"The abstract's unqualified 'diagnostic accuracy' and 'real-world impact' wording should be tempered, since Section 7 concedes that the patient actors do not capture incomplete or inaccurate patient input, linguistic diversity, and physical examination data; the results should be described as accuracy on simulated conversational vignettes.","section":"Abstract and Section 7"},{"comment":"There are minor language issues, including 'there exists no standardized and scalable framework' in the abstract, 'AIs training data' in Section 1, and 'we aim examine' in Section 4; a careful proofread is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The authors are employees of August AI and the manuscript evaluates their own product, yet no conflict-of-interest statement is included. If the journal's policy treats vendor self-evaluation as requiring independent replication, this may warrant a more severe editorial decision; otherwise, the requested revisions are extensive and should include re-measuring comparators under the benchmark protocol, adding statistical inference, and releasing benchmark artifacts. The absence of any comparator protocol is the single most serious issue and should be the primary focus of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The bottom line: the integrated loop—LLM patient actors, standardized vignettes, physician judges, strict matching criteria—is a sensible way to test conversational diagnostic AIs, and the paper describes it in more detail than most vendor reports. The internal arithmetic checks out (327/400 = 81.8%, 340/400 = 85.0%, 383/400 = 95.8%), and the limitations section is candid about the gap between vignettes and real patients. Credit where it's due: the design is reproducible in principle, and the patient-actor guidelines (simple language, most distressing symptom first, no volunteering) are concrete enough to actually build on.\n\nThe soft spot is the one the stress-test flags, and it's load-bearing. Section 5.1 reports Avey at 67.5%, Ada at 54.2%, and three clinicians at 49.7–72.5% without saying how those numbers were produced. Were the comparators run through the same 400 vignettes, the same patient actors, the same judges, and the same modified Gilbert rubric? If those scores are carried over from Hammoud et al. or elsewhere, then the comparison is apples to oranges: August's conversational top-1 on this protocol versus structured symptom-checker scores under a different protocol. The word 'significantly outperforming' is doing work that no significance test or confidence interval actually supports. That breaks the central claim as stated.\n\nThe closed-loop nature makes this worse. The patient actors were shaped by August's own anonymized user conversations, the rubric is the vendor's modification, and the system under test is the vendor's product. None of that makes the 81.8% figure implausible, but it means the test distribution is calibrated on the test-taker. Combined with small-count specialties like Dermatology (13 cases, 100%), the headline accuracy is best treated as unverified.\n\nA few other small things: no artifacts beyond a contact email, despite the 'reproducible framework' language, and the paper's own Section 7 concedes the patient actors don't fully replicate real patients.\n\nWho is this for? Anyone building or evaluating conversational diagnostic AIs, and anyone studying how to benchmark them responsibly. It deserves a serious referee rather than a desk reject, because the framework is worth engaging with and the flaws are fixable. My recommendation: send it to peer review but with a clear demand—either re-run the comparators under this exact protocol or drop the comparative superiority claims; also release the vignette subset, patient-actor prompts, judge agreement stats, and confidence intervals. As submitted, it's a conditional reject, not a definitive one.","headline":"The benchmark framework is a plausible, modest step, but the headline 'significantly outperforming' is unsupported because the comparator scores in Section 5.1 have no stated measurement protocol.","tokens_in":16350,"tokens_out":1691,"would_cite":false,"duration_ms":16928,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A conversational health AI reaches 81.8% top-one diagnostic accuracy by interviewing AI-simulated patients built from 400 validated clinical vignettes.","keywords":["diagnostic accuracy","health AI benchmarking","clinical vignettes","AI patient actors","conversational differential diagnosis","symptom checkers","large language models in medicine","referral accuracy"],"falsifier":"Re-run Avey and Ada on the same 400 vignettes through exactly the AI-patient-actor conversations and the same physician-judge rubric; if their top-one accuracies do not reproduce near 67.5% and 54.2%, or if August's margin shrinks, the outperformance claim is an artifact of mismatched baselines. A second check is to have trained humans act out the same vignettes with August: if top-one accuracy falls well below 81.8%, the LLM patient actors are inflating the score.","tokens_in":15084,"feed_emoji":"🩺","tokens_out":8044,"duration_ms":68620,"temperature":0.7,"pith_summary":"The paper argues that the right way to evaluate a conversational health AI is to let it conduct a live interview with a simulated patient, rather than feeding it a written case or a multiple-choice exam. To demonstrate this, the authors build AI patient actors from 400 validated clinical vignettes across 14 specialties, have their own health AI August converse with them, and ask physician judges to score the diagnoses August states aloud. On this protocol August reaches 81.8% top-one accuracy, 85.0% top-two accuracy, 95.8% correct specialist referrals, and a mean of 16 questions per consultation. The paper reports that these figures exceed published symptom-checker scores and the scores of three family physicians, and it offers the protocol as a reproducible, scalable benchmark for a field that has lacked one.","feed_headline":"A health AI hits 81.8% top-one on 400 conversational diagnoses","feed_subtitle":"AI patient actors on validated vignettes let August beat symptom checkers and three family doctors.","key_machinery":"The load-bearing object is the AI patient actor: an LLM initialized with a validated clinical vignette and prompted to speak in simple lay language, present the most distressing symptom first, volunteer nothing beyond what is asked, and stay strictly in character. This turns a static written case into a live, repeatable consultation that can be run at scale. The second mechanism is the matching rubric [41], adapted to be stricter by excluding near-matches, umbrella terms, and conditions with mere symptomatic overlap; human physician judges apply it to the diagnoses the health AI explicitly states in the conversation. Together these convert 'did the AI get the right answer' into 'did the AI elicit the right information and name the right condition from a genuine back-and-forth history.'","core_discovery":"The central discovery, on the paper's own terms, is that a health AI built for conversation can be held to a standardized diagnostic standard: when August interviews an LLM-based patient actor that is bound to the facts of a validated clinical vignette, a physician judge can reliably determine whether the AI's stated diagnosis matches the gold standard. Across 400 cases August's first diagnosis matched in 327 cases (81.8%) and one of its top two matched in 340 cases (85.0%). The paper also reports 95.8% accuracy in recommending the right specialist, and 47% fewer questions than symptom checkers (16 vs 29 on average). These numbers are presented as evidence both that August performs well and that the benchmark itself works as a scalable evaluation method for conversational diagnostic AI.","pith_inferences":["The comparative claim is only as strong as the comparators' protocol: the paper does not describe how the Avey, Ada, or physician baselines were measured, so a fair reading treats the outperformance as provisional until those systems are re-run under the same patient-actor and judge rules.","Because the authors' own patient actors answer only when asked and use grammatically correct English, the 81.8% figure is likely an upper bound for performance with real users; a small study with human standardized patients on the same vignettes would test this directly.","The judge step still depends on human physicians, so the framework's scalability claim will fully stand only once the proposed automated judge reproduces human-rubric scores on a held-out set of conversations.","If the vignette corpus and automated judge are released, the benchmark becomes a regression suite: future health AIs could be compared on identical conversations and identical scoring, which would give regulators and purchasers a common yardstick."],"forward_implications":["A conversational diagnostic AI can reach 81.8% top-one accuracy on validated vignettes while averaging 16 questions, suggesting that scripted short interviews can be both efficient and diagnostically productive.","Other health AI developers can run the same patient-actor protocol on the same vignette corpus, making head-to-head diagnostic comparisons possible without recruiting human standardized patients.","The choice of matching rubric materially changes reported accuracy; the paper's stricter exclusions make its 81.8% figure conservative relative to rubrics that count near-matches.","The 95.8% specialist-referral accuracy suggests the benchmark also captures triage quality, not just diagnosis naming, which matters for the AI's real-world role.","If the benchmark reflects real consultation skill, it implies a carefully prompted conversational AI can outperform practicing family physicians on vignette-based differential diagnosis in this specific setting."],"supporting_citations":[{"why":"Supplies the standard clinical vignette methodology and the 400-case corpus on which the entire benchmark is built.","marker":"[63]"},{"why":"Provides the diagnostic matching criteria that the paper adapts and tightens to decide when a predicted diagnosis counts as correct.","marker":"[41]"},{"why":"Establishes the vignette-based audit approach for symptom checkers and the argument that clinical utility drops beyond the top two diagnoses.","marker":"[36]"}],"fun_headline_variants":["Health AI August hits 81.8% on 400-vignette diagnostic test","Conversational AI diagnoses 327 of 400 cases correctly","AI patient actors put health chatbot to the test: 81.8% accurate","Benchmarking health AI: August tops symptom checkers at 81.8%","Scalable diagnostic benchmark: AI chatbot scores 81.8% on 400 cases"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparison assumes that the symptom-checker and doctor scores were produced under the same conversational protocol and the same matching rules; the paper describes August's protocol in detail but gives no such description for the baselines.","fun_headline_variants_meta":{"raw":{"variants":["Health AI August hits 81.8% on 400-vignette diagnostic test","Conversational AI diagnoses 327 of 400 cases correctly","AI patient actors put health chatbot to the test: 81.8% accurate","Benchmarking health AI: August tops symptom checkers at 81.8%","Scalable diagnostic benchmark: AI chatbot scores 81.8% on 400 cases"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000642,"raw_usage":{"total_tokens":2940,"prompt_tokens":917,"completion_tokens":2023,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1919}},"tokens_in":533,"tokens_out":2023,"duration_ms":11881,"temperature":1.0,"reasoning_tokens":1919,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:59:48.179879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run Avey and Ada on the same 400 vignettes through exactly the AI-patient-actor conversations and the same physician-judge rubric; if their top-one accuracies do not reproduce near 67.5% and 54.2%, or if August's margin shrinks, the outperformance claim is an artifact of mismatched baselines. A second check is to have trained humans act out the same vignettes with August: if top-one accuracy falls well below 81.8%, the LLM patient actors are inflating the score.","supporting_citations":[{"cited_title":"Eval- uating the Diagnostic Performance of Symp- tom Checkers: Clinical Vignette Study","cited_arxiv_id":null,"evidence_quote":"Supplies the standard clinical vignette methodology and the 400-case corpus on which the entire benchmark is built."}],"review_version":1}