{"id":"d48b0ba5-0830-40bb-9730-62d9aac08411","arxiv_id":"2608.12590","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An agentic, clinician-editable evidence workflow for thyroid ultrasound matches or beats specialist baselines on segmentation and classification benchmarks and improves report consistency and speed in reader studies.","lead":"This paper presents ThyroidXAgent, an AI system that coordinates specialized tools for thyroid ultrasound diagnosis and keeps a case-level record of the evidence behind each finding. It reports strong benchmark results and shows time savings and consistency gains in reader studies, though private-cohort and small-sample checks temper the headline numbers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's headline AUROC (0.9466) is not supported as evidence of multicentre generalization once the private NHC-MISD-TUS cohort is considered: its pooled AUROC is 0.819 and several centres perform near chance (Table S11).","rationale":"The reader's weakest assumption identifies the same load-bearing issue: the private NHC-MISD-TUS pooled metrics are treated as evidence of multicentre robustness despite near-chance per-centre results. This is the most consequential gap because the abstract's generalization claim is what a clinician or reviewer would carry forward. I do not see a more fundamental flaw in the agentic architecture or the evaluation of the public benchmarks; the paper is otherwise detailed and transparent, with per-centre tables and bootstrap confidence intervals. The concern reinforces the existing CONDITIONAL verdict rather than changing it: the system may well be useful, but the abstract overstates the strength of the multicentre evidence, and the private-cohort AUROC should be reported prominently. A random-effects summary of Table S11 is a concrete, feasible check that would settle whether the observed heterogeneity is compatible with the pooled metric. If the check shows high heterogeneity or a materially lower summary AUROC, the paper should downgrade the generalization claim to 'performance varies substantially across centres' and the verdict would remain conditional (or move toward reject if the abstract is not revised). Since the reader already reached CONDITIONAL, no adjustment is needed.","tokens_in":55955,"tokens_out":5092,"duration_ms":53424,"concrete_test":"Fit a random-effects meta-analysis (or a hierarchical logistic model) to the centre-level AUROC values in Table S11, rather than pooling all cases; report the summary AUROC, 95% CI, and I-squared heterogeneity statistic. As a sensitivity check, recompute the pooled AUROC after excluding centres with fewer than 50 classification cases (THYB_S_YN05, THYB_S_JL04, THYB_S_BJ09, and similar). If the summary estimate falls materially below 0.819 or I-squared exceeds roughly 80%, the abstract should state that private-cohort performance is heterogeneous across centres and should not be summarized by a single 'multicentre' AUROC.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that ThyroidXAgent generalizes across heterogeneous datasets. The only truly multicentre external validation is the private NHC-MISD-TUS cohort, where Fig. 2h reports an overall classification AUROC of 0.819. The abstract's headline AUROC of 0.9466 is the unweighted mean over five public test sets and omits this private-cohort number entirely. Table S11 shows that the pooled 0.819 conceals substantial per-centre variation: THYB_S_ZJ24 has AUROC 0.590 (N=269), THYB_S_BJ09 0.586 (N=20), THYB_S_JL04 0.407 (N=12), THYB_S_YN05 0.500 (N=12), and THYB_S_NM02 0.653 (N=49), while larger centres such as THYB_S_SH01, THYB_S_EN04, and THYB_S_JX06 reach 0.85-0.90. Pooling across cases treats this heterogeneity as sampling noise, but the spread is large enough to suggest systematic site effects. The abstract's generalization claim therefore depends on the assumption that a pooled AUROC over unequally sized centres is clinically meaningful; if per-centre performance varies this widely, 'multicentre robustness' is not established. The small centres have wide confidence intervals, so this is not proof of failure, but the pooled number alone is insufficient to support the headline claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ThyroidXAgent, an agentic AI system that coordinates specialized tools for thyroid ultrasound nodule segmentation, benign-malignant classification, malignant-lesion stratification, and report generation, storing intermediate outputs in an auditable case-level evidence record. The system was trained on a newly assembled OpenThyroidDB resource and evaluated on held-out public test sets plus a private 8,721-case, 35-centre NHC-MISD-TUS cohort. The main reported results are a mean Dice of 87.21% across seven segmentation test sets, a mean AUROC of 0.9466 across five public classification test sets, an AUROC of 0.819 on the private multicentre cohort, and a reader study showing improved physician classification accuracy, increased report diagnostic consistency (70.3% to 86.2%), and reduced segmentation/reporting time. The paper also introduces ThyClinScore, a lesion-level clinical semantic metric for report evaluation, and reports that it correlates better with an LLM judge than overlap-based metrics. The central claim is that an evidence-centred, clinician-correctable agentic workflow outperforms isolated predictive models and supports auditable clinical reporting.","tokens_in":56238,"tokens_out":4506,"duration_ms":45975,"significance":"If the results hold, the paper makes a useful contribution by reframing thyroid ultrasound AI around a case-level evidence record rather than isolated prediction endpoints. The manuscript is strong in its scale and methodology: it provides bootstrap confidence intervals throughout, uses genuinely held-out test sets, includes an independent private multicentre cohort, reports reader studies in a crossover design, and makes code and data publicly available. The strongest conceptual value is the demonstration that an agent can route cases to specialized tools and expose intermediate evidence for clinician correction, which is a plausible path toward auditable clinical AI. The main caveat is that the headline generalization claims are stronger than the evidence: the private-cohort AUROC of 0.819 is materially lower than the public-set mean of 0.9466, and per-centre variation is large; in addition, one of the 'external' public sets shares an institution with training data. The validation of ThyClinScore against an LLM judge rather than human expert scores is also indirect. These issues do not invalidate the paper, but they require careful revision of the claims.","major_comments":[{"comment":"The abstract's headline 'mean AUROC of 0.9466' is the unweighted mean over five public test sets and omits the private NHC-MISD-TUS result, which is 0.819 (Table S11, overall row). More importantly, the pooled private-cohort AUROC conceals very large per-centre heterogeneity: THYB_S_ZJ24 (N=269) has AUROC 0.590, THYB_S_BJ09 (N=20) 0.586, THYB_S_JL04 (N=12) 0.407, and THYB_S_YN05 (N=12) 0.500. Since the paper's key generalization claim rests on 'heterogeneous datasets' and 'multicentre robustness', the pooled number alone is insufficient. The authors should report a centre-level analysis, for example a random-effects estimate with per-centre CIs, and explicitly discuss which centres are near chance and why. They should also revise the abstract to include the private-cohort AUROC or temper the generalization language.","section":"§2.2, Fig. 2h, Table S11"},{"comment":"ZJH-8K is described as an independent external test set, but Table S1 states that it was collected at Zhujiang Hospital, Southern Medical University, the same institution that contributed the TN3K training data and the SMU-HMC report-training data. Institutional overlap with training data weakens the 'external' characterization and makes the cross-dataset generalization claim less strong than presented. The paper should explicitly acknowledge this overlap and re-evaluate the cross-dataset conclusions with ZJH-8K excluded or flagged as a same-institution holdout.","section":"§2.1, Table S1"},{"comment":"ThyClinScore is a central evaluation metric for report quality, but its parameters (λ, η, τ, α_s, β_f, the matching threshold, and the component weights w_q and u_r in §4.4) are described as predefined without any sensitivity analysis. The only external validation is a Pearson correlation with a location-aware LLM judge (Fig. 4c, r = 0.696), which is itself not a clinical ground truth. Since the paper uses ThyClinScore to claim that evidence-grounded assembly outperforms LLM baselines, the authors should demonstrate that the ordering of systems under ThyClinScore is stable across reasonable parameter choices, or validate the metric against human expert ratings on a sample of reports.","section":"§4.4, Fig. 4c, §2.5"},{"comment":"The reader study that supports the workflow-level claims of improved diagnostic consistency (70.3% to 86.2%) and reduced reporting time (2.5 to 1.8 min/case) involves only two physicians and 145 cases, and no confidence intervals or significance tests are reported for the consistency change. The crossover design is a strength, but the manuscript should provide at least bootstrap CIs or a paired test for the consistency and time endpoints, and should dampen the strength of the claim given the small number of readers.","section":"§2.6, Fig. 6"}],"minor_comments":[{"comment":"The abstract contains garbled characters such as 'benign ⚶malignant', and the introduction has a duplicated verb in 'ThyroidXAgent improved produced more clinically consistent reports'. These should be corrected.","section":"Abstract and Introduction"},{"comment":"The text states that ThyroidXAgent obtained the highest Dice on six of seven test sets and the lowest HD95 on all test sets, which is accurate. However, the PKTN result (82.99% Dice versus MedSAM2's 83.46%) is a clear exception and should be acknowledged in the main text when the 'six of seven' claim is made.","section":"§2.2 and Table S3"},{"comment":"The Code availability section says the source code is publicly available on GitHub but also says 'The website will be made available after acceptance', which is contradictory. Please clarify which resources are currently accessible.","section":"§5 and §6"},{"comment":"There are several typographical issues, such as 'Dice coeﬀicient' in Table S3 and 'Confilts'/'confilt' in Section 9; these should be fixed in proof.","section":"Table S3 and elsewhere"}],"recommendation":"major_revision","confidential_remarks":"The paper is substantial and the empirical apparatus is unusually thorough for this area. The main concern for the editor is the gap between the abstract's generalization claim and the actual private-cohort numbers: a pooled AUROC of 0.819 over centres that range from 0.407 to 0.903 requires a more careful discussion of site effects before publication. The ZJH-8K institutional overlap is also worth probing, as it affects how 'external' the external validation really is. I would not reject the paper on these grounds; the core system and the evidence-record concept are sound and the reported public-set numbers are internally consistent. However, the authors need to either add the targeted analyses (centre-level pooling, ThyClinScore sensitivity, reader-study inference) or soften the claims accordingly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the short version: this is a serious systems paper worth your time, but read the body rather than the abstract. The headline AUROC of 0.9466 is the mean over five public test sets, one of which (ZJH-8K) comes from the same institution as a training cohort. The only truly multicentre external validation, the private NHC-MISD-TUS cohort, gets a pooled AUROC of 0.819, and Table S11 shows several small centres near chance (0.407 to 0.653). The stress-test note gets this right. The paper does report all of this openly in the body, so it is more abstract overselling than concealment, but the generalization claim in the abstract is not supported as written.\n\nWhat is genuinely new: OpenThyroidDB is a real resource that integrates public and institutional data across segmentation, classification, reporting, and stratification tasks, with newly contributed cohorts. Wrapping that in a clinician-interactive agent workflow with an auditable evidence record is a sensible and reasonably original integration, and the report-assembly approach using BM25 template retrieval plus slot filling is a practical counterpoint to unconstrained LLM generation. ThyClinScore is a useful proposal for clinical semantic evaluation, though its hand-chosen weights and validation against an LLM judge rather than clinical labels mean it should be treated as a promising metric, not a settled one.\n\nThe soft spots beyond the abstract: the reader study used two physicians, which the authors acknowledge; the per-centre AUROC spread in the private cohort is large enough that pooling alone does not establish robustness across centres; and the same-institution external set weakens the external claim. None of these are fatal if the claims are softened. The paper is honest about limitations, provides extensive bootstrap CIs, and appears to be reproducible in principle with released code and weights.\n\nWho is this for: anyone working on medical AI workflows, thyroid ultrasound analysis, or clinically auditable agent systems. It is a strong candidate for serious peer review, but the revision should fix the abstract, report the private cohort number prominently, and temper the multicentre robustness language. I would cite it for the dataset and the workflow architecture, and I would probably bring it to a reading group as a case study in how to (and how not to) headline a systems paper.","headline":"A solid medical-AI systems paper with a genuinely useful new dataset and workflow, but the abstract's headline AUROC hides the private multicentre cohort's 0.819 and several near-chance centres.","tokens_in":56884,"tokens_out":2230,"would_cite":true,"duration_ms":26789,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Treating thyroid ultrasound as an auditable, coordinated evidence workflow lifts segmentation, classification, and reporting consistency across multicentre tests.","keywords":["agentic AI","thyroid ultrasound","case-level evidence record","nodule segmentation","benign-malignant classification","report generation","clinician-in-the-loop","auditability"],"falsifier":"Look at Table S11: on the private NHC-MISD-TUS cohort, pooled AUROC is 0.819, but several centres sit near chance, including THYB_S_ZJ24 (n=269, AUROC 0.590), THYB_S_JL04 (n=12, AUROC 0.407) and THYB_S_YN05 (n=12, AUROC 0.500). If a validation protocol counted per-centre failures rather than pooled means, and a meaningful fraction of larger centres stayed below 0.6, the multicentre robustness claim would be falsified.","tokens_in":55714,"feed_emoji":"🩺","tokens_out":5646,"duration_ms":49752,"temperature":0.7,"pith_summary":"Thyroid ultrasound diagnosis is a chain of linked steps: locate a nodule, measure it, characterise it, decide on risk, and write a report. This paper argues that AI should coordinate that whole chain around a case-level evidence record instead of firing off isolated predictions. The central claim is that an agentic workflow which routes cases to specialised tools, consolidates masks, probabilities and radiomic descriptors, and stores everything for clinician correction outperforms both standalone models and end-to-end multimodal language models on heterogeneous data. Reported results include a mean Dice of 87.21% for nodule segmentation, a mean AUROC of 0.9466 for benign-malignant classification, and an increase in report diagnostic consistency from 70.3% to 86.2% with reduced physician time. If the approach holds, clinical AI for multi-step diagnostic tasks can be built to be inspectable and correctable rather than a final black-box answer.","feed_headline":"Agentic AI lifts thyroid report consistency to 86.2%","feed_subtitle":"A coordinated evidence record beats isolated models and cuts reporting time by 27%","key_machinery":"The central object is the case-level evidence record, a structured store of lesion masks, measurements, class probabilities, radiomic descriptors, SHAP explanations, confidence signals and report clauses for one examination. The mechanism that carries the argument is the agent-as-workflow-controller: an LLM router reads structured summaries and metadata, selects the most reliable expert output for each case, writes the evidence to the shared record, and later converts that evidence into report text through template retrieval, slot filling and clause combination. This design makes auditability a property of the workflow itself, because every output traces back to editable evidence and corrections propagate forward.","core_discovery":"The paper's discovery is that the unit of AI assistance should be the auditable case-level evidence record, not a single prediction endpoint. ThyroidXAgent plans each case, routes inputs to specialised segmentation, classification, radiomics and measurement tools, and stores masks, class probabilities, radiomic features, uncertainty signals and report clauses in a shared record. That record can be inspected, corrected by a clinician, and reused downstream: a corrected mask updates SHAP attributions and classification confidence, and the same evidence drives evidence-grounded report assembly. Across seven segmentation test sets and five classification test sets the workflow led on most metrics, and in reader studies physicians using the evidence achieved higher accuracy and consistency while taking less time. The claim is that this coordination-and-evidence architecture, rather than any single new model, is what carries the improvement.","pith_inferences":["Editorial inference: pooling per-centre numbers may hide clinically useless performance at individual sites; a deployment-ready system should report and monitor centre-specific AUROC and Dice, especially where the paper's private cohort shows values near chance.","Editorial inference: the template-bank approach suggests reports can be adapted to new hospitals or languages by swapping retrieval libraries, without retraining image tools; that is a testable extension.","Editorial inference: clinician corrections written back to the evidence store could become a source of continuous model improvement, though privacy, provenance and versioning governance would be required.","Editorial inference: the same orchestration pattern could be transferred to other structured ultrasound examinations, such as breast or abdominal imaging, where the evidence objects and report clauses differ but the workflow controller stays the same."],"forward_implications":["A clinician can correct an intermediate lesion mask and the downstream radiomics, classification and report update automatically instead of restarting the pipeline.","Report generation can be constrained by structured facts rather than free-form text, lowering the risk of hallucinated lesion locations or measurements.","The same segmentation-radiomics-classification tools are reusable for different clinical questions, such as lymph-node metastasis and subtype classification, with only the classifier and task instructions changed.","If the reported workload reductions replicate, AI-assisted review could shorten segmentation and reporting time without sacrificing segmentation quality.","The evidence record provides a natural audit trail for retrospective review of AI-supported diagnoses."],"supporting_citations":[{"why":"Supplies the TN3K segmentation dataset and a thyroid-region-prior segmentation method that anchors the training pool.","marker":"[11]"},{"why":"Supplies the TN5K detection and classification dataset used in stacked training.","marker":"[55]"},{"why":"Supplies the DDTI external test cohort used to judge cross-dataset segmentation and classification.","marker":"[60]"},{"why":"Supplies the PyRadiomics descriptor extraction that produces the radiomic evidence in the case record.","marker":"[56]"},{"why":"Is the strongest segmentation baseline the workflow must beat, giving the comparison its benchmark.","marker":"[64]"},{"why":"Is the multimodal language-model baseline for classification and report generation that the evidence-grounded workflow outperforms.","marker":"[68]"},{"why":"Supplies SHAP attributions used for clinician-facing explanation and stored in the evidence record.","marker":"[72]"},{"why":"Supplies the ReAct reasoning-and-acting loop used by the planner-executor report workflow.","marker":"[52]"}],"fun_headline_variants":["Auditable evidence record lifts thyroid report consistency to 86%","Agentic AI cuts thyroid reporting time 27%, boosts consistency","Coordinated evidence record beats isolated AI for thyroid diagnosis","ThyroidXAgent: auditable AI improves accuracy, cuts time","Evidence-grounded agentic AI lifts thyroid report consistency"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The broad multicentre-generalization claim rests on pooled per-centre metrics; if pooling hides centres that perform at chance, the claim that the workflow is consistent across heterogeneous settings loses force.","fun_headline_variants_meta":{"raw":{"variants":["Auditable evidence record lifts thyroid report consistency to 86%","Agentic AI cuts thyroid reporting time 27%, boosts consistency","Coordinated evidence record beats isolated AI for thyroid diagnosis","ThyroidXAgent: auditable AI improves accuracy, cuts time","Evidence-grounded agentic AI lifts thyroid report consistency"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1320,"prompt_tokens":991,"completion_tokens":329,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":607,"tokens_out":329,"duration_ms":3206,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:04:29.793780+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Look at Table S11: on the private NHC-MISD-TUS cohort, pooled AUROC is 0.819, but several centres sit near chance, including THYB_S_ZJ24 (n=269, AUROC 0.590), THYB_S_JL04 (n=12, AUROC 0.407) and THYB_S_YN05 (n=12, AUROC 0.500). If a validation protocol counted per-centre failures rather than pooled means, and a meaningful fraction of larger centres stayed below 0.6, the multicentre robustness claim would be falsified.","supporting_citations":[{"cited_title":"Tn5000: An ultrasound image dataset for thyroid nodule detection and classification","cited_arxiv_id":null,"evidence_quote":"Supplies the TN5K detection and classification dataset used in stacked training."},{"cited_title":"Gpt-5 system card.https://openai.com/index/gpt-5-system-card/ (2025)","cited_arxiv_id":null,"evidence_quote":"Is the multimodal language-model baseline for classification and report generation that the evidence-grounded workflow outperforms."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies SHAP attributions used for clinician-facing explanation and stored in the evidence record."}],"review_version":1}