Pith. sign in

REVIEW 4 major objections 6 minor 61 references

A single CT model can cover many diseases at specialist accuracy by feeding segmentation priors into vision–language alignment and attention.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 05:09 UTC pith:VAPHBOXR

load-bearing objection Solid empirical CT VLM that internalizes multi-head specialist priors and posts real multi-benchmark gains, including beating fused nnU-Net on tumor AUCs; open-lesion generalization is the softest claim. the 4 major comments →

arxiv 2607.09135 v1 pith:VAPHBOXR submitted 2026-07-10 cs.CV

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

classification cs.CV
keywords super-generalistvision-language modelmedical image diagnosislesion groundingspecialist-generalist synergyCTattention calibrationclass-agnostic lesion segmentation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Specialist models are accurate on one organ or tumor but do not transfer; vision–language generalists cover many diseases but often miss fine lesion detail and cannot show where they looked. This paper claims those limits are not fundamental. Its Super-Generalist (SuG) framework trains a shared vision encoder with three specialist heads—anatomy, class-specific lesions, and class-agnostic lesions—then uses the resulting masks to pull anatomy-specific image tokens into contrastive alignment with report snippets and to supervise text-conditioned attention so disease words focus on lesion voxels. On large abdominal and chest CT benchmarks the model reaches state-of-the-art multi-disease diagnosis, beats fused specialist baselines on several critical tumors, and produces attention maps that still light up lesions never given class-specific labels. If the result holds, one trainable system can deliver both breadth and the spatial accountability clinicians need.

Core claim

SuG shows that specialist-derived spatial priors can be injected into generalist vision–language learning so that a single model simultaneously delivers broad multi-disease coverage, specialist-competitive (and sometimes superior) tumor diagnosis, and lesion grounding that generalizes beyond the supervised lesion classes.

What carries the argument

Specialist-enhanced vision–language alignment plus lesion-guided attention calibration: multi-scale anatomy masks pool visual tokens for anatomy-wise contrastive loss with report snippets, while class-specific and class-agnostic lesion masks directly supervise sigmoid attention maps between lesion prompt embeddings and image tokens.

Load-bearing premise

The approach assumes that LLM-parsed report snippets and specialist masks—especially the class-agnostic head trained only inside a validity mask that ignores unlabeled organs—supply clean enough spatial priors that alignment and attention improve open multi-disease diagnosis rather than merely memorizing the supervised tumors.

What would settle it

Train SuG with the same architecture but replace real specialist masks by random or anatomy-scrambled masks (or drop the class-agnostic head); if multi-disease AUC and unannotated-lesion grounding collapse to pure-generalist levels while supervised-tumor AUC stays high, the claimed transfer of spatial priors is falsified.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • One trainable checkpoint can replace separate specialist pipelines for many abdominal and chest CT diseases while still matching or beating them on key tumors.
  • Text-conditioned attention maps become a built-in spatial explanation that clinicians can inspect without a second localization model.
  • Class-agnostic lesion supervision can surface abnormalities that never received voxel labels during training, reducing the annotation burden for rare findings.
  • The same progressive three-stage recipe can be re-applied whenever new organ or lesion segmentors become available, extending coverage without redesigning the generalist backbone.
  • Downstream report generation benefits from the same lesion-aware visual encoder, improving clinical content of generated text.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The validity-mask design for class-agnostic lesions is a practical answer to the partial-label problem that also appears in multi-organ anomaly detection and open-vocabulary segmentation outside CT.
  • If the class-agnostic head continues to improve with larger unlabeled volumes, the same architecture could bootstrap lesion grounding for MRI or ultrasound where dense annotation is even scarcer.
  • Anatomy-wise contrastive alignment may be more data-efficient than global image–report CLIP-style losses for any imaging domain whose reports are already written organ by organ.
  • Once attention calibration is reliable, the same maps could serve as weak pseudo-labels to iteratively expand the specialist training set without additional radiologist time.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes SuG (Super-Generalist), a framework that unifies specialist multi-task segmentation (anatomy, class-specific lesion, and class-agnostic lesion heads with a validity-masked loss) and generalist vision–language learning via specialist-enhanced anatomy-wise contrastive alignment (Eqs. 7–9) plus lesion-guided text-conditioned attention calibration (Eq. 10). Training proceeds in three progressive stages. Evaluated on abdominal and chest CT (MedVL-CT69K, Merlin, CT-RATE, RAD-ChestCT, and in-house LesionSeg* tumor sets), SuG reports SOTA multi-disease diagnosis AUCs, surpasses fused nnU-Net specialists on several tumor diagnosis benchmarks (e.g., 96.1% vs 94.7% average AUC on LesionSegAbdomen), and shows improved voxel-level grounding on annotated lesions (Table 6) with qualitative generalization to unannotated lesion types (Fig. 3, Fig. A4). Ablations (Table 7) and bootstrap significance tests (Appendix F) support component contributions.

Significance. If the results hold under independent scrutiny, this is a substantial empirical contribution to medical vision–language modeling: it shows that internalizing specialist spatial priors (rather than treating specialists as external advisors, as in GSCo) can simultaneously improve broad multi-disease diagnosis, specialist-competitive tumor detection, and lesion grounding on 3D CT. Strengths include multi-population/multi-anatomy benchmarks, explicit specialist–generalist baseline adaptations under a shared protocol, progressive multi-stage optimization, quantitative grounding metrics (AUC/AUPR/Cohen’s d) for annotated lesions, ablations, and case-level bootstrap significance against second-best methods. The work is of clear interest to medical imaging and foundation-model communities seeking clinically trustworthy generalists.

major comments (4)
  1. The central “super-generalist” claim of robust generalization to lesion types lacking class-specific supervision is only weakly secured. Table 6 and the left panel of Fig. 3 quantify grounding solely on the four annotated tumor categories; unannotated lesions (pleural effusion, renal cyst, gallbladder/bladder stone, etc.) receive only qualitative attention maps (Fig. 3 right, Fig. A4) and diagnosis-AUC ablations on a report-derived subset (Table 7 “unannotated localized lesions”). No voxel-level AUC/AUPR/Cohen’s d analogous to Table 6 is reported for any unannotated category. Without that, the open-scope part of the abstract/intro claim can still be explained by correlated report language and supervised-tumor transfer rather than transferable lesion-aware representations. Please add quantitative grounding on held-out unannotated lesion masks (or substantially tone down the claim).
  2. §2.2 Eq. (6) and Fig. A1: the class-agnostic lesion head is optimized only inside validity mask Vj that excludes unlabeled anatomies. This is a reasonable incomplete-label strategy, but it means the CA prior is never supervised outside the four annotated organ systems. The paper then attributes multi-disease gains and unannotated grounding largely to this CA branch (Table 7 last row). A load-bearing analysis is missing: (i) how often CA predictions fire outside annotated anatomies on MedVL/Merlin/CT-RATE, (ii) false-positive rates on truly negative regions, and (iii) whether removing CA still preserves most of the broad-diagnosis lift. Without this, the “beyond anatomies annotated during training” claim remains under-supported.
  3. §2.3 and Appendix B.1: anatomy-level text tokens are obtained by prompting Qwen on free-text reports. Anatomy-wise contrastive learning (Eqs. 7–9) and subsequent diagnosis scores depend on the cleanliness of these extractions. There is no quantitative audit (agreement with radiologist-extracted snippets, failure modes, laterality/subregion errors) nor an ablation that replaces LLM snippets with whole-report or rule-based text. Given that the skeptic’s main risk is noisy priors amplifying supervised signal, a short extraction-quality study or sensitivity analysis is needed to underwrite the specialist-enhanced alignment results in Tables 2–5.
  4. §3.1 / Appendix C.2: OpenVocabCT, HCFNet, and VLWS were originally language-driven segmentation methods; they are re-purposed as S+G diagnosis baselines under the authors’ protocol. While the paper is transparent about this adaptation, the main tables present them as the primary specialist–generalist competitors. Please clarify in the main text what architectural/objective changes were required for fair multi-disease diagnosis (vs. their original segmentation objective) and, if possible, include at least one strong pure-generalist and one pure-specialist upper bound that was not re-engineered, so readers can separate synergy gains from re-implementation effects.
minor comments (6)
  1. Abstract and §1: “clincial” → “clinical” (typo repeated in abstract).
  2. Fig. 2 caption and §2.4: the attention map uses a sigmoid of the dot product (α = σ(h·f)); clarify whether this is applied per scale before or after any temperature/normalization, and whether gradients flow into the specialist masks or only into the visual tokens.
  3. Table 2 vs. Table A15: main text reports only AUC averages; pointing readers more explicitly to full SE/SP in the appendix would help interpret class imbalance.
  4. §2.5 progressive stages: report whether Stage-3 calibration can degrade pure segmentation Dice on LesionSeg* (or confirm it does not), since L_specialist remains active.
  5. Related work §A.2: the distinction from GSCo is useful; a short sentence in the main introduction (not only appendix) would help readers place the contribution.
  6. In-house LesionSegAbdomen/Lung statistics (Table A2) are welcome; if any subset is public or will be released, state so for reproducibility.

Circularity Check

0 steps flagged

No significant circularity: empirical multi-task training with external voxel labels, image-report pairs, and held-out diagnosis/grounding metrics; no derivation reduces to its inputs by construction.

full rationale

SuG is a standard multi-stage empirical ML framework (specialist segmentation losses on DS voxel annotations + anatomy-wise contrastive alignment on DG image-report pairs + attention calibration supervised by lesion masks), not a first-principles derivation. Specialist heads (Eqs. 1-6) are optimized against external masks; generalist contrastive (Eqs. 7-9) uses LLM-parsed report snippets and predicted anatomy masks as inputs but is evaluated on held-out multi-disease labels (MedVL-CT69K, Merlin, CT-RATE, RAD-ChestCT); lesion-guided calibration (Eq. 10) is trained then measured on held-out LesionSegAbdomen test masks (Table 6) or qualitatively on unannotated cases. Progressive stages and ablations (Table 7) are ordinary cumulative training, not self-definitional. Overlapping-author baselines (fVLM, ViSD-Boost) are re-trained and outperformed, not load-bearing uniqueness theorems. No fitted parameter is renamed a prediction, no ansatz is smuggled via self-citation, and no result equals its training objective by construction on the evaluation set. Mild train/test label-type overlap for supervised grounding is ordinary supervised evaluation, not circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 3 invented entities

SuG’s claims rest on standard deep-learning practice plus domain choices about how to inject segmentation into VLMs. Free parameters are training hyperparameters and the learnable contrastive temperature. Axioms are domain assumptions that organ/lesion masks and LLM-extracted report slices are valid supervision for diagnosis and grounding. Invented entities are architectural constructs (SuG pipeline, dual CS/CA lesion heads with validity masking, lesion-guided attention calibration), not new physical objects; independent evidence is only the paper’s own benchmarks.

free parameters (4)
  • contrastive temperature τ
    Learnable temperature in anatomy-wise image–text contrastive matching (Eqs. 7–9); scales similarities and is fit during training.
  • stage-wise learning rates and schedules
    Stage 1 SGD 1e-2 for 1000 epochs; Stages 2–3 AdamW 1e-4→1e-6 cosine for 30+5 epochs; batch sizes 2 / 48 — chosen hyperparameters that affect final metrics.
  • target spacing and patch size [5.0,1.0,1.0], 96×256×384
    nnU-Net resampling/patch configuration fixed to match prior VL CT work; not derived from first principles.
  • multi-scale count S and ℓ2-normalized anatomy pooling
    Design choices for aggregating anatomy tokens across scales; ablation shows material AUC impact (Table 7).
axioms (4)
  • domain assumption Voxel-level anatomy and lesion labels (with validity mask Vj for unlabeled anatomies) are reliable enough to supervise a shared encoder that transfers to report-level diagnosis.
    Core of specialist pathway §§2.2 and Fig. A1; without trustworthy masks the specialist-enhanced alignment collapses.
  • domain assumption Qwen LLM extractions of anatomy-level descriptions and lesion prompts faithfully represent report content for contrastive and attention losses.
    Appendix B.1–B.2 prompts define the text side of alignment and calibration; errors would systematically misalign modalities.
  • ad hoc to paper Text-conditioned sigmoid attention supervised by resized lesion masks improves diagnostic focus without harming multi-disease coverage.
    §2.4 L_calibration and Stage-3 schedule; justified by ablation but not a standard theorem.
  • standard math Standard segmentation (CE+Dice) and bidirectional InfoNCE-style contrastive losses are appropriate objectives for CT understanding.
    Uses conventional L_seg and cross-entropy over softmax similarities; standard ML toolkit.
invented entities (3)
  • SuG Super-Generalist framework (specialist decoder + specialist-enhanced VL alignment + lesion-guided attention calibration) no independent evidence
    purpose: Unify broad disease scope, specialist-level accuracy, and lesion grounding in one CT model.
    Primary proposed system; evidence is internal benchmarks only.
  • Class-agnostic lesion head with validity-masked loss no independent evidence
    purpose: Capture lesions beyond anatomies with class-specific labels while avoiding false negatives outside annotated organs.
    §2.2 Eq. 6 and Fig. A1; generalization claims for unannotated lesions depend on this construct.
  • Lesion-guided cross-modal attention calibration no independent evidence
    purpose: Force disease text embeddings to attend to clinically relevant voxels.
    §2.4; grounding tables and Fig. 3 are the only external handle.

pith-pipeline@v1.1.0-grok45 · 35643 in / 3626 out tokens · 38761 ms · 2026-07-13T05:09:11.597023+00:00 · methodology

0 comments
read the original abstract

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabilities, yet often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework that unifies generalist vision-language learning with specialist objectives, enabling both broad generalization and specialist-level diagnostic capability. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that captures lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision.

Figures

Figures reproduced from arXiv: 2607.09135 by Jianpeng Zhang, Kai Cao, Ling Zhang, Qi Zhang, Shaoteng Zhang, Tingbo Liang, Wanxing Chang, Weiwei Cao, Yong Xia, Yu Shi, Yutong Xie, Zaiyi Liu.

Figure 1
Figure 1. Figure 1: Comparison of medical AI paradigms: Specialist, Generalist, and the proposed Super Generalist (SuG). We evaluate these paradigms across three critical dimensions: disease scope (range of tasks), performance (diagnostic accuracy), and grounding (lesion localization). (a) Specialists excel in performance and grounding but are limited to a narrow disease scope. (b) Conversely, Generalists support a wide scope… view at source ↗
Figure 2
Figure 2. Figure 2: An overview of the proposed SuG framework. SuG consists of two fundamental branches and two synergistic mechanisms: (a) Vision branch extracts multi-scale features and performs anatomy/lesion segmentation via a specialist decoder. (b) Text branch encodes reports into anatomy-level tokens and lesion prompt tokens. (c) Specialist-enhanced vision-language alignment leverages predicted anatomy masks to guide t… view at source ↗
Figure 3
Figure 3. Figure 3: Attention visualization on annotated and unannotated lesions. The left and right panels display representative cases of four annotated (liver, colon, stomach, pancreas) and unannotated (pleural effusion, renal cyst, gallbladder stone, bladder stone) lesion categories, respectively. Yellow boxes highlight ROIs, while red arrows pinpoint lesion locations. Attention maps from our generalist baseline, the spec… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

61 extracted references · 1 canonical work pages

  1. [1]

    The financial, operational, and clinical advantages of generalist radiology ai.Radiology, 316(3):e242362, 2025

    Siddhant Dogra, Xiaoman Zhang, Ezequiel Silva III, and Pranav Rajpurkar. The financial, operational, and clinical advantages of generalist radiology ai.Radiology, 316(3):e242362, 2025

  2. [2]

    Foundation models for generalist medical artificial intelligence.Nature, 616(7956):259–265, 2023

    Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar. Foundation models for generalist medical artificial intelligence.Nature, 616(7956):259–265, 2023

  3. [3]

    Medical image segmentation review: The success of u-net.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):10076–10095, 2024

    Reza Azad, Ehsan Khodapanah Aghdam, Amelie Rauland, Yiwei Jia, Atlas Haddadi Avval, Afshin Bozorgpour, Sanaz Karimijafarbigloo, Joseph Paul Cohen, Ehsan Adeli, and Dorit Merhof. Medical image segmentation review: The success of u-net.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):10076–10095, 2024

  4. [4]

    A chain of diagnosis framework for accurate and explainable radiology report generation.IEEE Transactions on Medical Imaging, 2025

    Haibo Jin, Haoxuan Che, Sunan He, and Hao Chen. A chain of diagnosis framework for accurate and explainable radiology report generation.IEEE Transactions on Medical Imaging, 2025

  5. [5]

    A review of the application of deep learning in medical image classification and segmentation.Annals of translational medicine, 8(11):713, 2020

    Lei Cai, Jingyang Gao, and Di Zhao. A review of the application of deep learning in medical image classification and segmentation.Annals of translational medicine, 8(11):713, 2020

  6. [6]

    Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation

    Yutong Xie, Jianpeng Zhang, Chunhua Shen, and Yong Xia. Cotr: Efficiently bridging cnn and transformer for 3d medical image segmentation. InInternational conference on medical image computing and computer-assisted intervention, pages 171–180. Springer, 2021

  7. [7]

    nnu-net: Self-adapting framework for u-net-based medical image segmentation.arXiv preprint arXiv:1809.10486, 2018

    Fabian Isensee, Jens Petersen, Andre Klein, David Zimmerer, Paul F Jaeger, Simon Kohl, Jakob Wasserthal, Gregor Koehler, Tobias Norajitra, Sebastian Wirkert, et al. nnu-net: Self-adapting framework for u-net-based medical image segmentation.arXiv preprint arXiv:1809.10486, 2018

  8. [8]

    End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography

    Diego Ardila, Atilla P Kiraly, Sujeeth Bharadwaj, Bokyung Choi, Joshua J Reicher, Lily Peng, Daniel Tse, Mozziyar Etemadi, Wenxing Ye, Greg Corrado, et al. End-to-end lung cancer screening with three-dimensional deep learning on low-dose chest computed tomography. Nature medicine, 25(6):954–961, 2019

  9. [9]

    Large-scale pancreatic cancer detection via non-contrast ct and deep learning.Nature medicine, 29(12):3033–3043, 2023

    Kai Cao, Yingda Xia, Jiawen Yao, Xu Han, Lukas Lambert, Tingting Zhang, Wei Tang, Gang Jin, Hui Jiang, Xu Fang, et al. Large-scale pancreatic cancer detection via non-contrast ct and deep learning.Nature medicine, 29(12):3033–3043, 2023

  10. [10]

    Parse and recall: Towards accurate lung nodule ma- lignancy prediction like radiologists

    Jianpeng Zhang, Xianghua Ye, Jianfeng Zhang, Yuxing Tang, Minfeng Xu, Jianfei Guo, Xin Chen, Zaiyi Liu, Jingren Zhou, Le Lu, et al. Parse and recall: Towards accurate lung nodule ma- lignancy prediction like radiologists. InInternational Conference on Medical Image Computing and Computer-Assisted Intervention, pages 199–209. Springer, 2023. 10

  11. [11]

    Towards a comprehensive, efficient and promptable anatomic structure segmentation model using 3d whole-body ct scans

    Heng Guo, Jianfeng Zhang, Jiaxing Huang, Tony CW Mok, Dazhou Guo, Ke Yan, Le Lu, Dakai Jin, and Minfeng Xu. Towards a comprehensive, efficient and promptable anatomic structure segmentation model using 3d whole-body ct scans. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 3247–3256, 2025

  12. [12]

    Contrastive learning of medical visual representations from paired images and text

    Yuhao Zhang, Hang Jiang, Yasuhide Miura, Christopher D Manning, and Curtis P Langlotz. Contrastive learning of medical visual representations from paired images and text. InMachine learning for healthcare conference, pages 2–25. PMLR, 2022

  13. [14]

    Making the most of text semantics to improve biomedical vision–language processing

    Benedikt Boecking, Naoto Usuyama, Shruthi Bannur, Daniel C Castro, Anton Schwaighofer, Stephanie Hyland, Maria Wetscherek, Tristan Naumann, Aditya Nori, Javier Alvarez-Valle, et al. Making the most of text semantics to improve biomedical vision–language processing. In European conference on computer vision, pages 1–21. Springer, 2022

  14. [15]

    Merlin: A vision language foundation model for 3d computed tomography

    Louis Blankemeier, Joseph Paul Cohen, Ashwin Kumar, Dave Van Veen, Syed Jamal Safdar Gardezi, Magdalini Paschali, Zhihong Chen, Jean-Benoit Delbrouck, Eduardo Reis, Cesar Truyts, et al. Merlin: A vision language foundation model for 3d computed tomography. Research Square, pages rs–3, 2024

  15. [16]

    Umind-vl: A generalist ultrasound vision-language model for unified grounded perception and comprehensive interpretation.arXiv preprint arXiv:2511.22256, 2025

    Dengbo Chen, Ziwei Zhao, Kexin Zhang, Shishuang Zhao, Junjie Hou, Yaqian Wang, Nianxi Liao, Anlan Sun, Fei Gao, Jia Ding, et al. Umind-vl: A generalist ultrasound vision-language model for unified grounded perception and comprehensive interpretation.arXiv preprint arXiv:2511.22256, 2025

  16. [17]

    Generalist versus specialist vision foundation models for ocular disease and oculomics.arXiv preprint arXiv:2509.03421, 2025

    Yukun Zhou, Paul Nderitu, Jocelyn Hui Lin Goh, Justin Engelmann, Siegfried K Wagner, Anran Ran, Hongyang Jiang, Lie Ju, Ke Zou, Sahana Srinivasan, et al. Generalist versus specialist vision foundation models for ocular disease and oculomics.arXiv preprint arXiv:2509.03421, 2025

  17. [18]

    Clarify: A specialist-generalist framework for accurate and lightweight dermatological visual question answering.arXiv preprint arXiv:2508.18430, 2025

    Aranya Saha, Tanvir Ahmed Khan, Ismam Nur Swapnil, and Mohammad Ariful Haque. Clarify: A specialist-generalist framework for accurate and lightweight dermatological visual question answering.arXiv preprint arXiv:2508.18430, 2025

  18. [19]

    Gsco: Towards generalizable ai in medicine via generalist-specialist collaboration.arXiv preprint arXiv:2404.15127, 2024

    Sunan He, Yuxiang Nie, Hongmei Wang, Shu Yang, Yihui Wang, Zhiyuan Cai, Zhixuan Chen, Yingxue Xu, Luyang Luo, Huiling Xiang, et al. Gsco: Towards generalizable ai in medicine via generalist-specialist collaboration.arXiv preprint arXiv:2404.15127, 2024

  19. [20]

    Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen technical report.arXiv preprint arXiv:2309.16609, 2023

  20. [21]

    Large-scale and fine-grained vision-language pre-training for enhanced ct image understanding.arXiv preprint arXiv:2501.14548, 2025

    Zhongyi Shui, Jianpeng Zhang, Weiwei Cao, Sinuo Wang, Ruizhe Guo, Le Lu, Lin Yang, Xianghua Ye, Tingbo Liang, Qi Zhang, et al. Large-scale and fine-grained vision-language pre-training for enhanced ct image understanding.arXiv preprint arXiv:2501.14548, 2025

  21. [22]

    A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities.arXiv preprint arXiv:2403.17834, 5, 2024

    Ibrahim Ethem Hamamci, Sezgin Er, Furkan Almas, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Irem Dogan, Muhammed Furkan Dasdelen, Bastian Wittmann, Enis Simsar, Mehmet Simsar, et al. A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities.arXiv preprint arXiv:2403.17834, 5, 2024

  22. [23]

    Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes.Medical image analysis, 67:101857, 2021

    Rachel Lea Draelos, David Dov, Maciej A Mazurowski, Joseph Y Lo, Ricardo Henao, Geof- frey D Rubin, and Lawrence Carin. Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes.Medical image analysis, 67:101857, 2021

  23. [24]

    Yu, and Xiaofeng Yang

    Yuheng Li, Yuxiang Lai, Maria Thor, Deborah Marshall, Zachary Buchwald, David S. Yu, and Xiaofeng Yang. Towards universal text-driven ct image segmentation.arXiv preprint arXiv:2503.06030, 2025. 11

  24. [25]

    Hybrid cross-modality fusion network for medical image segmentation with contrastive learning.Engineering Applications of Artificial Intelligence, 144:110073, 2025

    Xichuan Zhou, Qianqian Song, Jing Nie, Yujie Feng, Haijun Liu, Fu Liang, Lihui Chen, and Jin Xie. Hybrid cross-modality fusion network for medical image segmentation with contrastive learning.Engineering Applications of Artificial Intelligence, 144:110073, 2025

  25. [26]

    Vision-language semantic grounding for multi-domain crop-weed segmentation.arXiv preprint arXiv:2602.23677, 2026

    Nazia Hossain, Xintong Jiang, Yu Tian, Philippe Seguin, O Grant Clark, and Shangpeng Sun. Vision-language semantic grounding for multi-domain crop-weed segmentation.arXiv preprint arXiv:2602.23677, 2026

  26. [27]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pages 8748–8763. PmLR, 2021

  27. [28]

    Joint learning of localized representations from medical images and reports

    Philip Müller, Georgios Kaissis, Congyu Zou, and Daniel Rueckert. Joint learning of localized representations from medical images and reports. InEuropean conference on computer vision, pages 685–701. Springer, 2022

  28. [29]

    Imitate: Clinical prior guided hierarchical vision-language pre-training.IEEE Transactions on Medical Imaging, 2024

    Che Liu, Sibo Cheng, Miaojing Shi, Anand Shah, Wenjia Bai, and Rossella Arcucci. Imitate: Clinical prior guided hierarchical vision-language pre-training.IEEE Transactions on Medical Imaging, 2024

  29. [30]

    Bootstrapping chest ct image understanding by distilling knowledge from x-ray expert models

    Weiwei Cao, Jianpeng Zhang, Yingda Xia, Tony CW Mok, Zi Li, Xianghua Ye, Le Lu, Jian Zheng, Yuxing Tang, and Ling Zhang. Bootstrapping chest ct image understanding by distilling knowledge from x-ray expert models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11238–11247, 2024

  30. [31]

    Boosting vision semantic density with anatomy normality mod- eling for medical vision-language pre-training

    Weiwei Cao, Jianpeng Zhang, Zhongyi Shui, Sinuo Wang, Zeli Chen, Xi Li, Le Lu, Xianghua Ye, Qi Zhang, Tingbo Liang, et al. Boosting vision semantic density with anatomy normality mod- eling for medical vision-language pre-training. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 23041–23050, 2025

  31. [32]

    A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities.CoRR, 2024

    Ibrahim Ethem Hamamci, Sezgin Er, Furkan Almas, Ayse Gulnihan Simsek, Sevval Nil Esirgun, Irem Dogan, Muhammed Furkan Dasdelen, Bastian Wittmann, Enis Simsar, Mehmet Simsar, et al. A foundation model utilizing chest ct volumes and radiology reports for supervised-level zero-shot detection of abnormalities.CoRR, 2024

  32. [33]

    Medclip-samv2: Towards universal text-driven medical image segmentation.Medical Image Analysis, page 103749, 2025

    Taha Koleilat, Hojat Asgariandehkordi, Hassan Rivaz, and Yiming Xiao. Medclip-samv2: Towards universal text-driven medical image segmentation.Medical Image Analysis, page 103749, 2025

  33. [34]

    U-kan makes strong backbone for medical image segmentation and generation

    Chenxin Li, Xinyu Liu, Wuyang Li, Cheng Wang, Hengyu Liu, Yifan Liu, Zhen Chen, and Yixuan Yuan. U-kan makes strong backbone for medical image segmentation and generation. InProceedings of the AAAI conference on artificial intelligence, volume 39, pages 4652–4660, 2025

  34. [35]

    Semisam+: rethinking semi-supervised medical image segmentation in the era of foundation models.Medical Image Analysis, page 103733, 2025

    Yichi Zhang, Bohao Lv, Le Xue, Wenbo Zhang, Yuchen Liu, Yu Fu, Yuan Cheng, and Yuan Qi. Semisam+: rethinking semi-supervised medical image segmentation in the era of foundation models.Medical Image Analysis, page 103733, 2025

  35. [36]

    Medianomaly: A comparative study of anomaly detection in medical images.Medical Image Analysis, 102:103500, 2025

    Yu Cai, Weiwen Zhang, Hao Chen, and Kwang-Ting Cheng. Medianomaly: A comparative study of anomaly detection in medical images.Medical Image Analysis, 102:103500, 2025

  36. [37]

    Medical imaging: a critical review on x-ray imaging for the detection of infection.Biomedical Materials & Devices, 4(1):1–45, 2026

    Egwonor Loveth Irede, Omowunmi Rebecca Aworinde, Ogunnaike Korede Lekan, Osemu- diamhen D Amienghemhen, Tochukwu Perpetua Okonkwo, Asishana Paul Onivefu, and Ik- hazuagbe H Ifijen. Medical imaging: a critical review on x-ray imaging for the detection of infection.Biomedical Materials & Devices, 4(1):1–45, 2026

  37. [38]

    Deep learning-based object detection algorithms in medical imaging: Systematic review.Heliyon, 11(1), 2025

    Carina Albuquerque, Roberto Henriques, and Mauro Castelli. Deep learning-based object detection algorithms in medical imaging: Systematic review.Heliyon, 11(1), 2025

  38. [39]

    Resvit fusionnet model: An explainable ai-driven approach for automated grading of diabetic retinopathy in retinal images.Computers in Biology and Medicine, 186:109656, 2025

    Amna Ikram and Azhar Imran. Resvit fusionnet model: An explainable ai-driven approach for automated grading of diabetic retinopathy in retinal images.Computers in Biology and Medicine, 186:109656, 2025. 12

  39. [40]

    Diffmic-v2: Medical image classification via improved diffusion network.IEEE Transactions on Medical Imaging, 44(5):2244–2255, 2025

    Yijun Yang, Huazhu Fu, Angelica I Aviles-Rivero, Zhaohu Xing, and Lei Zhu. Diffmic-v2: Medical image classification via improved diffusion network.IEEE Transactions on Medical Imaging, 44(5):2244–2255, 2025

  40. [41]

    A deep ensemble learning framework for glioma segmentation and grading prediction.Scientific Reports, 15(1):4448, 2025

    Liang Wen, Hui Sun, Guobiao Liang, and Yue Yu. A deep ensemble learning framework for glioma segmentation and grading prediction.Scientific Reports, 15(1):4448, 2025

  41. [42]

    A novel pd-1/pd-l1 pathway-related seven-gene signature for the development and validation of the prognosis prediction model for breast cancer

    Peng Zhang, Jingjing Yang, Xiaolong Zhong, Heloisa Sobreiro Selistre-de Araujo, Stergios Boussios, Yongneng Ma, and Hua Fang. A novel pd-1/pd-l1 pathway-related seven-gene signature for the development and validation of the prognosis prediction model for breast cancer. Translational cancer research, 13(3):1554–1566, 2024

  42. [43]

    Advancements in artificial intelligence for prostate cancer: Optimizing diagnosis, treatment, and prognostic assessment

    Yuki Arita, Christian Roest, Thomas C Kwee, Ramesh Paudyal, Alfonso Lema-Dopico, Stefan Fransen, Daisuke Hirahara, Eichi Takaya, Ryo Ueda, Lisa Ruby, et al. Advancements in artificial intelligence for prostate cancer: Optimizing diagnosis, treatment, and prognostic assessment. Asian Journal of Urology, 2025

  43. [44]

    Histo-genomic knowledge association for cancer prognosis from histopathology whole slide images.IEEE Transactions on Medical Imaging, 44(5):2170–2181, 2025

    Zhikang Wang, Yumeng Zhang, Yingxue Xu, Seiya Imoto, Hao Chen, and Jiangning Song. Histo-genomic knowledge association for cancer prognosis from histopathology whole slide images.IEEE Transactions on Medical Imaging, 44(5):2170–2181, 2025

  44. [45]

    U-net: Convolutional networks for biomedical image segmentation

    Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation. InInternational Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015

  45. [46]

    nn- former: Interleaved transformer for volumetric segmentation.arXiv preprint arXiv:2109.03201, 2021

    Hong-Yu Zhou, Jiansen Guo, Yinghao Zhang, Lequan Yu, Liansheng Wang, and Yizhou Yu. nn- former: Interleaved transformer for volumetric segmentation.arXiv preprint arXiv:2109.03201, 2021

  46. [47]

    Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images

    Ali Hatamizadeh, Vishwesh Nath, Yucheng Tang, Dong Yang, Holger R Roth, and Daguang Xu. Swin unetr: Swin transformers for semantic segmentation of brain tumors in mri images. In International MICCAI brainlesion workshop, pages 272–284. Springer, 2021

  47. [48]

    A visual–language foundation model for pathology image analysis using medical twitter.Nature medicine, 29(9):2307–2316, 2023

    Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual–language foundation model for pathology image analysis using medical twitter.Nature medicine, 29(9):2307–2316, 2023

  48. [49]

    Advancing radiograph repre- sentation learning with masked record modeling.arXiv preprint arXiv:2301.13155, 2023

    Hong-Yu Zhou, Chenyu Lian, Liansheng Wang, and Yizhou Yu. Advancing radiograph repre- sentation learning with masked record modeling.arXiv preprint arXiv:2301.13155, 2023

  49. [50]

    Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.arXiv preprint arXiv:2303.00915, 2023

    Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, et al. Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs.arXiv preprint arXiv:2303.00915, 2023

  50. [51]

    Pmc-clip: Contrastive language-image pre-training using biomedical documents

    Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Contrastive language-image pre-training using biomedical documents. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 525–536. Springer, 2023

  51. [52]

    Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning.Nature biomedical engineering, 6(12):1399–1406, 2022

    Ekin Tiu, Ellie Talius, Pujan Patel, Curtis P Langlotz, Andrew Y Ng, and Pranav Rajpurkar. Expert-level detection of pathologies from unannotated chest x-ray images via self-supervised learning.Nature biomedical engineering, 6(12):1399–1406, 2022

  52. [53]

    Multi- granularity cross-modal alignment for generalized medical visual representation learning.Ad- vances in neural information processing systems, 35:33536–33549, 2022

    Fuying Wang, Yuyin Zhou, Shujun Wang, Varut Vardhanabhuti, and Lequan Yu. Multi- granularity cross-modal alignment for generalized medical visual representation learning.Ad- vances in neural information processing systems, 35:33536–33549, 2022

  53. [54]

    Towards generalizable ai in medicine via generalist-specialist collaboration.Nature Biomedical Engineering, 2026

    Sunan He, Yuxiang Nie, Hongmei Wang, Shu Yang, Yihui Wang, Zhiyuan Cai, Zhixuan Chen, Yingxue Xu, Luyang Luo, Huiling Xiang, Xi Lin, Mingxiang Wu, Yifan Peng, George Shih, Ziyang Xu, Xian Wu, Qiong Wang, Ronald Cheong Kin Chan, Xiaohui Duan, Varut Vardhanabhuti, Winnie Chiu Wing Chu, Yefeng Zheng, Pranav Rajpurkar, Kang Zhang, and Hao Chen. Towards genera...

  54. [55]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), pages 4171–4186, 2019

  55. [56]

    Masked image modeling advances 3d medical image analysis

    Zekai Chen, Devansh Agarwal, Kshitij Aggarwal, Wiem Safta, Mariann Micsinai Balan, and Kevin Brown. Masked image modeling advances 3d medical image analysis. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1970–1980, 2023. 14 A Related work A.1 Medical specialist AI Specialist models have long been the dominant pa...

  56. [57]

    Extract explicit abnormal imaging findings from FINDINGS/DESCRIPTION

  57. [58]

    Keep only generic core signs; remove size, exact location, laterality, measurements, and comparison terms

  58. [59]

    Identify diseases from IMPRESSION/CONCLUSION if present; otherwise derive disease/problem labels from abnormal FINDINGS

  59. [60]

    For each disease, collect the supporting core signs from FINDINGS

  60. [61]

    Map each disease to exactly one target anatomy using predefined rules and standard medical knowledge

  61. [62]

    Localized Lesion

    Build a comma-separated English core sign string. Representative anatomy mapping rules: - large bowel: colon, rectum, cecum, appendix, anal canal - small bowel: jejunum, ileum - duodenum: duodenum only - gallbladder: gallbladder, cystic duct, and biliary duct terms if no separate duct anatomy exists - portal vein: portal vein system, including splenic vei...