REVIEW 4 major objections 5 minor 35 references
DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification
T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read A trimodal fusion model with task-specific modality gates achieves a mean macro-F1 of 0.801 on brain tumor subtype classification, narrowly beating the challenge baseline under fully multimodal input, but the advantage depends on histopatho
desk verdict Candid, well-documented challenge paper whose headline result (0.801 vs 0.796 on n=36) is within noise and rests on a possibly illegitimate class exclusion; the pathology-dependence finding is the real contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Task-specific gates: three independent linear layers, one per classification task, each computing softmax weights over the three modality latents (MRI, histopathology, and radiology-report embeddings). Instead of producing a single fused representation, each task receives its own convex combination of the modalities, with missing modalities masked to zero weight. This mechanism lets Level-1 molecular type draw heavily on histopathology while the grade tasks weight modalities differently. The system also uses a large language model as a frozen report encoder, applies modality dropout during training, and enforces label-hierarchy constraints in post-processing at high confidence.
What would settle it
Re-run the official evaluation with the six excluded 'Other/NEC' subjects included in the test pool, using the original label set, and check whether the Level-1 and mean macro-F1 still exceed the baseline; if the gap narrows or reverses, the reported advantage stems from the subject exclusion rather than the gating architecture.
Extended reading notes
Core claim
The central discovery is that giving each classification task its own learned gate over the three modality latents—MRI, histopathology, and radiology report embeddings—yields higher macro-F1 than fusing them through a single shared gate before all three heads. The primary system, trained with 5-fold cross-validation and logit ensembling, reaches a post-verification mean macro-F1 of 0.801 under the Fully Multimodal condition and ranks second among code-verified teams, exceeding the baseline's 0.796. The edge is concentrated in Level-1 Molecular Type, where the system scores 0.867 versus the baseline's 0.672. The paper further establishes that this advantage is not robust: when histopathology
Load-bearing premise
The evaluation drops all subjects whose molecular subtype is labeled 'Other/NEC', on the grounds that the class is not enumerated in the public task description; if the official task includes this class, the Level-1 macro-F1 comparisons are made on a different label distribution and the headline result may be an artifact of the filter.
Editorial extensions
If this is right
- Learned per-task gates outperform a shared gated representation on all three tasks, with the largest gap on Level-1 molecular type.
- The baseline still wins on the two grade-based tasks (WHO Grade and LGG vs HGG), so the fusion strategy is not uniformly better.
- Dropping histopathology at inference collapses the system's mean macro-F1 from 0.801 to 0.642, below the baseline's 0.744, and on a cohort without histopathology the system scores 0.493 vs 0.797.
- The claimed advantage over the baseline is therefore conditional on the histopathology modality being present in the input.
Reading between the lines
- The subject filter that drops six 'Other/NEC' subjects from the evaluation pool changes the label distribution for the Level-1 metric; if the challenge's official evaluation includes that class, the 0.867 vs 0.672 comparison—and the overall 'exceeds baseline' claim—may not transfer to the intended task.
- The strong dependence on histopathology undercuts the paper's motivating scenario of early diagnosis while awaiting the pathology report: the system is most competitive precisely when the slowest-to-obtain modality is already available.
- The per-task gate idea is a general recipe for multi-task multimodal problems: when different labels correspond to different biological or clinical questions, letting each head choose its own modality weights may be worth testing in other medical and non-medical fusion settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the DS@GT ARC submission to MEDIQA-CORE 2026 Task 1 (brain tumor subtype classification). The authors build a trimodal system combining pre-extracted MRI (NeuroVFM) and histopathology (Prov-GigaPath) embeddings with radiology report embeddings from Llama-3.1-8B-Instruct. Two fusion architectures are explored: a cross-modal self-attentive model and the primary submission with task-specific gates. Using 5-fold cross-validation, logit ensembling, and hierarchical post-processing, the system reports a post-verification mean macro-F1 of 0.801 on Test Set 1 versus the organizer baseline's 0.796, ranking second among verified teams. The authors also report ablations showing the advantage depends on histopathology availability and that the system underperforms the baseline when histopathology is missing.
Significance. If the reported result holds, the paper provides evidence that task-specific gating can improve molecular subtype classification through targeted modality weighting, and the modality ablation study is a useful contribution to understanding where a third modality helps. The paper also ships code on GitHub and relies on official post-verification scores, which is a strength. However, the central claim of outperforming the baseline rests on a very small aggregate difference without uncertainty quantification, and the exclusion of the 'Other/NEC' class from evaluation raises a protocol-validity concern. The inconsistency about the baseline's use of reports further undermines the comparison. The architectural idea is plausible but the empirical support as presented is not yet convincing.
major comments (4)
- [§4.2, Table 5] The headline claim that the submitted system 'exceeds' the organizer's baseline rests on a mean macro-F1 difference of 0.801 vs 0.796 (Δ=0.005) on Test Set 1 (n=36). The paper reports no confidence intervals, bootstrap, or significance test. Per-task, the system gains 0.195 on Level-1 but loses 0.058 on WHO Grade and 0.120 on LGG/HGG, so the aggregate difference is a small net of opposing effects. With roughly 12 subjects per macro-F1 class, one or two prediction changes can exceed 0.005. The official rerun confirms arithmetic, not that the difference is distinguishable from sampling noise. Please provide uncertainty quantification (e.g., bootstrap/permutation) or soften the 'exceeding' claim.
- [§3.3.1 Subject Filtering] The manuscript excludes all subjects with level1_label==3 ('Other/NEC') from training, validation, and the held-out evaluation pool, on the grounds that this class is 'not enumerated in the public task description' and that including it 'consistently hurt macro-F1 on the Level-1 task in preliminary experiments.' If the official task includes this class, the reported Level-1 macro-F1 (0.867 vs 0.672) and the resulting mean are computed on a different label distribution than the baseline, making the comparison invalid. Please state explicitly whether the organizers' official evaluation also excludes this class; if not, the headline comparison is an artifact of dropping difficult cases.
- [§3.2 vs Tables 5/6] Section 3.2 states the baseline 'does not use the radiology reports,' yet Table 6 shows the baseline's mean macro-F1 drops from 0.797 in MRI+Reports to 0.405 in MRI Images Only, and Table 5 shows the baseline's LGG/HGG collapses from 0.920 to 0.438 when Reports are dropped. If the baseline does not read reports, removing reports should have no effect. Please explain this apparent contradiction; if the baseline in Tables 5/6 is a different system or uses reports, the comparison narrative and Section 3.2 must be corrected.
- [§4.1, Table 4 vs §4.2, Table 5] The baseline's Fully Multimodal mean macro-F1 changes from 0.728 in the pre-verification leaderboard (Table 4) to 0.796 in the post-verification results (Table 5) without comment. The submitted system also changes from 0.824 to 0.801. If the post-verification rerun used a different protocol, baseline, or preprocessing, this should be stated; as written, the abstract's 'exceeding baseline' claim refers to the post-verification numbers while the pre-verification discussion draws a different conclusion (rank 4th of 5). Please reconcile these numbers.
minor comments (5)
- [Footnote (author email)] The corresponding author email contains corrupted characters ('envel⌢pe-⌢penhtruong47@gatech.edu'); fix the email address.
- [§3.1, Table 1] Text says 161 (91%) training subjects have MRI, but Table 1 gives 35+132=167 (91.3%). Please reconcile the numbers.
- [Figures 1 and 2] Figures 1 and 2 are referenced in the text but not present in the manuscript text provided; ensure they are included in the final version.
- [Abstract vs §4.1] The abstract says 'ranking second among the teams whose code passed verification' while §4.1 says 'ranking 4th of 5 teams on the final leaderboard.' Clarify which ranking is meant in each place.
- [§3.3.1] The phrase 'for architectural consistency across encoder backbones we ablated' is unclear; consider rephrasing.
Circularity Check
No circularity found: test evaluation is external, model selection uses validation folds, and no prediction reduces by construction to its inputs.
full rationale
The paper's derivation chain is self-contained with respect to the circularity tests. The submitted system is trained on organizer-provided frozen embeddings (NeuroVFM MRI, Prov-GigaPath histopathology, and Llama-3.1 report embeddings) and is evaluated on organizer-held-out test sets (Test #1, n=36; Test #2, n=18). Model selection and post-processing thresholds are chosen using stratified cross-validation on the training split, and the reported leaderboard scores are obtained by the organizers' independent code re-run (post-verification). No parameter is fitted to the held-out labels and then reported as a prediction of those same labels. The task-specific gating architecture is a genuine fusion design choice; the comparison against the cross-modal model is an ablation, not a tautology. The class-3 ('Other / NEC') subject exclusion is a sample-selection / label-distribution concern, not a circular-derivation concern: the exclusion changes the evaluation pool but is not a fitted quantity being renamed as a prediction. The internal inconsistency that the baseline is said not to use radiology reports while Tables 5/6 show its scores dropping when reports are removed, and the lack of confidence intervals on the 0.801 vs 0.796 mean macro-F1 gap, are correctness/validity issues, not circularity. The paper does not invoke self-citations as load-bearing evidence and does not import any uniqueness theorem or ansatz from the authors' prior work. The core claim is an empirical, externally evaluated result.
Assumptions & free parameters
free parameters (6)
- Post-processing confidence threshold τ =
0.80
- Class weights w_c =
Level-1: 1.00, 1.60, 7.77; LGG/HGG: 4.23, 1.00; WHO Grade: 3.43, 4.29, 1.00
- Modality dropout rate p_drop =
0.5
- Architecture capacities =
D=128, P=64, report bottleneck=256, self-attention blocks=2, heads=4
- Label smoothing ε =
0.1
- Early stopping and schedule constants =
5,000 iterations, cosine anneal to 1e-6, patience=5
assumptions (5)
- domain assumption The frozen NeuroVFM MRI and Prov-GigaPath histopathology embeddings are faithful, sufficient representations for the three classification tasks.
- domain assumption Mean-pooled final-layer Llama-3.1-8B hidden states capture radiology-report information useful for tumor typing.
- domain assumption The label hierarchy 'Level-1 prediction ⇒ LGG/HGG, LGG/HGG ⇒ WHO grade' holds in the CoRe-BT label set.
- ad hoc to paper Subjects with level1_label == 3 ('Other / NEC') may be removed from training, validation, and the held-out evaluation pool.
- domain assumption The organizers' post-verification re-runs and score tables are accurate and faithful to the submitted code.
Cite this review
Pith. "Pith review of DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification." pith.science (2026). https://pith.science/paper/GCJNCVR5
@misc{pith2026260800086,
author = {Pith},
title = {Pith review of: DS@GT ARC at MEDIQA-CORE-Task-1 2026: Trimodal Model Fusion with Task-Specific Gates for Brain Tumor Subtype Classification},
year = {2026},
howpublished = {\url{https://pith.science/paper/GCJNCVR5}},
note = {Machine review of arXiv:2608.00086}
}
read the original abstract
Brain tumor diagnosis is a time-sensitive process in which patients may wait weeks for a finalized pathology report. This problem motivates automated systems that classify tumor subtype from multimodal inputs. This paper details the DS@GT ARC team's work for ImageCLEFmed MEDIQA-CORE 2026 Task~1, Brain Tumor Subtype Classification. The task evaluates three glioma classification problems: Level-1 Molecular Type, LGG vs HGG, and WHO Grade. We combine pre-extracted MRI (NeuroVFM) and histopathology (Prov-GigaPath) embeddings with free-text radiology reports. Our team explored two trimodal fusion architectures, two report encoders (RadBERT and Llama-3.1-8B-Instruct), and a biologically motivated post-processing stage. We achieve a mean macro-F1 of 0.801 under the Fully Multimodal condition, exceeding the organizers' baseline of 0.796 and ranking second among the teams whose code passed verification. Additional evaluation across modality-dropping conditions shows that this advantage depends heavily on the availability of the histopathology modality, and that our system falls behind the baseline when modalities are missing. Our code is available on GitHub at https://github.com/dsgt-arc/imageclef-mediqacore-2026.
Figures
Reference graph
Works this paper leans on
-
[1]
B. Alther, V. Mylius, M. Weller, A. R. Gantenbein, From first symptoms to diagnosis: Initial clinical presentation of primary brain tumors, Clinical and Translational Neuroscience 4 (2020) 17. doi:10.1177/2514183X20968368
-
[2]
A. L. Stensjøen, O. Solheim, K. A. Kvistad, A. K. Håberg, Ø. Salvesen, E. M. Berntsen, Growth dynamics of untreated glioblastomas in vivo, Neuro-Oncology 17 (2015) 1402–1411. doi:10.1093/ neuonc/nov029
2015
-
[3]
J. E. Heras Rivera, D. K. Low, X. Xiong, J. J. Ruzevick, D. D. Child, W. wai Yim, M. Kurt, A. Ben Abacha, CoRe-BT: A multimodal radiology-pathology-text benchmark for robust brain tumor typing, CoRR abs/2603.03618 (2026). URL: https://arxiv.org/abs/2603.03618
arXiv 2026
-
[4]
Ben Abacha, J
A. Ben Abacha, J. E. Heras Rivera, D. K. Low, W. Yim, J. Ruzevick, D. Child, M. Kurt, Overview of the MEDIQA-CORE 2026 task 1 on brain tumor subtype classification, in: CLEF 2026 Working Notes, CEUR Workshop Proceedings, CEUR-WS.org, Jena, Germany, 2026
2026
-
[5]
Ionescu, H
B. Ionescu, H. Müller, D. Stanciu, A. Radu, R. Bolborici, M. Negru, A. Ene, V. Vasilescu, A.-A. Nicolae, L. Ştefan, M. Constantin, M. Dogariu, A. Andrei, H. Damm, T. M. G. Pakull, A. Ben Abacha, A. García Seco de Herrera, C. M. Friedrich, R. Brüngel, L. Reinartz, H. Schäfer, C. S. Schmidt, B. Bracke, P. Nath, B. Eryılmaz, M. Hjuler, D. Fabre, C. Lemaire, ...
2026
-
[6]
D. N. Louis, A. Perry, P. Wesseling, D. J. Brat, I. A. Cree, D. Figarella-Branger, C. Hawkins, H. K. Ng, S. M. Pfister, G. Reifenberger, R. Soffietti, A. von Deimling, D. W. Ellison, The 2021 WHO Classification of Tumors of the Central Nervous System: a summary, Neuro-Oncology 23 (2021) 1231–1251. doi:10.1093/neuonc/noab106
-
[7]
Kondepudi, et al., Health System Learning Achieves Generalist Neuroimaging Models, 2025
A. Kondepudi, et al., Health System Learning Achieves Generalist Neuroimaging Models, 2025. URL: https://arxiv.org/abs/2511.18640.arXiv:2511.18640
arXiv 2025
-
[8]
H. Xu, N. Usuyama, J. Bagga, S. Zhang, R. Rao, T. Naumann, et al., A whole-slide foundation model for digital pathology from real-world data, Nature 630 (2024) 181–188. doi: 10.1038/ s41586-024-07441-w
2024
Show all 35 references
-
[9]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, et al., The Llama 3 Herd of Models, 2024. URL: https://arxiv. org/abs/2407.21783.arXiv:2407.21783
2024 arXiv
-
[10]
Chang, H
K. Chang, H. X. Bai, H. Zhou, C. Su, et al., Residual Convolutional Neural Network for the Determination of IDH Status in Low- and High-Grade Gliomas from MR Imaging, Clinical Cancer Research 24 (2018) 1073–1081. doi:10.1158/1078-0432.CCR-17-2236
2018 doi
-
[11]
T. C. Hollon, B. Pandian, A. R. Adapa, E. Urias, A. V. Save, S. S. S. Khalsa, et al., Near real-time intraoperative brain tumor diagnosis using stimulated Raman histology and deep neural networks, Nature Medicine 26 (2020) 52–58. doi:10.1038/s41591-019-0715-9
2020 doi
-
[12]
Hollon, C
T. Hollon, C. Jiang, A. Chowdury, M. Nasir-Moin, A. Kondepudi, et al., Artificial-intelligence-based molecular classification of diffuse gliomas using rapid, label-free optical imaging, Nature Medicine 29 (2023) 828–832. doi:10.1038/s41591-023-02252-4
2023 doi
-
[13]
J. N. Acosta, G. J. Falcone, P. Rajpurkar, E. J. Topol, Multimodal biomedical AI, Nature Medicine 28 (2022) 1773–1784. doi:10.1038/s41591-022-01981-2
2022 doi
-
[14]
Lipkova, R
J. Lipkova, R. J. Chen, B. W. Chen, M. Y. Lu, M. Barbieri, D. Shao, et al., Artificial intelligence for multimodal data integration in oncology, Cancer Cell 40 (2022) 1095–1110. doi:10.1016/j. ccell.2022.09.012
2022 doi
-
[15]
R. J. Chen, M. Y. Lu, D. F. K. Williamson, T. Y. Chen, J. Lipkova, et al., Pan-cancer integrative histology-genomic analysis via multimodal deep learning, Cancer Cell 40 (2022) 865–878.e6. doi:10.1016/j.ccell.2022.07.004
2022 doi
-
[16]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, N. Houlsby, An Image is Worth 16x16 Words: Trans- formers for Image Recognition at Scale, in: International Conference on Learning Re...
2021
-
[17]
R. J. Chen, T. Ding, M. Y. Lu, D. F. K. Williamson, G. Jaume, et al., Towards a general-purpose foundation model for computational pathology, Nature Medicine 30 (2024) 850–862. doi:10.1038/ s41591-024-02857-3
2024
-
[18]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention Is All You Need, in: Advances in Neural Information Processing Systems, volume 30, Curran Associates, Inc., 2017, pp. 5998–6008
2017
-
[19]
M. Y. Lu, B. Chen, D. F. K. Williamson, R. J. Chen, I. Liang, T. Ding, et al., A visual-language foundation model for computational pathology, Nature Medicine 30 (2024) 863–874. doi:10.1038/ s41591-024-02856-4
2024
-
[20]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, et al., Learning Transferable Visual Models From Natural Language Supervision, in: Proceedings of the 38th International Conference on Machine Learning, volume 139 ofProceedings of Machine Learning Research, PML...
2021
-
[21]
Dibble, C
A. Dibble, C. Dalby, M. Sevegnani, A. Fracasso, D. M. Lyall, M. Harvey, M. Svanera, Alzheimer’s Disease Neuroimaging Initiative, Frontotemporal Lobar Degeneration Neuroimaging Initiative, NeuroFM: Toward Precision Neuroimaging with Foundation Models for Individualized Brain He...
2026 doi
-
[22]
D. Tak, B. A. Garomsa, A. Zapaishchykova, T. L. Chaunzwa, J. C. Climent Pardo, et al., A generalizable foundation model for analysis of human brain MRI, Nature Neuroscience (2026). URL: https://www.nature.com/articles/s41593-026-02202-6. doi:10.1038/s41593-026-02202-6
2026 doi
-
[23]
C. Wu, X. Zhang, Y. Zhang, Y. Wang, W. Xie, Towards Generalist Foundation Model for Radiology, arXiv preprint arXiv:2308.02463 (2023). URL: https://arxiv.org/abs/2308.02463
2023 arXiv
-
[24]
Q. Xie, Q. Chen, A. Chen, C. Peng, Y. Hu, F. Lin, X. Peng, et al., Me-llama: Medical foundation large language models for comprehensive text analysis and beyond, npj Digital Medicine 8 (2025) 141. URL: https://www.nature.com/articles/s41746-025-01533-1. doi: 10.1038/ s41746-02...
2025
-
[25]
Singhal, T
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, et al., Toward expert-level medical question answering with large language models, Nature Medicine 31 (2025) 943–950. URL: https://www.nature.com/articles/s41591-024-03423-7. doi:10.1038/s41591-024-03423-7
2025 doi
-
[26]
Havaei, N
M. Havaei, N. Guizard, N. Chapados, Y. Bengio, Hemis: Hetero-modal image segmentation, in: Medical Image Computing and Computer-Assisted Intervention – MICCAI 2016, volume 9901 of Lecture Notes in Computer Science, Springer, 2016, pp. 469–477. URL: https://link.springer.com/ c...
2016 doi
-
[27]
Dorent, S
R. Dorent, S. Joutard, M. Modat, S. Ourselin, T. Vercauteren, Hetero-Modal Variational Encoder-Decoder for Joint Modality Completion and Segmentation, Springer International Publishing, 2019, pp. 74–82. URL: http://dx.doi.org/10.1007/978-3-030-32245-8_9. doi: 10.1007/ 978-3-03...
2019 doi
-
[28]
Y. Wang, Y. Zhang, Y. Liu, Z. Lin, J. Tian, C. Zhong, Z. Shi, J. Fan, Z. He, Acn: Adversarial co-training network for brain tumor segmentation with missing modalities, arXiv preprint arXiv:2106.14591 (2021). URL: https://arxiv.org/abs/2106.14591
2021 arXiv
-
[29]
Y. Ding, X. Yu, Y. Yang, Rfnet: Region-aware fusion network for incomplete multi-modal brain tumor segmentation, in: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021, pp. 3975–3984. URL: https://openaccess.thecvf.com/content/ICCV2021/ papers...
2021
-
[30]
A. H. Abdelaziz, B.-J. Theobald, P. Dixon, R. Knothe, N. Apostoloff, S. Kajareker, Modality dropout for improved performance-driven talking faces, in: Proceedings of the 2020 International Conference on Multimodal Interaction (ICMI), ACM, 2020, pp. 378–386. URL: https://dl.acm...
2020
-
[31]
Y. Gu, K. Saito, J. Ma, Learning contrastive multimodal fusion with improved modality dropout for disease detection and prediction, in: Medical Image Computing and Computer Assisted Intervention – MICCAI 2025, Springer, 2025. URL: https://arxiv.org/abs/2509.18284
2025
-
[32]
A. Yan, J. McAuley, X. Lu, J. Du, E. Y. Chang, A. Gentili, C.-N. Hsu, Radbert: Adapting transformer- based language models to radiology, Radiology: Artificial Intelligence 4 (2022) e210258. doi:10. 1148/ryai.210258
2022
-
[33]
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov, Roberta: A robustly optimized bert pretraining approach, 2019. URL: https://arxiv.org/abs/1907. 11692.arXiv:1907.11692
2019 arXiv
-
[34]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, 2019. URL: https://arxiv.org/abs/1810.04805.arXiv:1810.04805
2019 arXiv
-
[35]
URL: http://www
PACE, Partnership for an Advanced Computing Environment (PACE), 2017. URL: http://www. pace.gatech.edu
2017
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.