REVIEW 4 major objections 4 minor 98 references
Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Visual in-context learners can follow multi-step medical imaging pipelines defined at test time.
desk verdict Compositional in-context learning with a single model is a real and new setting, and sequence-token masking works for generative/transformation tasks, but the discriminative results sit too close to the copy baseline to support the full claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a pipeline of three coupled components. First, a synthetic task-generation engine bootstraps diverse task sequences from 38 segmentation datasets by enriching each image–mask pair with generative tasks (super resolution, inpainting, denoising, color jitter, inversion, equalization, brightness), geometric transformations (flips, rotations), and discriminative outputs (segments, boxes, edges, skeletons, points), then sampling chains under guardrails: generative steps first, transformations second, discriminative tasks last, with class subsampling, random recoloring, parameter sampling, and sequence length capped at 30 images. Second, a task-informed codebook, a vector-quantized autoencoder (VQ-GAN) that maps each $200\times200$ image to 144 discrete tokens from a vocabulary of 16,384, is fine-tuned on the enriched data with data-balancing and random color remapping so that discrete tokens can actually represent task outputs. Third, the in-context learner, a GPT-2 Transformer, is trained with sequence-token masking: all tokens of the output sequence $O$ are replaced and must be predicted from the context–query token sequence, an objective that mirrors exactly what the model must do at test time and that clearly outperforms random token masking and image-token masking.
What would settle it
Train or evaluate on a compositional prompt whose task family or imaging modality is absent from the 38 MedSAM datasets—for example, a novel discriminative output type like cell counting or a pathology whole-slide modality—and check whether performance collapses to the copy baseline. A sharper in-distribution test: build a sequence where a class appears in the query's segmentation output but never in any context example, and verify whether the model still emits that class, which would reveal memorization rather than in-context following.
Extended reading notes
Core claim
The paper's central discovery is that visual in-context learning, previously demonstrated for single tasks, can be extended to compositional task sequences: a single GPT-style transformer, fed a context of example image-to-image step pairs, can generate the entire chain of outputs for a held-out query image across medical modalities. On a test set of 2,000 synthetic compositional tasks built from image and task outputs held out from training, the model with sequence-token masking reaches a segmentation IoU of 61.2% and box F-1 of 65.9%, far above a copy baseline (about 50.8% IoU) but below the codebook upper bound (about 86.7% IoU). The authors further establish that the representation bottleneck sits in the codebook: off-the-shelf VQ-GAN codebooks trained on natural images recover at most 29% IoU on a 142-class anatomy dataset, whereas fine-tuning the codebook on the task outputs themselves, with dataset/task balancing and color-remapping augmentation, raises the upper bound to near-90% IoU on binary and few-class tasks while generalizing to arbitrary class colors. Sequence-token masking—switching out all tokens of the output sequence during training—is identified as the objective that forces the model to learn long-range dependencies from context to output.
Load-bearing premise
The approach rests on the assumption that the synthetically generated task chains—constrained by hand-designed ordering, class subsampling, and recoloring rules—actually resemble the multi-step pipelines a clinician or user would define at test time; evaluation never leaves the 38 training datasets or the task families the generator can produce.
Editorial extensions
If this is right
- A medical user could define a bespoke pipeline—such as artifact removal, orientation correction, then structure segmentation—by showing the model a few example input–output pairs, and run it on new images without retraining or programming.
- The masking comparison establishes a direct recipe: to make an in-context learner solve a test-time task, train it to predict exactly that output region under masking, not generic token corruption.
- Codebook quality is the binding constraint; any attempt to scale this approach must address representation fidelity for fine structures and many-class semantics before the learner itself improves.
- Because intermediate outputs are generated and inspectable step by step, the approach offers a form of interpretability useful in medical imaging, where each pipeline stage can be verified.
- The gap between the best model and the codebook upper bound quantifies the remaining headroom—roughly 25 IoU points on segmentation—indicating where training objectives and data scale must improve.
Reading between the lines
- The guardrail ordering (generative before transformation before discriminative) implicitly encodes a prior on real clinical pipelines—restore, align, then label. A direct test would train on a permuted order and see whether the learners themselves impose the ordering, which would tell us whether the structure is learned or injected.
- The class-subsampling guardrail implies a concrete failure mode worth probing: if a query output contains a class never shown in the context, the model should be unable to produce it; this yields a diagnostic for true context adherence versus dataset-memorized class priors.
- Since the generation engine works off arbitrary segmentation datasets, the same recipe could transfer to non-medical domains; a natural extension is to test whether out-of-domain capability appears at larger data and model scale, as suggested by prior scaling results.
- The color-remapping augmentation points toward a more general principle: in-context learners generalize to unseen output appearances only when the tokenizer itself is invariant to those appearances; training the tokenizer with output-remapping noise may be a general precondition for visual in-context learning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether visual in-context learning can be extended from single tasks to compositional sequences of medical imaging tasks. The authors propose a synthetic task generation engine that bootstraps compositional task sequences from 38 MedSAM segmentation datasets, covering generative, geometric transformation, and discriminative tasks. They analyze codebook reconstruction bounds, propose a task-informed codebook with color remapping augmentation and data balancing, and compare three masking-based training objectives for a GPT-2 style in-context learner. The main quantitative result (Table 1) is that sequence-token masking outperforms token and image-token masking on held-out compositional tasks, with particular gains on generative and transformation tasks; however, discriminative task scores are close to or below a trivial copy baseline. Qualitative examples show coherent multi-step predictions on some sequences, while the appendix documents failures such as plausible but mislocated structures. The paper explicitly acknowledges several limitations, including generalization beyond the training datasets.
Significance. If the main claim were fully supported, this would be a useful first demonstration that a single transformer-based model can be trained to execute multi-step, test-time-defined compositional vision pipelines in the medical domain, with interpretable intermediate outputs. The experimental design is a genuine strength: the sandwich of a copy baseline and a measured codebook upper bound provides a clear frame for isolating the in-context learner's contribution, and the systematic comparison of masking strategies is informative even if the interpretation requires care. The codebook analysis (Section 4 and Table 2) is also valuable, showing that off-the-shelf VQ-GANs severely under-represent fine-grained discriminative outputs and that task-informed fine-tuning substantially raises the upper bound. The paper is honest about limitations in Appendix E, which strengthens its credibility. However, the central claim about mastering discriminative compositional steps is not yet established by the evidence, and the generalization claims are narrower than the abstract suggests.
major comments (4)
- [Table 1; Sec. 5.2.1; Appendix E]
- [Sec. 5.3; Table 1]
- [Sec. 3.1; Sec. 6.4; Appendix E]
- [Sec. 5.3; Sec. 6.4]
minor comments (4)
- [Sec. 6.2]
- [Sec. 5.2.2]
- [Table 1]
- [Appendix D.2]
Circularity Check
No circular derivation: codebook bounds are measured, the training objective matches the test protocol by design but does not feed the output back as input, and the paper's self-citations are not load-bearing.
full rationale
The paper's empirical chain is self-contained. Codebook upper bounds (Table 2 and the last row of Table 1) are obtained by directly encoding and decoding task outputs with the deployed codebook; this is a measured ceiling for any token-space predictor, not a prediction derived from the model. The in-context learner is trained with a masked-token objective (Sec. 5.3) and evaluated by decoding its predicted tokens; the quantity being predicted (O) is never used as an input to the model, so the result does not reduce to its inputs by construction. The sequence-token masking variant matches the test protocol by design, but this is standard supervised training rather than circularity: the model must still generate the full output sequence from the context C and query Q. The copy baseline is an external reference, and the paper explicitly reports that discriminative tasks remain near this baseline, which is a weakness in evidence strength, not a circular step. Self-citations (e.g., ATLAS [45], interactive segmentation taxonomy [57], prior segmentation work [70,71]) are used as data sources and related work; no uniqueness theorem or load-bearing claim rests on them. The main scope caveat, that test tasks are sampled from the same synthetic generation engine as training tasks, is explicitly disclosed in Appendix E as a limitation; it affects external validity, not the internal derivation chain.
Assumptions & free parameters
free parameters (7)
- Sequence length cap =
30 images (up to 15 tasks)
- Masking ratio p =
15%
- Number of END tokens k =
2
- Codebook tokens per image =
144 (VQ-GAN f16, vocabulary 16,384)
- Best training epochs =
52
- Data balance weights =
Equal task outputs vs images; dataset oversampling
- Image resolution =
200x200
assumptions (4)
- domain assumption The VQ-GAN token space preserves enough task-relevant structure for in-context prediction
- domain assumption Synthetic task sequences derived from segmentation annotations are a valid training distribution for compositional medical pipelines
- domain assumption Masked token prediction with a GPT2 transformer can learn the conditional structure of task sequences
- ad hoc to paper The enforced {generative} to {transformation} to {discriminative} ordering and class-subsampling guardrails yield informative tasks
invented entities (1)
-
i-END and c-END special tokens
independent evidence
Cite this review
Pith. "Pith review of Is Visual in-Context Learning for Compositional Medical Tasks within Reach?." pith.science (2026). https://pith.science/paper/7OEOIAPQ
@misc{pith2026250700868,
author = {Pith},
title = {Pith review of: Is Visual in-Context Learning for Compositional Medical Tasks within Reach?},
year = {2026},
howpublished = {\url{https://pith.science/paper/7OEOIAPQ}},
note = {Machine review of arXiv:2507.00868}
}
read the original abstract
In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time. To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Deep learning based automated detection of intraretinal cys- toid fluid
Zeeshan Ahmed, Shahbaz Qamar Panhwar, Attiya Baqai, Fahim Aziz Umrani, Munawar Ahmed, and Arbaaz Khan. Deep learning based automated detection of intraretinal cys- toid fluid. International Journal of Imaging Systems and Technology, 32(3):902–917, 2022. 16
2022
-
[2]
Dataset of breast ultrasound images
Walid Al-Dhabyani, Mohammed Gomaa, Hussien Khaled, and Aly Fahmy. Dataset of breast ultrasound images. Data in brief, 28:104863, 2020. 16
2020
-
[3]
P. An, S. Xu, S. A. Harmon, E. B. Turkbey, T. H. Sanford, A. Amalou, M. Kassin, N. Varble, M. Blain, V . Anderson, G. Patella, F.and Carrafiello, B. T. Turkbey, and B. J. Wood. Ct images in covid-19 [data set]., 2020. 16
2020
-
[4]
The medical segmentation decathlon.Nature communications, 13(1):4128, 2022
Michela Antonelli, Annika Reinke, Spyridon Bakas, Key- van Farahani, Annette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M Summers, et al. The medical segmentation decathlon.Nature communications, 13(1):4128, 2022. 16
2022
-
[5]
Data from lidc-idri [data set]
SG Armato III, G McLennan, L Bidaut, MF McNitt-Gray, CR Meyer, AP Reeves, B Zhao, DR Aberle, CI Henschke, EA Hoffman, et al. Data from lidc-idri [data set]. the cancer imaging archive, 2015. 16
2015
-
[6]
Masked siamese networks for label-efficient learning
Mahmoud Assran, Mathilde Caron, Ishan Misra, Piotr Bo- janowski, Florian Bordes, Pascal Vincent, Armand Joulin, Mike Rabbat, and Nicolas Ballas. Masked siamese networks for label-efficient learning. InEuropean Conference on Com- puter Vision, pages 456–473. Springer, 2022. 1, 6
2022
-
[7]
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. arXiv preprint arXiv:2301.08243, 2023. 6
arXiv 2023
-
[8]
Sequential modeling enables scal- able learning for large vision models
Yutong Bai, Xinyang Geng, Karttikeya Mangalam, Amir Bar, Alan Yuille, Trevor Darrell, Jitendra Malik, and Alexei A Efros. Sequential modeling enables scal- able learning for large vision models. arXiv preprint arXiv:2312.00785, 2023. 1, 2, 5, 6, 18
arXiv 2023
Show all 98 references
-
[9]
Segmentation labels and radiomic features for the pre-operative scans of the tcga- gbm collection (2017)
Spyridon Bakas, Hamed Akbari, Aristeidis Sotiras, Michel Bilello, Martin Rozycki, Justin Kirby, John Freymann, Key- van Farahani, and Christos Davatzikos. Segmentation labels and radiomic features for the pre-operative scans of the tcga- gbm collection (2017). DOI: https://doi...
2017
-
[10]
Segmentation labels and radiomic features for the pre-operative scans of the tcga-lgg collection [data set]
Spyridon Bakas, Hamed Akbari, Aristeidis Sotiras, Michel Bilello, Martin Rozycki, JS Kirby, JB Freymann, Keyvan Farahani, and Christos Davatzikos. Segmentation labels and radiomic features for the pre-operative scans of the tcga-lgg collection [data set]. the cancer imaging ar...
2017
-
[11]
Advancing the cancer genome atlas glioma mri collections with expert seg- mentation labels and radiomic features
Spyridon Bakas, Hamed Akbari, Aristeidis Sotiras, Michel Bilello, Martin Rozycki, Justin S Kirby, John B Freymann, Keyvan Farahani, and Christos Davatzikos. Advancing the cancer genome atlas glioma mri collections with expert seg- mentation labels and radiomic features. Scient...
2017
-
[12]
Identifying the best machine learning algorithms for brain tumor segmentation, progression assessment, and over- all survival prediction in the brats challenge
Spyridon Bakas, Mauricio Reyes, Andras Jakab, Stefan Bauer, Markus Rempfler, Alessandro Crimi, Russell Takeshi Shinohara, Christoph Berger, Sung Min Ha, Martin Rozycki, et al. Identifying the best machine learning algorithms for brain tumor segmentation, progression assessment...
2018 arXiv
-
[13]
Visual prompting via image inpaint- ing
Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Glober- son, and Alexei Efros. Visual prompting via image inpaint- ing. Advances in Neural Information Processing Systems , 35:25005–25017, 2022. 1, 2
2022
-
[14]
Label-efficient se- mantic segmentation with diffusion models
Dmitry Baranchuk, Andrey V oynov, Ivan Rubachev, Valentin Khrulkov, and Artem Babenko. Label-efficient se- mantic segmentation with diffusion models. In International Conference on Learning Representations, 2021. 1
2021
-
[15]
xl- stm: Extended long short-term memory
Maximilian Beck, Korbinian P ¨oppel, Markus Spanring, An- dreas Auer, Oleksandra Prudnikova, Michael Kopp, G ¨unter Klambauer, Johannes Brandstetter, and Sepp Hochreiter. xl- stm: Extended long short-term memory. arXiv preprint arXiv:2405.04517, 2024. 6
2024 arXiv
-
[16]
Retrieval-augmented diffusion models, 2022
Andreas Blattmann, Robin Rombach, Kaan Oktay, and Bj¨orn Ommer. Retrieval-augmented diffusion models, 2022. 14
2022
-
[17]
Retouch: The retinal oct fluid detection and segmenta- tion benchmark and challenge
Hrvoje Bogunovi ´c, Freerk Venhuizen, Sophie Klimscha, Ste- fanos Apostolopoulos, Alireza Bab-Hadiashar, Ulas Bagci, Mirza Faisal Beg, Loza Bekalo, Qiang Chen, Carlos Ciller, et al. Retouch: The retinal oct fluid detection and segmenta- tion benchmark and challenge. IEEE trans...
2019
-
[18]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Sub- biah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakan- tan, Pranav Shyam, Girish Sastry, Amanda Askell, Sand- hini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler...
2005 arXiv
-
[19]
Uni- verseg: Universal medical image segmentation
Victor Ion Butoi, Jose Javier Gonzalez Ortiz, Tianyu Ma, Mert R Sabuncu, John Guttag, and Adrian V Dalca. Uni- verseg: Universal medical image segmentation. arXiv preprint arXiv:2304.06131, 2023. 2
2023 arXiv
-
[20]
Lung segmentation in chest radiographs using anatomical atlases with nonrigid registration
Sema Candemir, Stefan Jaeger, Kannappan Palaniappan, Jonathan P Musco, Rahul K Singh, Zhiyun Xue, Alexandros Karargyris, Sameer Antani, George Thoma, and Clement J McDonald. Lung segmentation in chest radiographs using anatomical atlases with nonrigid registration. IEEE transa...
2013
-
[21]
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T. Freeman. Maskgit: Masked generative image transformer. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
2022
-
[22]
Muse: Text-to-image generation via masked generative transform- ers
Huiwen Chang, Han Zhang, Jarred Barber, AJ Maschinot, Jose Lezama, Lu Jiang, Ming-Hsuan Yang, Kevin Mur- phy, William T Freeman, Michael Rubinstein, et al. Muse: Text-to-image generation via masked generative transform- ers. arXiv preprint arXiv:2301.00704, 2023. 2
2023 arXiv
-
[23]
Can ai help in screening viral and covid-19 pneumonia? Ieee Access, 8:132665–132676, 2020
Muhammad EH Chowdhury, Tawsifur Rahman, Amith Khandakar, Rashid Mazhar, Muhammad Abdul Kadir, Zaid Bin Mahbub, Khandakar Reajul Islam, Muham- mad Salman Khan, Atif Iqbal, Nasser Al Emadi, et al. Can ai help in screening viral and covid-19 pneumonia? Ieee Access, 8:132665–13267...
2020
-
[24]
Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the interna- tional skin imaging collaboration (isic)
Noel Codella, Veronica Rotemberg, Philipp Tschandl, M Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, et al. Skin lesion analysis toward melanoma detection 2018: A challenge hosted by the interna- tional skin imagin...
2018 arXiv
-
[25]
Noel CF Codella, David Gutman, M Emre Celebi, Brian Helba, Michael A Marchetti, Stephen W Dusza, Aadi Kalloo, Konstantinos Liopyris, Nabin Mishra, Harald Kit- tler, et al. Skin lesion analysis toward melanoma detection: A challenge at the 2017 international symposium on biomed...
2017
-
[26]
Neuralizer: General neuroimage analysis without re-training
Steffen Czolbe and Adrian V Dalca. Neuralizer: General neuroimage analysis without re-training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6217–6230, 2023. 2
2023
-
[27]
Vision transformers need registers
Timoth ´ee Darcet, Maxime Oquab, Julien Mairal, and Pi- otr Bojanowski. Vision transformers need registers. arXiv preprint arXiv:2309.16588, 2023. 6
2023 arXiv
-
[28]
Covid-19 in- fection map generation and detection from chest x-ray im- ages
Aysen Degerli, Mete Ahishali, Mehmet Yamac, Serkan Ki- ranyaz, Muhammad EH Chowdhury, Khalid Hameed, Tahir Hamid, Rashid Mazhar, and Moncef Gabbouj. Covid-19 in- fection map generation and detection from chest x-ray im- ages. Health information science and systems, 9(1):15, 2021. 16
2021
-
[29]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 7
2009
-
[31]
BERT: pre-training of deep bidirectional trans- formers for language understanding.CoRR, abs/1810.04805,
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional trans- formers for language understanding.CoRR, abs/1810.04805,
-
[32]
Cross- moda 2021 challenge: Benchmark of cross-modality do- main adaptation techniques for vestibular schwannoma and cochlea segmentation
Reuben Dorent, Aaron Kujawa, Marina Ivory, Spyridon Bakas, Nicola Rieke, Samuel Joutard, Ben Glocker, Jorge Cardoso, Marc Modat, Kayhan Batmanghelich, et al. Cross- moda 2021 challenge: Benchmark of cross-modality do- main adaptation techniques for vestibular schwannoma and co...
2021
-
[33]
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer. Taming transformers for high-resolution image synthesis. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12873–12883, 2021. 2, 3, 6, 7
2021
-
[34]
An annotated test-retest collection of prostate mul- tiparametric mri
Andriy Fedorov, Michael Schwier, David Clunie, Christian Herz, Steve Pieper, Ron Kikinis, Clare Tempany, and Fiona Fennessy. An annotated test-retest collection of prostate mul- tiparametric mri. Scientific data, 5(1):1–13, 2018. 16
2018
-
[35]
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. Mamba: Linear-time sequence modeling with selective state spaces. arXiv preprint arXiv:2312.00752, 2023. 6
2023 arXiv
-
[36]
Analogist: Out-of-the-box visual in-context learning with image diffusion model
Zheng Gu, Shiyuan Yang, Jing Liao, Jing Huo, and Yang Gao. Analogist: Out-of-the-box visual in-context learning with image diffusion model. ACM Transactions on Graphics (TOG), 43(4):1–15, 2024. 2
2024
-
[37]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. 8
2025 arXiv
-
[38]
Data-efficient large vision models through sequential autoregression.arXiv preprint arXiv:2402.04841, 2024
Jianyuan Guo, Zhiwei Hao, Chengcheng Wang, Yehui Tang, Han Wu, Han Hu, Kai Han, and Chang Xu. Data-efficient large vision models through sequential autoregression.arXiv preprint arXiv:2402.04841, 2024. 2, 5, 6
2024 arXiv
-
[39]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 6
2022
-
[40]
Whole-cell organelle segmentation in volume electron microscopy
Larissa Heinrich, Davis Bennett, David Ackerman, Woohyun Park, John Bogovic, Nils Eckstein, Alyson Petruncio, Jody Clements, Song Pang, C Shan Xu, et al. Whole-cell organelle segmentation in volume electron microscopy. Nature, 599(7883):141–146, 2021. 3, 4, 14
2021
-
[41]
Isles 2022: A multi-center magnetic resonance imag- 10 ing stroke lesion segmentation dataset
Moritz R Hernandez Petzsche, Ezequiel de la Rosa, Uta Hanning, Roland Wiest, Waldo Valenzuela, Mauricio Reyes, Maria Meyer, Sook-Lei Liew, Florian Kofler, Ivan Ezhov, et al. Isles 2022: A multi-center magnetic resonance imag- 10 ing stroke lesion segmentation dataset. Scientif...
2022
-
[42]
Automatic lung segmentation in routine imaging is primarily a data di- versity problem, not a methodology problem
Johannes Hofmanninger, Forian Prayer, Jeanny Pan, Sebas- tian R ¨ohrich, Helmut Prosch, and Georg Langs. Automatic lung segmentation in routine imaging is primarily a data di- versity problem, not a methodology problem. European ra- diology experimental, 4:1–13, 2020. 16
2020
-
[43]
Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80
W-Y Hong, C-L Kao, Y-H Kuo, J-R Wang, W-L Chang, and C-S Shih. Cholecseg8k: a semantic segmentation dataset for laparoscopic cholecystectomy based on cholec80. arXiv preprint arXiv:2012.12453, 2020. 16
2012 arXiv
-
[44]
Automatic tuberculosis screening using chest radio- graphs
Stefan Jaeger, Alexandros Karargyris, Sema Candemir, Les Folio, Jenifer Siegelman, Fiona Callaghan, Zhiyun Xue, Kannappan Palaniappan, Rahul K Singh, Sameer Antani, et al. Automatic tuberculosis screening using chest radio- graphs. IEEE transactions on medical imaging , 33(2):...
2013
-
[45]
Towards unifying anatomy segmentation: Automated generation of a full-body ct dataset
Alexander Jaus, Constantin Seibold, Kelsey Hermann, Ne- gar Shahamiri, Alexandra Walter, Kristina Giske, Johannes Haubold, Jens Kleesiek, and Rainer Stiefelhagen. Towards unifying anatomy segmentation: Automated generation of a full-body ct dataset. In 2024 IEEE International ...
2024
-
[46]
Kvasir-seg: A segmented polyp dataset
Debesh Jha, Pia H Smedsrud, Michael A Riegler, P ˚al Halvorsen, Thomas De Lange, Dag Johansen, and H˚avard D Johansen. Kvasir-seg: A segmented polyp dataset. In In- ternational conference on multimedia modeling, pages 451–
-
[47]
Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation
Yuanfeng Ji, Haotian Bai, Jie Yang, Chongjian Ge, Ye Zhu, Ruimao Zhang, Zhen Li, Lingyan Zhang, Wanling Ma, Xi- ang Wan, et al. Amos: A large-scale abdominal multi-organ benchmark for versatile medical image segmentation. arXiv preprint arXiv:2206.08023, 2022. 16
2022 arXiv
-
[48]
Categorized digital database for low en- ergy and subtracted contrast enhanced spectral mammogra- phy images
R Khaled et al. Categorized digital database for low en- ergy and subtracted contrast enhanced spectral mammogra- phy images. The Cancer Imaging Archive, 2021. 16
2021
-
[49]
Standardized assessment of automatic segmentation of white matter hyperintensities and results of the wmh segmentation challenge
Hugo J Kuijf, J Matthijs Biesbroek, Jeroen De Bresser, Rut- ger Heinen, Simon Andermatt, Mariana Bento, Matt Berseth, Mikhail Belyaev, M Jorge Cardoso, Adria Casamitjana, et al. Standardized assessment of automatic segmentation of white matter hyperintensities and results of t...
2019
-
[50]
Decoupled con- text processing for context augmented language modeling
Zonglin Li, Ruiqi Guo, and Sanjiv Kumar. Decoupled con- text processing for context augmented language modeling. Advances in Neural Information Processing Systems , 35: 21698–21710, 2022. 1
2022
-
[51]
Unified-io: A unified model for vision, language, and multi-modal tasks
Jiasen Lu, Christopher Clark, Rowan Zellers, Roozbeh Mot- taghi, and Aniruddha Kembhavi. Unified-io: A unified model for vision, language, and multi-modal tasks. arXiv preprint arXiv:2206.08916, 2022. 2
2022 arXiv
-
[52]
Toward data-efficient learning: A benchmark for covid- 19 ct lung and infection segmentation
Jun Ma, Yixin Wang, Xingle An, Cheng Ge, Ziqi Yu, Jianan Chen, Qiongjie Zhu, Guoqiang Dong, Jian He, Zhiqiang He, et al. Toward data-efficient learning: A benchmark for covid- 19 ct lung and infection segmentation. Medical physics, 48 (3):1197–1210, 2021. 16
2021
-
[53]
Fast and low-gpu-memory abdomen ct organ seg- mentation: the flare challenge
Jun Ma, Yao Zhang, Song Gu, Xingle An, Zhihe Wang, Cheng Ge, Congcong Wang, Fan Zhang, Yu Wang, Yinan Xu, et al. Fast and low-gpu-memory abdomen ct organ seg- mentation: the flare challenge. Medical Image Analysis, 82: 102616, 2022. 16
2022
-
[54]
Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):6695–6714, 2022
Jun Ma, Yao Zhang, Song Gu, Cheng Zhu, Cheng Ge, Yichi Zhang, Xingle An, Congcong Wang, Qiyuan Wang, Xin Liu, Shucheng Cao, Qi Zhang, Shangqing Liu, Yunpeng Wang, Yuhui Li, Jian He, and Xiaoping Yang. Abdomenct-1k: Is abdominal organ segmentation a solved problem? IEEE Transac...
2022
-
[55]
Segment anything in medical images
Jun Ma, Yuting He, Feifei Li, Lin Han, Chenyu You, and Bo Wang. Segment anything in medical images. Nature Communications, 15(1):654, 2024. 5, 6, 14, 15
2024
-
[56]
m2caiseg: Semantic segmentation of laparoscopic images using convolutional neural networks
Salman Maqbool, Aqsa Riaz, Hasan Sajid, and Osman Hasan. m2caiseg: Semantic segmentation of laparoscopic images using convolutional neural networks. arXiv preprint arXiv:2008.10134, 2020. 16
2008 arXiv
-
[57]
Deep interactive segmentation of medical images: A systematic review and taxonomy
Zdravko Marinov, Paul F J ¨ager, Jan Egger, Jens Kleesiek, and Rainer Stiefelhagen. Deep interactive segmentation of medical images: A systematic review and taxonomy. IEEE transactions on pattern analysis and machine intelligence ,
-
[58]
Martin, C
D. Martin, C. Fowlkes, D. Tal, and J. Malik. A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecologi- cal statistics. In Proc. 8th Int’l Conf. Computer Vision, pages 416–423, 2001. 3, 14
2001
-
[59]
Data structures for statistical comput- ing in python
Wes McKinney et al. Data structures for statistical comput- ing in python. In Proceedings of the 9th Python in Science Conference, pages 51–56. Austin, TX, 2010. 20
2010
-
[60]
Factors of influence for transfer learning across diverse appearance domains and task types
Thomas Mensink, Jasper Uijlings, Alina Kuznetsova, Michael Gygli, and Vittorio Ferrari. Factors of influence for transfer learning across diverse appearance domains and task types. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12):9298–9314, 2021. 1
2021
-
[61]
The multimodal brain tumor image segmentation benchmark (brats)
Bjoern H Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, et al. The multimodal brain tumor image segmentation benchmark (brats). IEEE transactions on medical imaging , 34(...
1993
-
[62]
Supervised transfer learning at scale for medical imaging
Basil Mustafa, Aaron Loh, Jan Freyberg, Patricia MacWilliams, Megan Wilson, Scott Mayer McKinney, Marcin Sieniek, Jim Winkens, Yuan Liu, Peggy Bui, et al. Supervised transfer learning at scale for medical imaging. arXiv preprint arXiv:2101.05913, 2021. 1
2021 arXiv
-
[63]
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zem- ing Lin, Natalia Gimelshein, Luca Antiga, Alban Desmai- son, Andreas Kopf, Edward Yang, Zachary DeVito, Mar- tin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steine...
2019
-
[64]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V . Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V . Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learni...
2011
-
[65]
Kvasir: A multi- class image dataset for computer aided gastrointestinal dis- ease detection
Konstantin Pogorelov, Kristin Ranheim Randel, Carsten Gri- wodz, Sigrun Losada Eskeland, Thomas de Lange, Dag Johansen, Concetto Spampinato, Duc-Tien Dang-Nguyen, Mathias Lux, Peter Thelin Schmidt, et al. Kvasir: A multi- class image dataset for computer aided gastrointestinal...
2017
-
[66]
Language models are unsu- pervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsu- pervised multitask learners. OpenAI blog, 1(8):9, 2019. 2, 6, 7
2019
-
[67]
Exploring the effect of image enhancement techniques on covid-19 detection using chest x-ray images.Computers in biology and medicine, 132: 104319, 2021
Tawsifur Rahman, Amith Khandakar, Yazan Qiblawey, Anas Tahir, Serkan Kiranyaz, Saad Bin Abul Kashem, Moham- mad Tariqul Islam, Somaya Al Maadeed, Susu M Zughaier, Muhammad Salman Khan, et al. Exploring the effect of image enhancement techniques on covid-19 detection using ches...
2021
-
[68]
Tyche: Stochastic in-context learning for medical image segmenta- tion
Marianne Rakic, Hallee E Wong, Jose Javier Gonzalez Ortiz, Beth A Cimini, John V Guttag, and Adrian V Dalca. Tyche: Stochastic in-context learning for medical image segmenta- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 1...
-
[69]
Every annotation counts: Multi-label deep supervision for medical image segmen- tation
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, and Rainer Stiefelhagen. Every annotation counts: Multi-label deep supervision for medical image segmen- tation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9532–9542,
-
[70]
Graph-constrained con- trastive regularization for semi-weakly volumetric segmen- tation
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, and Rainer Stiefelhagen. Graph-constrained con- trastive regularization for semi-weakly volumetric segmen- tation. In European Conference on Computer Vision, pages 401–419. Springer, 2022. 1
2022
-
[71]
Decoupled semantic proto- types enable learning from diverse annotation types for semi- weakly segmentation in expert-driven domains
Simon Reiß, Constantin Seibold, Alexander Freytag, Erik Rodner, and Rainer Stiefelhagen. Decoupled semantic proto- types enable learning from diverse annotation types for semi- weakly segmentation in expert-driven domains. In Proceed- ings of the IEEE/CVF Conference on Compute...
2023
-
[72]
Medi- cal vision generalist: Unifying medical imaging tasks in con- text
Sucheng Ren, Xiaoke Huang, Xianhang Li, Junfei Xiao, Jieru Mei, Zeyu Wang, Alan Yuille, and Yuyin Zhou. Medi- cal vision generalist: Unifying medical imaging tasks in con- text. arXiv preprint arXiv:2406.05565, 2024. 2
2024 arXiv
-
[73]
Ct-org, a new dataset for mul- tiple organ segmentation in computed tomography.Scientific Data, 7(1):381, 2020
Blaine Rister, Darvin Yi, Kaushik Shivakumar, Tomomi Nobashi, and Daniel L Rubin. Ct-org, a new dataset for mul- tiple organ segmentation in computed tomography.Scientific Data, 7(1):381, 2020. 16
2020
-
[74]
High-resolution image syn- thesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models, 2021. 3, 14
2021
-
[75]
Rapid artificial intelli- gence solutions in a pandemic—the covid-19-20 lung ct le- sion segmentation challenge
Holger R Roth, Ziyue Xu, Carlos Tor-D ´ıez, Ramon Sanchez Jacob, Jonathan Zember, Jose Molto, Wenqi Li, Sheng Xu, Baris Turkbey, Evrim Turkbey, et al. Rapid artificial intelli- gence solutions in a pandemic—the covid-19-20 lung ct le- sion segmentation challenge. Medical image...
2022
-
[76]
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR, abs/1910.01108,
1910 arXiv
-
[77]
Amp: Adaptive masked proxies for few-shot segmentation
Mennatullah Siam, Boris N Oreshkin, and Martin Jagersand. Amp: Adaptive masked proxies for few-shot segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5249–5258, 2019. 1
2019
-
[78]
A large annotated medical image dataset for the development and evaluation of segmentation algo- rithms
Amber L Simpson, Michela Antonelli, Spyridon Bakas, Michel Bilello, Keyvan Farahani, Bram Van Ginneken, An- nette Kopp-Schneider, Bennett A Landman, Geert Litjens, Bjoern Menze, et al. A large annotated medical image dataset for the development and evaluation of segmentation a...
1902 arXiv
-
[79]
Exploring effective factors for improving visual in-context learning
Yanpeng Sun, Qiang Chen, Jian Wang, Jingdong Wang, and Zechao Li. Exploring effective factors for improving visual in-context learning. arXiv preprint arXiv:2304.04748, 2023. 2, 4
2023
-
[80]
Learning to compare: Re- lation network for few-shot learning
Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip HS Torr, and Timothy M Hospedales. Learning to compare: Re- lation network for few-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 1199–1208, 2018. 1
2018
-
[81]
Covid-19 infection localization and severity grading from chest x-ray images
Anas M Tahir, Muhammad EH Chowdhury, Amith Khan- dakar, Tawsifur Rahman, Yazan Qiblawey, Uzair Khurshid, Serkan Kiranyaz, Nabil Ibtehaz, M Sohel Rahman, Somaya Al-Maadeed, et al. Covid-19 infection localization and severity grading from chest x-ray images. Computers in bi- olo...
2021
-
[82]
Tahir, Muhammad E
Anas M. Tahir, Muhammad E. H. Chowdhury, Yazan Qi- blawey, Amith Khandakar, Tawsifur Rahman, Serkan Ki- ranyaz, Uzair Khurshid, Nabil Ibtehaz, Sakib Mahmud, and Maymouna Ezeddin. Covid-qu-ex dataset, 2022. 16
2022
-
[83]
The ham10000 dataset, a large collection of multi-source der- matoscopic images of common pigmented skin lesions
Philipp Tschandl, Cliff Rosendahl, and Harald Kittler. The ham10000 dataset, a large collection of multi-source der- matoscopic images of common pigmented skin lesions. Sci- entific data, 5(1):1–9, 2018. 3, 4, 14, 16
2018
-
[84]
En- donet: a deep architecture for recognition tasks on laparo- scopic videos
Andru P Twinanda, Sherif Shehata, Didier Mutter, Jacques Marescaux, Michel De Mathelin, and Nicolas Padoy. En- donet: a deep architecture for recognition tasks on laparo- scopic videos. IEEE transactions on medical imaging , 36 (1):86–97, 2016. 16
2016
-
[85]
Automated measurement of fetal head circumference using 2d ultrasound images
Thomas LA van den Heuvel, Dagmar de Bruijn, Chris L de Korte, and Bram van Ginneken. Automated measurement of fetal head circumference using 2d ultrasound images. PloS one, 13(8):e0200412, 2018. 16
2018
-
[86]
Neural discrete representation learning
Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information pro- cessing systems, 30, 2017. 2, 3
2017
-
[87]
Gomez, Lukasz Kaiser, and Illia 12 Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia 12 Polosukhin. Attention is all you need. In Advances in Neu- ral Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2...
2017
-
[88]
Images speak in images: A generalist painter for in-context visual learning
Xinlong Wang, Wen Wang, Yue Cao, Chunhua Shen, and Tiejun Huang. Images speak in images: A generalist painter for in-context visual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6830–6839, 2023. 1, 2
2023
-
[89]
Seggpt: Segmenting ev- erything in context
Xinlong Wang, Xiaosong Zhang, Yue Cao, Wen Wang, Chunhua Shen, and Tiejun Huang. Seggpt: Segmenting ev- erything in context. arXiv preprint arXiv:2304.03284, 2023. 2
2023 arXiv
-
[90]
To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images
Jakob Wasserthal, Hanns-Christian Breit, Manfred T Meyer, Maurice Pradella, Daniel Hinck, Alexander W Sauter, Tobias Heye, Daniel T Boll, Joshy Cyriac, Shan Yang, et al. To- talsegmentator: robust segmentation of 104 anatomic struc- tures in ct images. Radiology: Artificial In...
2023
-
[91]
Thinking llms: General instruction following with thought generation, 2024a
Tianhao Wu, Janice Lan, Weizhe Yuan, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar. Thinking llms: General instruction following with thought generation, 2024a. URL https://arxiv. org/abs/2410.10630, 2023. 8
2023 arXiv
-
[92]
Feature generating networks for zero-shot learning
Yongqin Xian, Tobias Lorenz, Bernt Schiele, and Zeynep Akata. Feature generating networks for zero-shot learning. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5542–5551, 2018. 1
2018
-
[93]
J. Yang, G. Sharp, H. Veeraraghavan, W. Van Elmpt, A. Dekker, T. Lustberg, and M. Gooding. Data from lung ct segmentation challenge (lctsc) (version 3) [data set]., 2017. 16
2017
-
[94]
Imagebrush: Learning visual in-context instructions for exemplar-based image ma- nipulation
Yifan Yang, Houwen Peng, Yifei Shen, Yuqing Yang, Han Hu, Lili Qiu, Hideki Koike, et al. Imagebrush: Learning visual in-context instructions for exemplar-based image ma- nipulation. Advances in Neural Information Processing Sys- tems, 36, 2024. 2
2024
-
[95]
An image is worth 32 tokens for reconstruction and generation
Qihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen, Daniel Cremers, and Liang-Chieh Chen. An image is worth 32 tokens for reconstruction and generation. arXiv preprint arXiv:2406.07550, 2024. 2
2024 arXiv
-
[96]
Siim-acr pneumothorax segmen- tation
Anna Zawacki, Carol Wu, George Shih, Julia Elliott, Mikhail Fomitchev, Mohannad Hussain, ParasLakhani, Phil Culli- ton, and Shunxing Bao. Siim-acr pneumothorax segmen- tation. https : / / kaggle . com / competitions / siim- acr- pneumothorax- segmentation, 2019. Kaggle. 16
2019
-
[97]
Video in-context learning
Wentao Zhang, Junliang Guo, Tianyu He, Li Zhao, Linli Xu, and Jiang Bian. Video in-context learning. arXiv preprint arXiv:2407.07356, 2024. 2
2024 arXiv
-
[98]
What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36:17773–17794,
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. What makes good examples for visual in-context learning? Advances in Neural Information Processing Systems, 36:17773–17794,
-
[99]
Robust detection and segmentation for diagnosis of vertebral diseases using routine mr images
D ˇzenan Zuki ´c, Ale ˇs Vlas ´ak, Jan Egger, Daniel Ho ˇr´ınek, Christopher Nimsky, and Andreas Kolb. Robust detection and segmentation for diagnosis of vertebral diseases using routine mr images. In Computer Graphics Forum , pages 190–204. Wiley Online Library, 2014. 16 13 g...
2014
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.