Pith. sign in

REVIEW 4 major objections 5 minor 65 references

Source-dataset selection in transfer learning is a social and intuitive process, not a systematic one.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

arxiv 2510.00902 v2 pith:U6LR7OWD submitted 2025-10-01 cs.CV cs.CYcs.HC

Intuitions of Machine Learning Researchers about Transfer Learning for Medical Image Classification

classification cs.CV cs.CYcs.HC
keywords transfer learningsource dataset selectionmedical image classificationpractitioner intuitionhuman-computer interactionmixed-methods surveysimilarity heuristicscommunity practices
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to show that when machine-learning researchers choose which dataset to pretrain on for medical-image classification, they are guided more by intuition, community habits, and vague notions of similarity than by systematic, measurable criteria. The authors support this with a task-based survey of 15 practitioners who, for two very different target tasks (colorectal tissue patches and chest X-rays) and the same three candidate source datasets, mostly justified choices through established baselines, the availability of pretrained models, personal experience, and perceived 'domain similarity' or 'image quality'—terms they invoked without precise definitions. Their quantitative and qualitative results challenge the conventional rule that 'more similar is better': similarity ratings and expected fine-tuning performance were not consistently aligned, and fairness considerations were almost absent from the reasoning. A sympathetic reader should care because source selection is a high-stakes, under-documented decision that affects the generalizability and equity of medical AI; if this description of practice is right, better tooling and clearer conceptual definitions could shift the field from intuition toward more deliberate, transparent choices.

Core claim

The central claim is that source-dataset selection for transfer learning in medical imaging is not a rational, evidence-driven optimization but a situated practice in which community influence (what colleagues use, what reviewers expect, what baselines are standard), practical dataset attributes (size, ease of use, pretrained-model availability), and perceived similarity—semantic, visual, or in the learned embedding space—jointly determine choices. The study's paired case-study design shows this concretely: switching the target from H&E tissue patches to chest X-rays moved practitioners' preference toward the radiological source, yet their similarity ratings did not consistently track their

What carries the argument

The argument is carried by a task-based survey combined with an interactive dataset browser. Each participant was presented with two deliberately different target tasks (a nine-class colorectal H&E patch classification and a multi-label chest X-ray classification) against the same three candidate pretraining sources (ImageNet-1K, RadImageNet, and Ecoset), and had to rate willingness, expected fine-tuning performance, and the expected effects of pretraining across six dimensions (domain, visual, and embedding similarity; dataset scale; fairness; robustness). This repeated-measures design isolates how the change in target task alters choices, while a qualitative content analysis guided by a th

Load-bearing premise

The findings rest on the answers of 15 self-selected practitioners recruited through the authors' own networks; if these volunteers' hypothetical judgments are not representative of how machine-learning researchers at large choose source datasets, the paper's general claims about researcher intuition do not follow.

What would settle it

A preregistered replication with a much larger, more diverse sample that tracks real source-dataset choices in actual medical-imaging transfer-learning projects, and measures whether practitioners' similarity ratings actually predict fine-tuning performance, would settle whether these patterns are robust or artifacts of the small, networked sample.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If source choice is task-dependent and community-influenced, then guidance for transfer learning must be contextual rather than a one-size-fits-all recommendation.
  • Because availability of pretrained models and established baselines strongly shape choice, new domain-specific pretraining datasets need to become community standards before they displace generic ones like ImageNet.
  • Since perceived similarity and expected performance are not always aligned, researchers should empirically validate transferability rather than trusting intuition that 'more similar is better.'
  • Because fairness and robustness were rarely linked to source choice in the responses, dataset documentation and reporting norms should make these dimensions explicit when justifying source selection.
  • The pervasive vague use of terms like 'domain gap' and 'good image quality' indicates that operational definitions and interactive tools are needed to make these concepts usable in practice.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • A direct behavioral test—logging real researchers' source-dataset choices in actual projects and comparing them with their stated heuristics—would separate what practitioners say they do from what they do.
  • The similarity-performance misalignment might be explained by a distinction between visual surface similarity (color, texture, shape) and task-relevant feature reuse; a controlled manipulation of these cues could identify which notion is actually driving expectations.
  • If community influence is a genuine driver, then publication and review norms are direct levers on transfer-learning practice: requiring authors to justify source choice with defined criteria could reshape behavior more effectively than new algorithms.
  • The near-total absence of fairness reasoning in source selection suggests that pretraining biases can propagate silently, since practitioners do not frame source choice as an equity-relevant decision.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper reports a mixed-methods survey of 15 ML practitioners on how they select source datasets for transfer learning in medical image classification. Participants answered questions about a recent transfer-learning project and two controlled case studies (H&E patch classification and chest X-ray classification) with three candidate source datasets (ImageNet-1K, RadImageNet, Ecoset). The authors find that choices are task-dependent and shaped by community practices, dataset attributes, and perceived visual/semantic similarity, and they claim that similarity ratings and expected performance are not always aligned, challenging the 'more similar is better' view. The qualitative analysis identifies community influence, source-dataset attributes, and source-target similarity as central themes, alongside frequent vague use of terms like 'domain gap'. The paper is positioned as an HCI contribution to making tacit knowledge in transfer learning explicit.

Significance. If the findings are credited, the paper makes a useful contribution by shifting attention from purely technical transferability metrics to the social and intuitive processes behind source dataset selection, an underexplored area in HCI and medical imaging. The mixed-methods design, including the interactive dataset browser and the transparent reporting of the codebook, is a strength. The paper also makes a concrete, falsifiable claim—that perceived similarity and expected performance can decouple—which is important for the design of decision-support tools for transfer learning. However, the small, network-recruited sample and the statistical fragility of the core 'not aligned' claim limit the strength of the conclusions as currently stated.

major comments (4)
  1. [§5.3 and Abstract] The central claim that similarity ratings and expected performance are 'not always aligned', challenging 'more similar is better', is not securely supported by the reported Spearman correlations. For RadImageNet in the H&E case all similarity correlations are |ρ|≤0.3, and in the chest X-ray case all are |ρ|≤0.2 (§5.1, Fig. 5). With N=15, the 95% CI for ρ=0.3 is approximately (−0.24, 0.70), so the data are compatible with no association and with a moderate positive association. Range restriction—participants nearly uniformly rating RadImageNet as similar and high-performing—can attenuate ρ regardless of the true relationship. The disclaimer in §4.4 that p-values are 'for completeness' does not address the interpretive problem, because §5.3 treats near-zero correlations as substantive evidence. Please report confidence intervals or individual-level trajectories, and either soften the abstr
  2. [§4.3, §6.1] The paper's abstract and discussion make general claims about 'researchers' ('choices are task-dependent and influenced by community practices...'). The evidence base is 15 self-selected participants recruited through the authors' professional networks, including direct email invitations to researchers who had previously engaged with the authors' work. This creates a substantial risk that the sample is biased toward people familiar with the authors' perspective or with above-average interest in medical imaging transfer learning. A sample of this size and recruitment strategy is appropriate for an exploratory qualitative study, but the generalizing language in the abstract and conclusions overreaches. Please explicitly frame the study as exploratory and revise the abstract and §6 to avoid implying population-level conclusions.
  3. [§4.5] The qualitative coding follows directed content analysis, but no inter-coder reliability statistic (e.g., Cohen's κ or percentage agreement) is reported. Because the codebook was built from the same literature that informed the questionnaire, the coding process risks confirming the authors' prior taxonomy rather than independently discovering participants' categories. The paper would be strengthened by reporting agreement scores and by describing how many codes emerged inductively versus were imposed a priori. This matters for the qualitative claims about community influence and vagueness, which are otherwise plausible but not quantifiably reliable.
  4. [§5.1] The interpretation of the Friedman and Wilcoxon results is inconsistent with the reported significance levels. In case study 2, the Friedman test is significant, but for the key RadImageNet vs. ImageNet-1K contrast the paper reports only 'paired difference is positive (W=3.0, r=0.8)' without a p-value, while earlier it notes the across-case RadImageNet shift has 'adjusted p=0.2'. The text then states 'the ordering is RadImageNet > ImageNet-1K ≈ Ecoset' as if this were established. Please report all p-values and effect sizes with confidence intervals, and use language proportional to the evidence, e.g., 'tendency' rather than 'ordering' where results are not significant.
minor comments (5)
  1. [§4.3] 'Table??' is a broken cross-reference; the demographics table is not properly cited. Please fix.
  2. [§5.1] Figures 3–5 are hard to read in the current rendering and the caption for Fig. 5 does not explain the radar chart axes clearly. Please ensure high-resolution figures and clarify the meaning of the scale in the caption.
  3. [§4.4] The text says 'we report the types of statistical significance tests for completeness, rather than basing our conclusions on the (here not reported) p-values' but then uses language like 'the only signal is a hint for robustness' (§5.1). Please either report the p-values in the text or consistently avoid interpreting non-significant correlations as signals.
  4. [§4.2] The target dataset is referred to as 'CRC-VAL-HE-7K' but reference [17] is titled 'NCT-CRC-HE'. Please reconcile the naming and ensure the dataset version/identifier is accurate.
  5. [Appendix A] The appendix tables are dense and the checkmark categories are not aligned precisely with the quoted text in some rows (e.g., Table A6, Raghu et al.). Consider using a column for the quote's relevant factor rather than global checkmarks.

Circularity Check

0 steps flagged

No significant circularity: the survey results are empirical and not derived from their own inputs.

full rationale

The paper is a task-based survey of 15 ML practitioners; its central empirical findings (task-dependent source choice, influence of community practice, ambiguous use of 'domain gap'/'good image quality', and the non-alignment between similarity ratings and expected performance) are grounded in participant responses rather than in a fitted parameter, a self-citation chain, or an equation that reproduces its own input. The taxonomy of transfer learning factors (Table 1, Appendix A) was used to design the questionnaire and as the entry point for directed content analysis (Section 4.5), which is a recognized qualitative method; the categories found are not 'predictions', and the paper explicitly leaves room for emergent codes and indeed reports community influence as a novel factor. The 'more similar is better' challenge is a descriptive correlation between two subjective rating scales (Section 4.4, Section 5.1, Fig. 5), not a derivation. Self-citations ([21], [22], [23], [57], [63]) are ancillary and are accompanied by external references where they are used for domain-specific claims. Concerns about statistical power (N=15) or range restriction are validity or correctness risks, not circularity.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

No numbers are fitted to produce the qualitative findings; the free parameters listed are design choices in the case studies and exclusion rules. The central claim rests on the domain assumptions above about self-report, coding reliability, and representativeness.

free parameters (3)
  • CheXpert label-inclusion threshold = classes with >=100 images retained
    Author-chosen cutoff for the Case Study 2 target subset; affects task composition and therefore source-choice intuitions.
  • Case-study fine-tune subset sizes = 250 patches/class (H&E); 50 images/class (CheXpert)
    Author-chosen training/validation splits define the low-data scenario presented to participants.
  • Participant exclusion rule = one participant removed
    Author decision to exclude a participant with an impossible years-of-experience entry and repeated identical open answers; no sensitivity analysis reported for this exclusion.
axioms (4)
  • domain assumption Self-reported hypothetical judgments reflect real source-selection behavior
    Participants were asked what they would do in hypothetical case studies; no behavioral or historical validation is provided.
  • domain assumption Two-author coding without reported inter-rater reliability captures reliable themes
    The directed content analysis used two coders but no agreement statistics are reported, so coding reliability is unverified.
  • domain assumption Recruitment via professional networks yields a representative sample
    Participants were self-selected from the authors' networks, Slack channels, and direct invitations to researchers who had engaged with their work (§4.3), so generalizability to all ML researchers is assumed.
  • domain assumption Participants complied with the instruction not to use web search or AI tools
    The authors instructed participants to answer from intuition only, but compliance was not verified.

pith-pipeline@v1.3.0-alltime-deepseek · 25037 in / 10412 out tokens · 144641 ms · 2026-08-04T12:59:31.349486+00:00 · methodology

0 comments
read the original abstract

Transfer learning is crucial for medical imaging, yet the selection of source datasets often relies on researchers' intuition rather than systematic principles, which can impact the generalizability of algorithms and, thus, patient outcomes. This study investigates these decisions through a task-based survey with machine learning practitioners. Unlike prior work that benchmarks models and experimental setups, we take a human-computer interaction (HCI) perspective on how practitioners select source datasets. Our findings indicate that choices are task-dependent and influenced by community practices, dataset properties, and computational (data embedding), or perceived visual or semantic similarity. However, similarity ratings and expected performance are not always aligned, challenging a traditional "more similar is better" view. Moreover, ethical and fairness considerations remain largely absent from source dataset sections. Participants often used ambiguous terminology, which suggests a need for clearer definitions and tools to make them explicit and usable. By clarifying these heuristics and introducing a conceptual framework of transfer learning factors, this work provides practical insights for more systematic source selection in transfer learning.

Figures

Figures reproduced from arXiv: 2510.00902 by Amelia Jim\'enez-S\'anchez, Hubert Dariusz Zaj\k{a}c, Veronika Cheplygina, Yucheng Lu.

Figure 1
Figure 1. Figure 1: Screenshot of our interactive dataset browser. [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Areas of research expertise among participants. [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Participants’ willingness to use different source datasets. (a) Case study 1: H&E patch classification. (b) Case study 2: [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Participants’ subjective assessment of the expected fine-tuning performance for each source dataset. (a) Case study 1: [PITH_FULL_IMAGE:figures/full_fig_p014_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ratings of expected pretraining effects for a successful fine-tuning outcome presented by a 5-point scale (1 = very poor, 5 = [PITH_FULL_IMAGE:figures/full_fig_p015_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 4 canonical work pages · 1 internal anchor

  1. [1]

    Adriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, and Marisol Wong-Villacres. 2025. Emerging Data Practices: Data Work in the Era of Large Language Models. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY , USA, 1–21. doi:10.1145/3706598.3714069

  2. [2]

    Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao. 2022. The values encoded in machine learning research. In ACM Conference on Fairness, Accountability, and Transparency (FAccT). ACM

  3. [3]

    Bowker and Susan Leigh Star

    Geoffery C. Bowker and Susan Leigh Star. 2000. Sorting things out: classification and its consequences. MIT Press, Cambridge, MA, USA

  4. [4]

    Ángel Alexander Cabrera, Marco Tulio Ribeiro, Bongshin Lee, Robert Deline, Adam Perer, and Steven M. Drucker. 2023. What Did My AI Learn? How Data Scientists Make Sense of Model Behavior. ACM Trans. Comput.-Hum. Interact. 30, 1 (March 2023), 1:1–1:27. doi:10.1145/3542921

  5. [5]

    Inha Cha, Juhyun Oh, Cheul Young Park, Jiyoon Han, and Hwalsuk Lee. 2023. Unlocking the Tacit Knowledge of Data Work in Machine Learning. In Extended Abstracts of the 2023 CHI Conference on Human Factors in Computing Systems (CHI EA ’23). Association for Computing Machinery, New York, NY , USA, 1–7. doi:10.1145/3544549.3585616

  6. [6]

    Levy Chaves, Alceu Bissoto, Eduardo Valle, and Sandra Avila. 2023. The performance of transferability metrics does not translate to medical tasks. In MICCAI Workshop on Domain Adaptation and Representation Transfer. Springer, 105–114

  7. [7]

    Sihong Chen, Kai Ma, and Yefeng Zheng. 2019. Med3d: Transfer learning for 3d medical image analysis. arXiv preprint arXiv:1904.00625 (2019)

  8. [8]

    Veronika Cheplygina. 2019. Cats or CAT scans: Transfer learning from natural or medical image source data sets? Current Opinion in Biomedical Engineering 9 (2019), 21–27

  9. [9]

    Mehdi Cherti and Jenia Jitsev. 2022. Effect of pre-training scale on intra-and inter-domain, full and few-shot transfer learning for natural and X-ray chest images. In 2022 International Joint Conference on Neural Networks (IJCNN). IEEE, 1–9

  10. [10]

    Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. 2014. Describing textures in the wild. In Proceedings of the IEEE conference on computer vision and pattern recognition. 3606–3613

  11. [11]

    Line H Clemmensen and Rune D Kjærsgaard. 2022. Data representativity for machine learning and ai systems. arXiv preprint arXiv:2203.04706 (2022)

  12. [12]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. IEEE, 248–255

  13. [13]

    Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. 2018. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. In International conference on learning representations

  14. [14]

    Seung Seog Han, Gyeong Hun Park, Woohyung Lim, Myoung Shin Kim, Jung Im Na, Ilwoo Park, and Sung Eun Chang. 2018. Deep neural networks show an equivalent and often superior performance to dermatologists in onychomycosis diagnosis: Automatic construction of onychomycosis datasets Manuscript submitted to ACM Intuitions of Machine Learning Researchers about ...

  15. [15]

    Ruben Hemelings, Bart Elen, Joao Barbosa-Breda, Sophie Lemmens, Maarten Meire, Sayeh Pourjavan, Evelien Vandewalle, Sara Van de Veire, Matthew B Blaschko, Patrick De Boever, et al. 2020. Accurate prediction of glaucoma from colour fundus images with a convolutional neural network that relies on active and transfer learning. Acta ophthalmologica 98, 1 (202...

  16. [16]

    Hsiu-Fang Hsieh and Sarah E. Shannon. 2005. Three Approaches to Qualitative Content Analysis. Qualitative Health Research 15, 9 (Nov. 2005), 1277–1288. doi:10.1177/1049732305276687 Publisher: SAGE Publications Inc

  17. [17]

    Andrey Ignatov and Grigory Malivenko. 2024. NCT-CRC-HE: Not All Histopathological Datasets are Equally Useful. In European Conference on Computer Vision. Springer, 300–317

  18. [18]

    Jeremy Irvin, Pranav Rajpurkar, Michael Ko, Yifan Yu, Silviana Ciurea-Ilcus, Chris Chute, Henrik Marklund, Behzad Haghgoo, Robyn Ball, Katie Shpanskaya, et al. 2019. Chexpert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence, V ol. 33. 590–597

  19. [19]

    Saachi Jain, Hadi Salman, Alaa Khaddaj, Eric Wong, Sung Min Park, and Aleksander M˛ adry. 2023. A data-based perspective on transfer learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3613–3622

  20. [20]

    Christian Janiesch, Patrick Zschech, and Kai Heinrich. 2021. Machine learning and deep learning. Electronic Markets 31, 3 (Sept. 2021), 685–695. doi:10.1007/s12525-021-00475-2

  21. [21]

    Amelia Jiménez-Sánchez, Natalia-Rozalia Avlona, Dovile Juodelyte, Théo Sourget, Caroline Vang-Larsen, Anna Rogers, Hubert Dariusz Zaj˛ ac, and Veronika Cheplygina. 2024. Copycats: the many lives of a publicly available medical imaging dataset. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J....

  22. [22]

    Dovile Juodelyte, Enzo Ferrante, Yucheng Lu, Prabhant Singh, Joaquin Vanschoren, and Veronika Cheplygina. 2024. On dataset transferability in medical image classification. arXiv preprint arXiv:2412.20172 (2024)

  23. [23]

    Dovile Juodelyte, Amelia Jiménez Sánchez, and Veronika Cheplygina. 2023. Revisiting Hidden Representations in Transfer Learning for Medical Imaging. Transactions on Machine Learning Research(2023)

  24. [24]

    Alexander Ke, William Ellsworth, Oishi Banerjee, Andrew Y Ng, and Pranav Rajpurkar. 2021. CheXtransfer: performance and parameter efficiency of ImageNet models for chest X-Ray interpretation. In Proceedings of the conference on health, inference, and learning. 116–124

  25. [25]

    Hee E Kim, Alejandro Cosa-Linan, Nandhini Santhanam, Mahboubeh Jannesari, Mate E Maros, and Thomas Ganslandt. 2022. Transfer learning for medical image classification: a literature review. BMC medical imaging 22, 1 (2022), 69

  26. [26]

    Haijun Lei, Tao Han, Feng Zhou, Zhen Yu, Jing Qin, Ahmed Elazab, and Baiying Lei. 2018. A deeply supervised residual network for HEp-2 cell classification via cross-modal transfer learning. Pattern Recognition 79 (2018), 290–302

  27. [27]

    Johann Li, Guangming Zhu, Cong Hua, Mingtao Feng, Basheer Bennamoun, Ping Li, Xiaoyuan Lu, Juan Song, Peiyi Shen, Xu Xu, et al. 2023. A systematic collection of medical image datasets for deep learning. Comput. Surveys 56, 5 (2023), 1–51

  28. [28]

    Wenxuan Li, Alan Yuille, and Zongwei Zhou. 2025. How well do supervised 3d models transfer to medical imaging tasks? arXiv preprint arXiv:2501.11253 (2025)

  29. [29]

    Geert Litjens, Thijs Kooi, Babak Ehteshami Bejnordi, Arnaud Arindra Adiyoso Setio, Francesco Ciompi, Mohsen Ghafoorian, Jeroen AWM van der Laak, Bram Van Ginneken, and Clara I Sánchez. 2017. A survey on deep learning in medical image analysis. Medical image analysis 42 (2017), 60–88

  30. [30]

    Nahiyan Malik and Danilo Bzdok. 2022. From YouTube to the brain: Transfer learning can improve brain-imaging predictions with deep learning. Neural Networks 153 (2022), 325–338

  31. [31]

    Christos Matsoukas, Johan Fredin Haslum, Moein Sorkhei, Magnus Söderberg, and Kevin Smith. 2022. What makes transfer learning work for medical images: Feature reuse & other factors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9225–9234

  32. [32]

    Johannes Mehrer, Courtney J Spoerer, Emer C Jones, Nikolaus Kriegeskorte, and Tim C Kietzmann. 2021. An ecologically motivated image dataset for deep learning yields better models of human vision. Proceedings of the National Academy of Sciences 118, 8 (2021), e2011417118

  33. [33]

    Xueyan Mei, Zelong Liu, Philip M Robson, Brett Marinelli, Mingqian Huang, Amish Doshi, Adam Jacobi, Chendi Cao, Katherine E Link, Thomas Yang, et al. 2022. RadImageNet: an open radiologic deep learning research dataset for effective transfer learning. Radiology: Artificial Intelligence 4, 5 (2022), e210315

  34. [34]

    Afonso Menegola, Michel Fornaciali, Ramon Pires, Flávia Vasques Bittencourt, Sandra Avila, and Eduardo Valle. 2017. Knowledge transfer for melanoma screening with deep learning. In 2017 IEEE 14th international symposium on biomedical imaging (ISBI 2017). IEEE, 297–300

  35. [35]

    Thomas Mensink, Jasper Uijlings, Alina Kuznetsova, Michael Gygli, and Vittorio Ferrari. 2021. Factors of influence for transfer learning across diverse appearance domains and task types. IEEE Transactions on Pattern Analysis and Machine Intelligence44, 12 (2021), 9298–9314

  36. [36]

    Milagros Miceli and Julian Posada. 2022. The Data-Production Dispositif. Proc. ACM Hum.-Comput. Interact. 6, CSCW2 (Nov. 2022), 460:1–460:37. doi:10.1145/3555561

  37. [37]

    Shervin Minaee, Rahele Kafieh, Milan Sonka, Shakib Yazdani, and Ghazaleh Jamalipour Soufi. 2020. Deep-COVID: Predicting COVID-19 from chest X-ray images using deep transfer learning. Medical image analysis 65 (2020), 101794

  38. [38]

    Swati Mishra and Jeffrey M Rzeszotarski. 2021. Designing Interactive Transfer Learning Tools for ML Non-Experts. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21) . Association for Computing Machinery, New York, NY , USA, 1–15. doi:10.1145/3411764.3445096 Manuscript submitted to ACM 20 Yucheng Lu, Hubert Dariusz Zaj...

  39. [39]

    Sedir Mohammed, Lisa Ehrlinger, Hazar Harmouch, Felix Naumann, and Divesh Srivastava. 2024. Data Quality Assessment: Challenges and Opportunities. doi:10.48550/arXiv.2403.00526 arXiv:2403.00526 [cs]

  40. [40]

    Inês C Moreira, Igor Amaral, Inês Domingues, António Cardoso, Maria João Cardoso, and Jaime S Cardoso. 2012. Inbreast: toward a full-field digital mammographic database. Academic radiology 19, 2 (2012), 236–248

  41. [41]

    Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How data science workers work with data: Discovery, capture, curation, design, creation. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–15

  42. [42]

    Wolf, Josh Andres, Michael Desmond, Narendra Nath Joshi, Zahra Ashktorab, Aabhas Sharma, Kristina Brimijoin, Qian Pan, Evelyn Duesterwald, and Casey Dugan

    Michael Muller, Christine T. Wolf, Josh Andres, Michael Desmond, Narendra Nath Joshi, Zahra Ashktorab, Aabhas Sharma, Kristina Brimijoin, Qian Pan, Evelyn Duesterwald, and Casey Dugan. 2021. Designing Ground Truth and the Social Life of Labels. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). Association for Compu...

  43. [43]

    Lauren Oakden-Rayner. 2019. Exploring large scale public medical image datasets. arXiv preprint arXiv:1907.12720 (2019)

  44. [44]

    Sinno Jialin Pan and Qiang Yang. 2009. A survey on transfer learning. IEEE Transactions on knowledge and data engineering 22, 10 (2009), 1345–1359

  45. [45]

    Samira Pouyanfar, Saad Sadiq, Yilin Yan, Haiman Tian, Yudong Tao, Maria Presa Reyes, Mei-Ling Shyu, Shu-Ching Chen, and S. S. Iyengar. 2018. A Survey on Deep Learning: Algorithms, Techniques, and Applications. ACM Comput. Surv. 51, 5 (Sept. 2018), 92:1–92:36. doi:10.1145/3234150

  46. [46]

    Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. 2019. Transfusion: Understanding transfer learning with applications to medical imaging. arXiv preprint arXiv:1902.07208 (2019)

  47. [47]

    Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021. AI and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366 (2021)

  48. [48]

    Eduardo Ribeiro, Michael Häfner, Georg Wimmer, Toru Tamaki, JJW Tischendorf, Shigeto Yoshida, Shinji Tanaka, and Andreas Uhl. 2017. Exploring texture transfer learning for colonic polyp classification via convolutional neural networks. In International Symposium on Biomedical Imaging (ISBI). IEEE, 1044–1048

  49. [49]

    Everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–15

  50. [50]

    Kjeld Schmidt. 2012. The Trouble with ‘Tacit Knowledge’. Computer Supported Cooperative Work (CSCW) 21, 2 (June 2012), 163–225. doi:10.1007/s10606-012-9160-8

  51. [51]

    Bibo Shi, Rui Hou, Maciej A Mazurowski, Lars J Grimm, Yinhao Ren, Jeffrey R Marks, Lorraine M King, Carlo C Maley, E Shelley Hwang, and Joseph Y Lo. 2018. Learning better deep features for the prediction of occult invasive disease in ductal carcinoma in situ through transfer learning. In Medical Imaging 2018: Computer-Aided Diagnosis, V ol. 10575. Interna...

  52. [52]

    Hoo-Chang Shin, Holger R Roth, Mingchen Gao, Le Lu, Ziyue Xu, Isabella Nogues, Jianhua Yao, Daniel Mollura, and Ronald M Summers. 2016. Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer Learning. IEEE Transactions on Medical Imaging 35, 5 (2016), 1285–1298

  53. [53]

    Shinde and Seema Shah

    Pramila P. Shinde and Seema Shah. 2018. A Review of Machine Learning and Deep Learning Applications. In 2018 Fourth International Conference on Computing Communication Control and Automation (ICCUBEA). 1–6. doi:10.1109/ICCUBEA.2018.8697857

  54. [54]

    Nima Tajbakhsh, Jae Y Shin, Suryakanth R Gurudu, R Todd Hurst, Christopher B Kendall, Michael B Gotway, and Jianming Liang. 2016. Convolutional neural networks for medical image analysis: full training or fine tuning? IEEE Transactions on Medical Imaging 35, 5 (2016), 1299–1312

  55. [55]

    Jamie Thompson. 2022. A Guide to Abductive Thematic Analysis. The Qualitative Report (May 2022). doi:10.46743/2160-3715/2022.5340

  56. [56]

    Mira Valkonen, Jorma Isola, Onni Ylinen, Ville Muhonen, Anna Saxlin, Teemu Tolonen, Matti Nykter, and Pekka Ruusuvuori. 2019. Cytokeratin- supervised deep learning for automatic recognition of epithelial cells in breast cancers stained for ER, PR, and Ki-67. IEEE transactions on medical imaging 39, 2 (2019), 534–542

  57. [57]

    Gaël Varoquaux and Veronika Cheplygina. 2022. Machine learning for medical imaging: methodological failures and recommendations for the future. Nature Digital Medicine 5, 1 (2022), 1–8

  58. [58]

    Ding Wang, Shantanu Prabhat, and Nithya Sambasivan. 2022. Whose AI Dream? In search of the aspiration in data annotation.. In CHI Conference on Human Factors in Computing Systems. ACM, New Orleans LA USA, 1–16. doi:10.1145/3491102.3502121

  59. [59]

    Ken CL Wong, Tanveer Syeda-Mahmood, and Mehdi Moradi. 2018. Building medical image classifiers with very limited data using segmentation networks. Medical image analysis 49 (2018), 105–116

  60. [60]

    Linshan Wu, Jiaxin Zhuang, and Hao Chen. 2024. Large-scale 3d medical image pre-training with geometric context priors. arXiv preprint arXiv:2410.09890 (2024)

  61. [61]

    Yiting Xie and David Richmond. 2018. Pre-training on grayscale imagenet improves medical image classification. In Proceedings of the European conference on computer vision (ECCV) workshops. 0–0

  62. [62]

    Yuncheng Yang, Meng Wei, Junjun He, Jie Yang, Jin Ye, and Yun Gu. 2023. Pick the best pre-trained model: Towards transferability estimation for medical image segmentation. In International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 674–683

  63. [63]

    Hubert Dariusz Zaj˛ ac, Natalia Rozalia Avlona, Finn Kensing, Tariq Osman Andersen, and Irina Shklovski. 2023. Ground Truth Or Dare: Factors Affecting The Creation Of Medical Datasets For Training AI. In Conference on AI, Ethics, and Society (AIES). 351–362. Manuscript submitted to ACM Intuitions of Machine Learning Researchers about Transfer Learning for...

  64. [64]

    Xingchen Zeng, Ziyao Gao, Yilin Ye, and Wei Zeng. 2024. IntentTuner: An Interactive Framework for Integrating Human Intentions in Fine-tuning Text-to-Image Generative Models. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM, Honolulu HI USA, 1–18. doi:10.1145/3613904.3642165

  65. [65]

    present” (indicated by a check mark) or “absent

    Dora Zhao, Jerone Andrews, Orestis Papakyriakopoulos, and Alice Xiang. 2024. Position: Measure Dataset Diversity, Don’t Just Claim It. InForty-first International Conference on Machine Learning. A TRANSFER LEARNING NOTIONS: PAPER ANNOTATIONS This appendix shows examples of prior literature in machine learning in medical imaging, that discusses different c...