{"id":"cea13a4d-a9d9-4be1-85a0-1f0bf8c557f7","arxiv_id":"2510.00902","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Source-dataset selection for medical transfer learning is driven by community practice and perceived similarity, and 'more similar is better' does not consistently hold.","lead":"Researchers surveyed 15 ML practitioners to see how they choose pretraining datasets for medical image classification. The answers show choices follow community norms and intuition as much as technical similarity, and that people's similarity judgments don't always match their performance expectations.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'not aligned' challenge to 'more similar is better' rests on Spearman correlations from N=15 that cannot distinguish zero from moderate association; range restriction may explain the null.","rationale":"The reader's weakest assumption was the representativeness of the 15 self-selected participants. That is a legitimate external-validity concern, but it is not the most load-bearing issue for the paper's headline. The paper's distinct contribution is the claim that similarity and expected performance are 'not always aligned', and that claim is internally insecure: it is inferred from correlations that are uninformative at N=15 and possibly attenuated by range restriction. Even if the participants were perfectly representative, the reported statistics would not establish the misalignment claim. My concrete test targets that internal inferential gap. The overall verdict remains CONDITIONAL: the qualitative findings about community influence, task-dependence, and vague terminology are plausible and well-illustrated, but the central quantitative challenge to 'more similar is better' should be conditional on reporting confidence intervals, range restriction diagnostics, or additional data. I disagree with the reader's choice of weakest assumption because it locates the risk in sampling rather than in the statistical interpretation of the within-sample data, which is where the claim is actually fragile.","tokens_in":25257,"tokens_out":4133,"duration_ms":45509,"concrete_test":"Re-analyze the raw paired ratings from Q17–Q18 for each source–task combination: (1) compute Spearman ρ with bootstrap 95% CIs and a permutation-based null for every dimension–performance pair; (2) compute the marginal distribution of each similarity rating. If a similarity dimension has ≥80% of responses in two adjacent Likert categories, explicitly test for range restriction and do not interpret ρ≈0 as evidence of misalignment. The 'not aligned' claim would be supported only if the CI excludes ρ>0.5 for at least one similarity dimension while the corresponding performance ratings show meaningful variance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's most novel claim is that similarity ratings and expected performance are 'not always aligned', challenging 'more similar is better' (Abstract; §5.3). The evidence for this is per-source Spearman correlations between participants' dimension ratings and expected fine-tuning performance, computed with N=15 (§4.4, §5.1, Fig. 5). For RadImageNet in the H&E case, all similarity correlations are |ρ| ≤ 0.3; for RadImageNet in the chest X-ray case, all are |ρ| ≤ 0.2. With N=15, the 95% CI for a sample ρ of 0.3 is roughly (−0.24, 0.70), and for ρ of 0.2 it is roughly (−0.34, 0.65). These data are compatible with no association and with a moderate positive association. Moreover, if participants nearly uniformly rate RadImageNet as both similar and high-performing, Spearman's ρ is attenuated by range restriction; a null correlation would then mean only that the ratings did not vary, not that similarity and expected performance are decoupled. The paper's disclaimer that p-values are reported only for completeness (§4.4) does not solve this: the qualitative interpretation in §5.3 treats near-zero ρ as substantive evidence of misalignment. Thus the central 'more similar is better is challenged' claim is not securely supported by the reported statistics.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods survey of 15 ML practitioners on how they select source datasets for transfer learning in medical image classification. Participants answered questions about a recent transfer-learning project and two controlled case studies (H&E patch classification and chest X-ray classification) with three candidate source datasets (ImageNet-1K, RadImageNet, Ecoset). The authors find that choices are task-dependent and shaped by community practices, dataset attributes, and perceived visual/semantic similarity, and they claim that similarity ratings and expected performance are not always aligned, challenging the 'more similar is better' view. The qualitative analysis identifies community influence, source-dataset attributes, and source-target similarity as central themes, alongside frequent vague use of terms like 'domain gap'. The paper is positioned as an HCI contribution to making tacit knowledge in transfer learning explicit.","tokens_in":25567,"tokens_out":2857,"duration_ms":402118,"significance":"If the findings are credited, the paper makes a useful contribution by shifting attention from purely technical transferability metrics to the social and intuitive processes behind source dataset selection, an underexplored area in HCI and medical imaging. The mixed-methods design, including the interactive dataset browser and the transparent reporting of the codebook, is a strength. The paper also makes a concrete, falsifiable claim—that perceived similarity and expected performance can decouple—which is important for the design of decision-support tools for transfer learning. However, the small, network-recruited sample and the statistical fragility of the core 'not aligned' claim limit the strength of the conclusions as currently stated.","major_comments":[{"comment":"The central claim that similarity ratings and expected performance are 'not always aligned', challenging 'more similar is better', is not securely supported by the reported Spearman correlations. For RadImageNet in the H&E case all similarity correlations are |ρ|≤0.3, and in the chest X-ray case all are |ρ|≤0.2 (§5.1, Fig. 5). With N=15, the 95% CI for ρ=0.3 is approximately (−0.24, 0.70), so the data are compatible with no association and with a moderate positive association. Range restriction—participants nearly uniformly rating RadImageNet as similar and high-performing—can attenuate ρ regardless of the true relationship. The disclaimer in §4.4 that p-values are 'for completeness' does not address the interpretive problem, because §5.3 treats near-zero correlations as substantive evidence. Please report confidence intervals or individual-level trajectories, and either soften the abstr","section":"§5.3 and Abstract"},{"comment":"The paper's abstract and discussion make general claims about 'researchers' ('choices are task-dependent and influenced by community practices...'). The evidence base is 15 self-selected participants recruited through the authors' professional networks, including direct email invitations to researchers who had previously engaged with the authors' work. This creates a substantial risk that the sample is biased toward people familiar with the authors' perspective or with above-average interest in medical imaging transfer learning. A sample of this size and recruitment strategy is appropriate for an exploratory qualitative study, but the generalizing language in the abstract and conclusions overreaches. Please explicitly frame the study as exploratory and revise the abstract and §6 to avoid implying population-level conclusions.","section":"§4.3, §6.1"},{"comment":"The qualitative coding follows directed content analysis, but no inter-coder reliability statistic (e.g., Cohen's κ or percentage agreement) is reported. Because the codebook was built from the same literature that informed the questionnaire, the coding process risks confirming the authors' prior taxonomy rather than independently discovering participants' categories. The paper would be strengthened by reporting agreement scores and by describing how many codes emerged inductively versus were imposed a priori. This matters for the qualitative claims about community influence and vagueness, which are otherwise plausible but not quantifiably reliable.","section":"§4.5"},{"comment":"The interpretation of the Friedman and Wilcoxon results is inconsistent with the reported significance levels. In case study 2, the Friedman test is significant, but for the key RadImageNet vs. ImageNet-1K contrast the paper reports only 'paired difference is positive (W=3.0, r=0.8)' without a p-value, while earlier it notes the across-case RadImageNet shift has 'adjusted p=0.2'. The text then states 'the ordering is RadImageNet > ImageNet-1K ≈ Ecoset' as if this were established. Please report all p-values and effect sizes with confidence intervals, and use language proportional to the evidence, e.g., 'tendency' rather than 'ordering' where results are not significant.","section":"§5.1"}],"minor_comments":[{"comment":"'Table??' is a broken cross-reference; the demographics table is not properly cited. Please fix.","section":"§4.3"},{"comment":"Figures 3–5 are hard to read in the current rendering and the caption for Fig. 5 does not explain the radar chart axes clearly. Please ensure high-resolution figures and clarify the meaning of the scale in the caption.","section":"§5.1"},{"comment":"The text says 'we report the types of statistical significance tests for completeness, rather than basing our conclusions on the (here not reported) p-values' but then uses language like 'the only signal is a hint for robustness' (§5.1). Please either report the p-values in the text or consistently avoid interpreting non-significant correlations as signals.","section":"§4.4"},{"comment":"The target dataset is referred to as 'CRC-VAL-HE-7K' but reference [17] is titled 'NCT-CRC-HE'. Please reconcile the naming and ensure the dataset version/identifier is accurate.","section":"§4.2"},{"comment":"The appendix tables are dense and the checkmark categories are not aligned precisely with the quoted text in some rows (e.g., Table A6, Raghu et al.). Consider using a column for the quote's relevant factor rather than global checkmarks.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable exploratory HCI/qualitative study, but the abstract and conclusions overstate the quantitative support for the 'not aligned' claim. A revision that reframes the central contribution as qualitative and exploratory, adds appropriate uncertainty, and tones down the generalization would be within scope. The lack of inter-coder reliability and the network recruitment are additional but addressable concerns."},"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Source-dataset selection in transfer learning is a social and intuitive process, not a systematic one.","keywords":["transfer learning","source dataset selection","medical image classification","practitioner intuition","human-computer interaction","mixed-methods survey","similarity heuristics","community practices"],"falsifier":"A preregistered replication with a much larger, more diverse sample that tracks real source-dataset choices in actual medical-imaging transfer-learning projects, and measures whether practitioners' similarity ratings actually predict fine-tuning performance, would settle whether these patterns are robust or artifacts of the small, networked sample.","tokens_in":25152,"feed_emoji":"🔬","tokens_out":6638,"duration_ms":56679,"temperature":0.7,"pith_summary":"The paper sets out to show that when machine-learning researchers choose which dataset to pretrain on for medical-image classification, they are guided more by intuition, community habits, and vague notions of similarity than by systematic, measurable criteria. The authors support this with a task-based survey of 15 practitioners who, for two very different target tasks (colorectal tissue patches and chest X-rays) and the same three candidate source datasets, mostly justified choices through established baselines, the availability of pretrained models, personal experience, and perceived 'domain similarity' or 'image quality'—terms they invoked without precise definitions. Their quantitative and qualitative results challenge the conventional rule that 'more similar is better': similarity ratings and expected fine-tuning performance were not consistently aligned, and fairness considerations were almost absent from the reasoning. A sympathetic reader should care because source selection is a high-stakes, under-documented decision that affects the generalizability and equity of medical AI; if this description of practice is right, better tooling and clearer conceptual definitions could shift the field from intuition toward more deliberate, transparent choices.","feed_headline":"Community habits, not transfer metrics, steer pretraining picks","feed_subtitle":"A 15-practitioner survey finds source-dataset choices rely on intuition, vague similarity, and reviewer expectations.","key_machinery":"The argument is carried by a task-based survey combined with an interactive dataset browser. Each participant was presented with two deliberately different target tasks (a nine-class colorectal H&E patch classification and a multi-label chest X-ray classification) against the same three candidate pretraining sources (ImageNet-1K, RadImageNet, and Ecoset), and had to rate willingness, expected fine-tuning performance, and the expected effects of pretraining across six dimensions (domain, visual, and embedding similarity; dataset scale; fairness; robustness). This repeated-measures design isolates how the change in target task alters choices, while a qualitative content analysis guided by a th","core_discovery":"The central claim is that source-dataset selection for transfer learning in medical imaging is not a rational, evidence-driven optimization but a situated practice in which community influence (what colleagues use, what reviewers expect, what baselines are standard), practical dataset attributes (size, ease of use, pretrained-model availability), and perceived similarity—semantic, visual, or in the learned embedding space—jointly determine choices. The study's paired case-study design shows this concretely: switching the target from H&E tissue patches to chest X-rays moved practitioners' preference toward the radiological source, yet their similarity ratings did not consistently track their","pith_inferences":["A direct behavioral test—logging real researchers' source-dataset choices in actual projects and comparing them with their stated heuristics—would separate what practitioners say they do from what they do.","The similarity-performance misalignment might be explained by a distinction between visual surface similarity (color, texture, shape) and task-relevant feature reuse; a controlled manipulation of these cues could identify which notion is actually driving expectations.","If community influence is a genuine driver, then publication and review norms are direct levers on transfer-learning practice: requiring authors to justify source choice with defined criteria could reshape behavior more effectively than new algorithms.","The near-total absence of fairness reasoning in source selection suggests that pretraining biases can propagate silently, since practitioners do not frame source choice as an equity-relevant decision."],"forward_implications":["If source choice is task-dependent and community-influenced, then guidance for transfer learning must be contextual rather than a one-size-fits-all recommendation.","Because availability of pretrained models and established baselines strongly shape choice, new domain-specific pretraining datasets need to become community standards before they displace generic ones like ImageNet.","Since perceived similarity and expected performance are not always aligned, researchers should empirically validate transferability rather than trusting intuition that 'more similar is better.'","Because fairness and robustness were rarely linked to source choice in the responses, dataset documentation and reporting norms should make these dimensions explicit when justifying source selection.","The pervasive vague use of terms like 'domain gap' and 'good image quality' indicates that operational definitions and interactive tools are needed to make these concepts usable in practice."],"fun_headline_variants":["Why ML experts pick pretraining data: gut feel, not metrics","Transfer learning choices driven by habit, not evidence","Survey: researchers rely on intuition for transfer learning","Pretraining picks: community bias beats systematic metrics","Medical transfer learning: intuition wins over science"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The findings rest on the answers of 15 self-selected practitioners recruited through the authors' own networks; if these volunteers' hypothetical judgments are not representative of how machine-learning researchers at large choose source datasets, the paper's general claims about researcher intuition do not follow.","fun_headline_variants_meta":{"raw":{"variants":["Why ML experts pick pretraining data: gut feel, not metrics","Transfer learning choices driven by habit, not evidence","Survey: researchers rely on intuition for transfer learning","Pretraining picks: community bias beats systematic metrics","Medical transfer learning: intuition wins over science"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":968,"prompt_tokens":686,"completion_tokens":282,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":208}},"tokens_in":430,"tokens_out":282,"duration_ms":3277,"temperature":1.0,"reasoning_tokens":208,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T12:59:31.349486+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A preregistered replication with a much larger, more diverse sample that tracks real source-dataset choices in actual medical-imaging transfer-learning projects, and measures whether practitioners' similarity ratings actually predict fine-tuning performance, would settle whether these patterns are robust or artifacts of the small, networked sample.","supporting_citations":[],"review_version":1}