REVIEW 2 major objections 5 minor 18 references
The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read This paper claims that using public datasets is associated with 60.8% more citations per year for medical image computing papers, after stratifying by code release and open access.
desk verdict A transparent, well-scoped measurement of public-data use and citation outcomes at MICCAI; the 60.8% association is credible as a descriptive estimate but should not be read as causal. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a stratified ratio-of-means analysis. Each of the 500 papers is manually coded for three binary attributes—data type (public versus private-only), code release, and open-access status—and papers are sorted into the four groups formed by the latter two attributes. Within each group the mean citations per year of public-data papers is divided by that of private-data papers, and the four ratios are combined into a weighted average by group prevalence. Winsorizing citation rates at 50 per year protects against a handful of extremely cited papers, and the bootstrap supplies the confidence interval around the final 60.8% estimate.
What would settle it
Have two independent, blinded coders re-classify the same 500 papers using the paper's own definitions, and retrieve citation counts at a later date; if inter-coder agreement on data status is low, or if re-measured citation counts shrink the public-data advantage toward zero, the central association fails.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a quantitative association: among accepted computer-vision and machine-learning papers at the annual medical image computing conference from 2014 to 2018, using a publicly available dataset (whether an existing benchmark or data released with the paper) is associated with 60.8% more citations per year than using only private data, with a 95% bootstrap confidence interval of 28.1% to 110.2%, after controlling for two known confounders: code release and open-access publication. The paper also establishes that 54.2% of these papers used only private data and that this proportion declined from 64.0% in 2014 to 44% in 2018. It further reports that 21.6% of papers using public datasets did not cite the dataset in a formal way, and that in 5.0% of such cases no citable entity existed at all.
Load-bearing premise
The whole result rests on the manual labels that put each paper into 'public data' or 'private data' and on one-time citation counts; if those labels or counts are systematically wrong in a way that tracks citation success, the 60.8% gap and the 54.2% private-data share could be artifacts.
Editorial extensions
If this is right
- If the association holds, papers that reuse public datasets receive materially more attention than methodologically similar work built on private data, making dataset choice a de facto impact lever.
- A required data-availability statement, along the lines recommended by the authors, could shift the field's private-data share quickly without requiring new infrastructure.
- Dataset creators who provide a citable, indexed entity remove the main structural excuse for informal data references.
- The decline in private-data-only papers over the five years can be read as early evidence that the norm is already moving toward openness, so policy changes may be reinforcing rather than reversing a trend.
Reading between the lines
- By the authors' own reasoning, a similar citation advantage likely holds in other fields where data collection dominates method development, such as clinical natural-language processing; this is an extrapolation, not something the paper measures.
- If the association reflects causality, the field's citation economy systematically undervalues private-data work regardless of quality, so evaluation and funding criteria that emphasize citations would add another incentive toward data sharing.
- A direct test of the paper's mechanism would be a before/after study of a conference that introduces a required data-availability statement, tracking citation rates while holding methods otherwise constant.
- The coding protocol could be applied to proceedings after 2018 to test whether the private-data share and the citation gap converge, a prediction that follows from the paper's narrative but is not tested here.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript analyzes 500 MICCAI papers from 2014 to 2018 that use machine learning for computer vision tasks, manually labelling each paper for data usage (public, released, or private-only), code release, open access status, and citation counts. The authors report three main findings: (1) 54.2% of the sampled papers used only private data, with a decline from 64% to 44% across the five years; (2) after stratifying by open access and code release, papers using public data received 60.8% more citations per year than private-data-only papers (95% CI 28.1%–110.2%); and (3) 21.6% of papers using public data did not cite the dataset itself. The paper concludes with policy recommendations for MICCAI, including a data availability statement requirement and stronger reviewer scrutiny of data references.
Significance. If the headline association is taken at face value, it is a novel and useful contribution to the bibliometric literature on data sharing in medical image computing, complementing prior work by Piwowar and Vision, Drachen et al., and Colavizza et al. The analysis is transparent and reproducible: the authors make their code and data publicly available, describe their manual protocol, use a stratified ratio estimator with a bootstrap confidence interval, and explicitly report a winsorization robustness check. The paper is also careful to stop short of claiming causality, acknowledging the unmeasured confounder of author reputation. The main value lies in quantifying a large disparity in citation impact between public-data and private-data papers, which is relevant to MICCAI policy discussions and to researchers studying reproducibility incentives.
major comments (2)
- [§4.3] The central estimate of a 60.8% citation advantage for public-data papers is plausibly confounded by author reputation, a concern the authors explicitly acknowledge but do not address quantitatively. Since the paper's title and policy recommendations lean on this estimate, the manuscript should either control for reputation (for example, by stratifying on the senior author's prior citation record or h-index) or perform a sensitivity analysis showing how strongly a reputation effect would need to be to explain the observed ratio. Without such analysis, the interpretation should be substantially softened to a purely descriptive association, and the recommendations should be framed as conditional on this limitation.
- [§3.1] All outcome and exposure variables—data usage category, code release, open access, and citation counts—were assigned through manual review by the authors, but no inter-rater reliability is reported. If coding errors are correlated with citation counts or with public-data status, the estimated ratio and the reported prevalence could be biased. The authors should provide a codebook and have a second coder independently label a random subsample (e.g., 20% of the 500 papers), reporting agreement statistics such as Cohen's kappa, and show that disagreements are not associated with the outcome. This is a load-bearing issue because the entire analysis depends on the accuracy of these manual labels.
minor comments (5)
- [§3.1] The date on which Google Scholar citation counts were collected is not stated. Because citation counts change over time and the study spans five publication years, the authors should report the collection date and, ideally, verify stability by re-collecting a subsample at a later date.
- [§3.2] The winsorization threshold of 50 citations per year is described as a robustness check, but only the direction and significance are commented on. Reporting the bootstrap confidence interval with and without the two affected papers, or with alternative thresholds (e.g., 25 and 100), would make the robustness claim more concrete.
- [§4.4] The section title says "More than quarter of data references were not citations," but the reported proportion is 21.6%, which is less than one quarter. The title should be corrected to "More than one in five" or the reported percentage should be replaced with the actual percentage of 21.6% in a correctly worded heading.
- [§4.3] The sentence "papers based on public data were cited over 60% more per year" appears without the confidence interval in the abstract and in the opening of Section 4.3. Stating the interval (28.1%–110.2%) in both places would help readers assess the precision of the estimate.
- [§4.1] The notable increase in private-data-only papers in 2017 is described as anomalous, but no explanation or analysis is offered. A brief discussion of possible reasons (e.g., conference theme, author pool, or data collection artifact) would strengthen the temporal description.
Circularity Check
No circularity: the paper reports an observational bibliometric measurement, not a derivation or fitted prediction.
full rationale
This paper makes no derived prediction that reduces to its own inputs. Its central quantities are direct measurements: the proportion of sampled MICCAI CV/ML papers using only private data (54.2%, Section 4.1) is a count from manually labeled papers, and the citation advantage (60.8% more citations per year, 95% CI 28.1%–110.2%, Section 4.3) is a stratified ratio estimator computed from external Google Scholar citation counts, with bootstrap confidence intervals. No parameter is fitted to a subset and then renamed as a prediction; no uniqueness theorem or prior claim by the same authors is imported to force a conclusion; and the only data-dependent choice, Winsorization at 50 citations per year, is disclosed and stated not to change the direction or significance of the result (Section 3.2). The acknowledged inability to control for author reputation (Section 4.3) is a confounding/validity limitation of an observational association, not a circular step, because the estimate is explicitly not advanced as a causal first-principles result. The paper is self-contained against external citation data, and therefore warrants a circularity score of 0.
Assumptions & free parameters
free parameters (1)
- Winsorization threshold for citations per year =
50 citations/year
assumptions (3)
- domain assumption Random ordering of accepted papers and manual screening yields a representative sample of MICCAI CV/ML papers.
- domain assumption Google Scholar citation counts are a valid and accurate proxy for scientific impact.
- domain assumption Manual classification of public data use, code release, and reference type is correct.
Cite this review
Pith. "Pith review of The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018." pith.science (2026). https://pith.science/paper/H7UMTS7X
@misc{pith2026190806830,
author = {Pith},
title = {Pith review of: The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018},
year = {2026},
howpublished = {\url{https://pith.science/paper/H7UMTS7X}},
note = {Machine review of arXiv:1908.06830}
}
read the original abstract
Widely-used public benchmarks are of huge importance to computer vision and machine learning research, especially with the computational resources required to reproduce state of the art results quickly becoming untenable. In medical image computing, the wide variety of image modalities and problem formulations yields a huge task-space for benchmarks to cover, and thus the widespread adoption of standard benchmarks has been slow, and barriers to releasing medical data exacerbate this issue. In this paper, we examine the role that publicly available data has played in MICCAI papers from the past five years. We find that more than half of these papers are based on private data alone, although this proportion seems to be decreasing over time. Additionally, we observed that after controlling for open access publication and the release of code, papers based on public data were cited over 60% more per year than their private-data counterparts. Further, we found that more than 20% of papers using public data did not provide a citation to the dataset or associated manuscript, highlighting the "second-rate" status that data contributions often take compared to theoretical ones. We conclude by making recommendations for MICCAI policies which could help to better incentivise data sharing and move the field toward more efficient and reproducible science.
Figures
Reference graph
Works this paper leans on
-
[1]
In: International conference on medical image computing and computer-assisted intervention
C ¸ i¸ cek,¨O., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O.: 3d u-net: learning dense volumetric segmentation from sparse annotation. In: International conference on medical image computing and computer-assisted intervention. pp. 424–432. Springer (2016)
2016
-
[2]
Journal of digital imaging 26(6), 1045–1057 (2013)
Clark, K., Vendt, B., Smith, K., Freymann, J., Kirby, J., Koppel, P., Moore, S., Phillips, S., Maffitt, D., Pringle, M., et al.: The cancer imaging archive (tcia): main- taining and operating a public information repository. Journal of digital imaging 26(6), 1045–1057 (2013)
work page 2013
-
[3]
The citation advantage of linking publications to research data
Colavizza, G., Hrynaszkiewicz, I., Staden, I., Whitaker, K., McGillivray, B.: The citation advantage of linking publications to research data. arXiv preprint arXiv:1907.02565 (2019)
work page Pith review arXiv 2019
-
[4]
In: 2009 IEEE conference on computer vision and pattern recognition
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large- scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. pp. 248–255. Ieee (2009)
2009
-
[5]
Statistische Hefte 15(2-3), 157–170 (1974)
Dixon, W.J., Yuen, K.K.: Trimming and winsorization: A review. Statistische Hefte 15(2-3), 157–170 (1974)
work page 1974
-
[6]
Drachen, T., Ellegaard, O., Larsen, A., Dorch, S.: Sharing data increases citations. Liber Quarterly 26(2) (2016)
work page 2016
-
[7]
Efron, B., Tibshirani, R.J.: An introduction to the bootstrap. CRC press (1994)
work page 1994
-
[8]
Radiographics 37(2), 505–515 (2017)
Erickson, B.J., Korfiatis, P., Akkus, Z., Kline, T.L.: Machine learning for medical imaging. Radiographics 37(2), 505–515 (2017)
work page 2017
Show all 18 references
-
[9]
PLoS biology 4(5), e157 (2006)
Eysenbach, G.: Citation advantage of open access articles. PLoS biology 4(5), e157 (2006)
2006
-
[10]
Circulation 101(23), e215–e220 (2000)
Goldberger, A.L., Amaral, L.A., Glass, L., Hausdorff, J.M., Ivanov, P.C., Mark, R.G., Mietus, J.E., Moody, G.B., Peng, C.K., Stanley, H.E.: Physiobank, phys- iotoolkit, and physionet: components of a new research resource for complex phys- iologic signals. Circulation 101(23), ...
2000
-
[11]
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images. Tech. rep., Citeseer (2009)
2009
-
[12]
In: European conference on computer vision
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ ar, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014)
2014
-
[13]
PeerJ 1, e175 (2013)
Piwowar, H.A., Vision, T.J.: Data reuse and the open data citation advantage. PeerJ 1, e175 (2013)
2013
-
[14]
In: International conference on medical image computing and computer- assisted intervention
Roth, H.R., Lu, L., Farag, A., Shin, H.C., Liu, J., Turkbey, E.B., Summers, R.M.: Deeporgan: Multi-level deep convolutional networks for automated pancreas seg- mentation. In: International conference on medical image computing and computer- assisted intervention. pp. 556–564....
2015
-
[15]
Proceedings of the National Academy of Sciences 115(50), 12603–12607 (2018)
Sekara, V., Deville, P., Ahnert, S.E., Barab´ asi, A.L., Sinatra, R., Lehmann, S.: The chaperone effect in scientific publishing. Proceedings of the National Academy of Sciences 115(50), 12603–12607 (2018)
2018
-
[16]
Journal of Informetrics 8(4), 963–971 (2014)
Thelwall, M., Wilson, P.: Regression for citation data: An evaluation of different methods. Journal of Informetrics 8(4), 963–971 (2014)
2014
-
[17]
Computing in Science & Engineering 14(4), 42–47 (2012)
Vandewalle, P.: Code sharing is associated with research impact in image process- ing. Computing in Science & Engineering 14(4), 42–47 (2012)
2012
-
[18]
Scien- tific data 3 (2016)
Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.W., da Silva Santos, L.B., Bourne, P.E., et al.: The fair guiding principles for scientific data management and stewardship. Scien- tific data 3 (2016)
2016
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.