Pith. sign in

REVIEW 2 major objections 5 minor 18 references

The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018

T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read This paper claims that using public datasets is associated with 60.8% more citations per year for medical image computing papers, after stratifying by code release and open access.

desk verdict A transparent, well-scoped measurement of public-data use and citation outcomes at MICCAI; the 60.8% association is credible as a descriptive estimate but should not be read as causal. read the letter →

arxiv 1908.06830 v1 pith:H7UMTS7X submitted 2019-08-12 cs.LG eess.IVstat.ML

classification cs.LGeess.IVstat.ML
keywords publicdatacitationadvantageprivatemedicalimagecomputingsharingreproducibilitybibliometricscoderelease
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to measure how often medical image computing research is built on datasets that anyone can download, and whether that choice is associated with how often the work is cited. Reviewing a random sample of 500 accepted papers at the annual medical image computing conference between 2014 and 2018, it finds that more than half relied on private data alone, though that share fell from 64% to 44% over the period. After stratifying by code release and open-access status, papers using public data were cited about 60.8% more per year than those using only private data. The authors read this as evidence that reusable data is a major catalyst in this field and recommend policy changes, such as required data-availability statements, to make sharing the norm.

What carries the argument

The central mechanism is a stratified ratio-of-means analysis. Each of the 500 papers is manually coded for three binary attributes—data type (public versus private-only), code release, and open-access status—and papers are sorted into the four groups formed by the latter two attributes. Within each group the mean citations per year of public-data papers is divided by that of private-data papers, and the four ratios are combined into a weighted average by group prevalence. Winsorizing citation rates at 50 per year protects against a handful of extremely cited papers, and the bootstrap supplies the confidence interval around the final 60.8% estimate.

What would settle it

Have two independent, blinded coders re-classify the same 500 papers using the paper's own definitions, and retrieve citation counts at a later date; if inter-coder agreement on data status is low, or if re-measured citation counts shrink the public-data advantage toward zero, the central association fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a quantitative association: among accepted computer-vision and machine-learning papers at the annual medical image computing conference from 2014 to 2018, using a publicly available dataset (whether an existing benchmark or data released with the paper) is associated with 60.8% more citations per year than using only private data, with a 95% bootstrap confidence interval of 28.1% to 110.2%, after controlling for two known confounders: code release and open-access publication. The paper also establishes that 54.2% of these papers used only private data and that this proportion declined from 64.0% in 2014 to 44% in 2018. It further reports that 21.6% of papers using public datasets did not cite the dataset in a formal way, and that in 5.0% of such cases no citable entity existed at all.

Load-bearing premise

The whole result rests on the manual labels that put each paper into 'public data' or 'private data' and on one-time citation counts; if those labels or counts are systematically wrong in a way that tracks citation success, the 60.8% gap and the 54.2% private-data share could be artifacts.

Editorial extensions

If this is right

  • If the association holds, papers that reuse public datasets receive materially more attention than methodologically similar work built on private data, making dataset choice a de facto impact lever.
  • A required data-availability statement, along the lines recommended by the authors, could shift the field's private-data share quickly without requiring new infrastructure.
  • Dataset creators who provide a citable, indexed entity remove the main structural excuse for informal data references.
  • The decline in private-data-only papers over the five years can be read as early evidence that the norm is already moving toward openness, so policy changes may be reinforcing rather than reversing a trend.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • By the authors' own reasoning, a similar citation advantage likely holds in other fields where data collection dominates method development, such as clinical natural-language processing; this is an extrapolation, not something the paper measures.
  • If the association reflects causality, the field's citation economy systematically undervalues private-data work regardless of quality, so evaluation and funding criteria that emphasize citations would add another incentive toward data sharing.
  • A direct test of the paper's mechanism would be a before/after study of a conference that introduces a required data-availability statement, tracking citation rates while holding methods otherwise constant.
  • The coding protocol could be applied to proceedings after 2018 to test whether the private-data share and the citation gap converge, a prediction that follows from the paper's narrative but is not tested here.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This manuscript analyzes 500 MICCAI papers from 2014 to 2018 that use machine learning for computer vision tasks, manually labelling each paper for data usage (public, released, or private-only), code release, open access status, and citation counts. The authors report three main findings: (1) 54.2% of the sampled papers used only private data, with a decline from 64% to 44% across the five years; (2) after stratifying by open access and code release, papers using public data received 60.8% more citations per year than private-data-only papers (95% CI 28.1%–110.2%); and (3) 21.6% of papers using public data did not cite the dataset itself. The paper concludes with policy recommendations for MICCAI, including a data availability statement requirement and stronger reviewer scrutiny of data references.

Significance. If the headline association is taken at face value, it is a novel and useful contribution to the bibliometric literature on data sharing in medical image computing, complementing prior work by Piwowar and Vision, Drachen et al., and Colavizza et al. The analysis is transparent and reproducible: the authors make their code and data publicly available, describe their manual protocol, use a stratified ratio estimator with a bootstrap confidence interval, and explicitly report a winsorization robustness check. The paper is also careful to stop short of claiming causality, acknowledging the unmeasured confounder of author reputation. The main value lies in quantifying a large disparity in citation impact between public-data and private-data papers, which is relevant to MICCAI policy discussions and to researchers studying reproducibility incentives.

major comments (2)
  1. [§4.3] The central estimate of a 60.8% citation advantage for public-data papers is plausibly confounded by author reputation, a concern the authors explicitly acknowledge but do not address quantitatively. Since the paper's title and policy recommendations lean on this estimate, the manuscript should either control for reputation (for example, by stratifying on the senior author's prior citation record or h-index) or perform a sensitivity analysis showing how strongly a reputation effect would need to be to explain the observed ratio. Without such analysis, the interpretation should be substantially softened to a purely descriptive association, and the recommendations should be framed as conditional on this limitation.
  2. [§3.1] All outcome and exposure variables—data usage category, code release, open access, and citation counts—were assigned through manual review by the authors, but no inter-rater reliability is reported. If coding errors are correlated with citation counts or with public-data status, the estimated ratio and the reported prevalence could be biased. The authors should provide a codebook and have a second coder independently label a random subsample (e.g., 20% of the 500 papers), reporting agreement statistics such as Cohen's kappa, and show that disagreements are not associated with the outcome. This is a load-bearing issue because the entire analysis depends on the accuracy of these manual labels.
minor comments (5)
  1. [§3.1] The date on which Google Scholar citation counts were collected is not stated. Because citation counts change over time and the study spans five publication years, the authors should report the collection date and, ideally, verify stability by re-collecting a subsample at a later date.
  2. [§3.2] The winsorization threshold of 50 citations per year is described as a robustness check, but only the direction and significance are commented on. Reporting the bootstrap confidence interval with and without the two affected papers, or with alternative thresholds (e.g., 25 and 100), would make the robustness claim more concrete.
  3. [§4.4] The section title says "More than quarter of data references were not citations," but the reported proportion is 21.6%, which is less than one quarter. The title should be corrected to "More than one in five" or the reported percentage should be replaced with the actual percentage of 21.6% in a correctly worded heading.
  4. [§4.3] The sentence "papers based on public data were cited over 60% more per year" appears without the confidence interval in the abstract and in the opening of Section 4.3. Stating the interval (28.1%–110.2%) in both places would help readers assess the precision of the estimate.
  5. [§4.1] The notable increase in private-data-only papers in 2017 is described as anomalous, but no explanation or analysis is offered. A brief discussion of possible reasons (e.g., conference theme, author pool, or data collection artifact) would strengthen the temporal description.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the paper reports an observational bibliometric measurement, not a derivation or fitted prediction.

full rationale

This paper makes no derived prediction that reduces to its own inputs. Its central quantities are direct measurements: the proportion of sampled MICCAI CV/ML papers using only private data (54.2%, Section 4.1) is a count from manually labeled papers, and the citation advantage (60.8% more citations per year, 95% CI 28.1%–110.2%, Section 4.3) is a stratified ratio estimator computed from external Google Scholar citation counts, with bootstrap confidence intervals. No parameter is fitted to a subset and then renamed as a prediction; no uniqueness theorem or prior claim by the same authors is imported to force a conclusion; and the only data-dependent choice, Winsorization at 50 citations per year, is disclosed and stated not to change the direction or significance of the result (Section 3.2). The acknowledged inability to control for author reputation (Section 4.3) is a confounding/validity limitation of an observational association, not a circular step, because the estimate is explicitly not advanced as a causal first-principles result. The paper is self-contained against external citation data, and therefore warrants a circularity score of 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper's claims rest on the accuracy of manual annotation and on the validity of citations-per-year as an impact proxy. There are no invented entities. The only hand-set numeric choice is the Winsorization cap, which the authors report does not change the direction or significance of the result. No circular derivation is present.

free parameters (1)
  • Winsorization threshold for citations per year = 50 citations/year
    Chosen by hand to reduce the influence of two high-citation papers (Sections 3.2 and 4.3). The authors report that trimming at this threshold did not change the direction or significance of the results, so it is a robustness parameter rather than a fitted constant.
assumptions (3)
  • domain assumption Random ordering of accepted papers and manual screening yields a representative sample of MICCAI CV/ML papers.
    Section 3.1: the authors randomly ordered all accepted papers and manually selected the first 100 per year meeting inclusion criteria. This assumes the manual screening and inclusion criteria do not introduce bias.
  • domain assumption Google Scholar citation counts are a valid and accurate proxy for scientific impact.
    Section 3.1 records citation count from Google Scholar as the key outcome variable. The analysis treats citations per year as a meaningful measure of impact without reporting retrieval dates or validation against other databases.
  • domain assumption Manual classification of public data use, code release, and reference type is correct.
    Section 3.1 describes one-time manual annotation by the authors without inter-rater reliability or independent validation. The central statistics depend on these classifications being accurate and unbiased.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018." pith.science (2026). https://pith.science/paper/H7UMTS7X

@misc{pith2026190806830,
  author       = {Pith},
  title        = {Pith review of: The Role of Publicly Available Data in MICCAI Papers from 2014 to 2018},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H7UMTS7X}},
  note         = {Machine review of arXiv:1908.06830}
}
read the original abstract

Widely-used public benchmarks are of huge importance to computer vision and machine learning research, especially with the computational resources required to reproduce state of the art results quickly becoming untenable. In medical image computing, the wide variety of image modalities and problem formulations yields a huge task-space for benchmarks to cover, and thus the widespread adoption of standard benchmarks has been slow, and barriers to releasing medical data exacerbate this issue. In this paper, we examine the role that publicly available data has played in MICCAI papers from the past five years. We find that more than half of these papers are based on private data alone, although this proportion seems to be decreasing over time. Additionally, we observed that after controlling for open access publication and the release of code, papers based on public data were cited over 60% more per year than their private-data counterparts. Further, we found that more than 20% of papers using public data did not provide a citation to the dataset or associated manuscript, highlighting the "second-rate" status that data contributions often take compared to theoretical ones. We conclude by making recommendations for MICCAI policies which could help to better incentivise data sharing and move the field toward more efficient and reproducible science.

Figures

Figures reproduced from arXiv: 1908.06830 by the authors.

Figure 1
Figure 1. The various modes of data used by MICCAI CV/ML papers from 2014 through 2018. The lightest region (top) represents papers that used at least one existing public dataset, the middle region represents papers that used their own data but publicly released it with their paper, and the darkest region (bottom) represents papers using only private data that was not released with publication [PITH_FULL_IMAGE:figures/full_f… view at source ↗
Figure 2
Figure 2. The prevalence of MICCAI CV/ML papers releasing code over time. The lighter region (top) represents papers that did release their code with publications, and the darker region represents papers than did not (bottom). Even rarer was the practice of releasing one’s data. Of the 309 papers that used their own data, only 15 (4.9%) released this data by the time of publication. We believe that this illustrates the high b… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

18 extracted references · 15 canonical work pages

  1. [1]

    In: International conference on medical image computing and computer-assisted intervention

    C ¸ i¸ cek,¨O., Abdulkadir, A., Lienkamp, S.S., Brox, T., Ronneberger, O.: 3d u-net: learning dense volumetric segmentation from sparse annotation. In: International conference on medical image computing and computer-assisted intervention. pp. 424–432. Springer (2016)

  2. [2]

    Journal of digital imaging 26(6), 1045–1057 (2013)

    Clark, K., Vendt, B., Smith, K., Freymann, J., Kirby, J., Koppel, P., Moore, S., Phillips, S., Maffitt, D., Pringle, M., et al.: The cancer imaging archive (tcia): main- taining and operating a public information repository. Journal of digital imaging 26(6), 1045–1057 (2013)

  3. [3]

    The citation advantage of linking publications to research data

    Colavizza, G., Hrynaszkiewicz, I., Staden, I., Whitaker, K., McGillivray, B.: The citation advantage of linking publications to research data. arXiv preprint arXiv:1907.02565 (2019)

  4. [4]

    In: 2009 IEEE conference on computer vision and pattern recognition

    Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large- scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. pp. 248–255. Ieee (2009)

  5. [5]

    Statistische Hefte 15(2-3), 157–170 (1974)

    Dixon, W.J., Yuen, K.K.: Trimming and winsorization: A review. Statistische Hefte 15(2-3), 157–170 (1974)

  6. [6]

    Liber Quarterly 26(2) (2016)

    Drachen, T., Ellegaard, O., Larsen, A., Dorch, S.: Sharing data increases citations. Liber Quarterly 26(2) (2016)

  7. [7]

    CRC press (1994)

    Efron, B., Tibshirani, R.J.: An introduction to the bootstrap. CRC press (1994)

  8. [8]

    Radiographics 37(2), 505–515 (2017)

    Erickson, B.J., Korfiatis, P., Akkus, Z., Kline, T.L.: Machine learning for medical imaging. Radiographics 37(2), 505–515 (2017)

Show all 18 references
  1. [9]

    PLoS biology 4(5), e157 (2006)

    Eysenbach, G.: Citation advantage of open access articles. PLoS biology 4(5), e157 (2006)

  2. [10]

    Circulation 101(23), e215–e220 (2000)

    Goldberger, A.L., Amaral, L.A., Glass, L., Hausdorff, J.M., Ivanov, P.C., Mark, R.G., Mietus, J.E., Moody, G.B., Peng, C.K., Stanley, H.E.: Physiobank, phys- iotoolkit, and physionet: components of a new research resource for complex phys- iologic signals. Circulation 101(23), ...

  3. [11]

    Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images. Tech. rep., Citeseer (2009)

  4. [12]

    In: European conference on computer vision

    Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll´ ar, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: European conference on computer vision. pp. 740–755. Springer (2014)

  5. [13]

    PeerJ 1, e175 (2013)

    Piwowar, H.A., Vision, T.J.: Data reuse and the open data citation advantage. PeerJ 1, e175 (2013)

  6. [14]

    In: International conference on medical image computing and computer- assisted intervention

    Roth, H.R., Lu, L., Farag, A., Shin, H.C., Liu, J., Turkbey, E.B., Summers, R.M.: Deeporgan: Multi-level deep convolutional networks for automated pancreas seg- mentation. In: International conference on medical image computing and computer- assisted intervention. pp. 556–564....

  7. [15]

    Proceedings of the National Academy of Sciences 115(50), 12603–12607 (2018)

    Sekara, V., Deville, P., Ahnert, S.E., Barab´ asi, A.L., Sinatra, R., Lehmann, S.: The chaperone effect in scientific publishing. Proceedings of the National Academy of Sciences 115(50), 12603–12607 (2018)

  8. [16]

    Journal of Informetrics 8(4), 963–971 (2014)

    Thelwall, M., Wilson, P.: Regression for citation data: An evaluation of different methods. Journal of Informetrics 8(4), 963–971 (2014)

  9. [17]

    Computing in Science & Engineering 14(4), 42–47 (2012)

    Vandewalle, P.: Code sharing is associated with research impact in image process- ing. Computing in Science & Engineering 14(4), 42–47 (2012)

  10. [18]

    Scien- tific data 3 (2016)

    Wilkinson, M.D., Dumontier, M., Aalbersberg, I.J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.W., da Silva Santos, L.B., Bourne, P.E., et al.: The fair guiding principles for scientific data management and stewardship. Scien- tific data 3 (2016)

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.