{"id":"7ce5bacc-33eb-4e58-b29f-372b06e20780","arxiv_id":"1908.06830","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Across 500 sampled MICCAI papers (2014-2018), 54.2% used only private data, while public-data use was associated with a 60.8% citation-per-year advantage after controlling for open access and code release.","lead":"This study reviewed 500 machine-learning papers from the MICCAI medical image computing conference and found that more than half relied only on private data. Papers that used public datasets were cited about 60% more per year, and the authors argue for policies that reward data sharing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unmeasured author reputation, acknowledged in §4.3, is the main threat: the 60.8% public-data citation advantage may be mostly a senior-author effect, and only a reputation-adjusted reanalysis can settle it.","rationale":"I read this as a careful, transparent bibliometric study with openly released data and code, and the reader's CONDITIONAL verdict is appropriate. The reader identified manual coding reliability as the weakest assumption, and that is a real concern: the public/private data classification and one-time Google Scholar counts are unvalidated by a second coder, and systematic coding errors could bias both the prevalence and the citation-ratio estimates. My stress-test, however, places the most weight on a different load-bearing point: even with perfect coding, the 60.8% association is adjusted only for open access and code release, and the paper itself acknowledges the omission of author reputation. Because reputation is strongly associated with both citation counts and the propensity to use public benchmarks, this unmeasured confounder could explain a substantial part of the effect. The paper is appropriately cautious and stops short of causal language, so the literal association claim may survive, but the title and policy recommendations treat the association as the 'role' of public data, which requires the reputation-adjusted estimate to remain nontrivial. A single reanalysis with author-reputation covariates would settle this. Since the reader's conditional verdict already captures the need for such verification, I recommend no change to the verdict, but I would ground the conditionality in the reputation confounder as much as in coding reliability.","tokens_in":5781,"tokens_out":11380,"duration_ms":129839,"concrete_test":"Using the released data and code (github.com/neheller/labels19), add a measure of senior-author reputation at publication time (e.g., last author's h-index or prior-5-year citation count) and re-estimate the public-vs-private citation ratio with year and paper type as additional strata or covariates. A practical benchmark: if the adjusted ratio's 95% CI includes 1 or the point estimate falls below about 1.2, the concern lands; if it remains near 1.6 with CI excluding 1, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim is that, after stratifying by open access and code release, papers using public data receive 60.8% more citations per year (95% CI 28.1%–110.2%; §4.3). The estimate is descriptive, not causal, but the paper's title and policy recommendations interpret it as evidence of the role of public data. The weakest link is that only two binary confounders are controlled. Author reputation is a known, powerful predictor of citations (ref. [15]) and is plausibly correlated with public-data use: established labs are more likely to build on benchmarks and to attract citations for any paper. The authors explicitly concede in §4.3 that they could not control for author reputation. If reputation explains even half of the observed ratio, the 60.8% effect cannot support the stated role of public data. Year of publication and paper type (challenge vs regular) are additional unmodeled factors, and the 'anomalous' 2017 private-data spike is unexplained. None of this makes the association false, but it makes the headline estimate's interpretation insecure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript analyzes 500 MICCAI papers from 2014 to 2018 that use machine learning for computer vision tasks, manually labelling each paper for data usage (public, released, or private-only), code release, open access status, and citation counts. The authors report three main findings: (1) 54.2% of the sampled papers used only private data, with a decline from 64% to 44% across the five years; (2) after stratifying by open access and code release, papers using public data received 60.8% more citations per year than private-data-only papers (95% CI 28.1%–110.2%); and (3) 21.6% of papers using public data did not cite the dataset itself. The paper concludes with policy recommendations for MICCAI, including a data availability statement requirement and stronger reviewer scrutiny of data references.","tokens_in":5977,"tokens_out":2685,"duration_ms":30101,"significance":"If the headline association is taken at face value, it is a novel and useful contribution to the bibliometric literature on data sharing in medical image computing, complementing prior work by Piwowar and Vision, Drachen et al., and Colavizza et al. The analysis is transparent and reproducible: the authors make their code and data publicly available, describe their manual protocol, use a stratified ratio estimator with a bootstrap confidence interval, and explicitly report a winsorization robustness check. The paper is also careful to stop short of claiming causality, acknowledging the unmeasured confounder of author reputation. The main value lies in quantifying a large disparity in citation impact between public-data and private-data papers, which is relevant to MICCAI policy discussions and to researchers studying reproducibility incentives.","major_comments":[{"comment":"The central estimate of a 60.8% citation advantage for public-data papers is plausibly confounded by author reputation, a concern the authors explicitly acknowledge but do not address quantitatively. Since the paper's title and policy recommendations lean on this estimate, the manuscript should either control for reputation (for example, by stratifying on the senior author's prior citation record or h-index) or perform a sensitivity analysis showing how strongly a reputation effect would need to be to explain the observed ratio. Without such analysis, the interpretation should be substantially softened to a purely descriptive association, and the recommendations should be framed as conditional on this limitation.","section":"§4.3"},{"comment":"All outcome and exposure variables—data usage category, code release, open access, and citation counts—were assigned through manual review by the authors, but no inter-rater reliability is reported. If coding errors are correlated with citation counts or with public-data status, the estimated ratio and the reported prevalence could be biased. The authors should provide a codebook and have a second coder independently label a random subsample (e.g., 20% of the 500 papers), reporting agreement statistics such as Cohen's kappa, and show that disagreements are not associated with the outcome. This is a load-bearing issue because the entire analysis depends on the accuracy of these manual labels.","section":"§3.1"}],"minor_comments":[{"comment":"The date on which Google Scholar citation counts were collected is not stated. Because citation counts change over time and the study spans five publication years, the authors should report the collection date and, ideally, verify stability by re-collecting a subsample at a later date.","section":"§3.1"},{"comment":"The winsorization threshold of 50 citations per year is described as a robustness check, but only the direction and significance are commented on. Reporting the bootstrap confidence interval with and without the two affected papers, or with alternative thresholds (e.g., 25 and 100), would make the robustness claim more concrete.","section":"§3.2"},{"comment":"The section title says \"More than quarter of data references were not citations,\" but the reported proportion is 21.6%, which is less than one quarter. The title should be corrected to \"More than one in five\" or the reported percentage should be replaced with the actual percentage of 21.6% in a correctly worded heading.","section":"§4.4"},{"comment":"The sentence \"papers based on public data were cited over 60% more per year\" appears without the confidence interval in the abstract and in the opening of Section 4.3. Stating the interval (28.1%–110.2%) in both places would help readers assess the precision of the estimate.","section":"§4.3"},{"comment":"The notable increase in private-data-only papers in 2017 is described as anomalous, but no explanation or analysis is offered. A brief discussion of possible reasons (e.g., conference theme, author pool, or data collection artifact) would strengthen the temporal description.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a straightforward bibliometric study with a clearly stated methodology and publicly available code and data. The main concern is that the headline claim—even though carefully labeled as an association—is used to support policy recommendations, and the text does not sufficiently address the magnitude of possible confounding by author reputation. In my view this is fixable within the manuscript's scope by adding a sensitivity analysis or by explicitly reducing the strength of the conclusions. I do not see a basis for rejection, but the current version is not ready for acceptance as-is."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nick: You should know two things about this one. First, it is the first bibliometric study I know of that measures the citation advantage of simply using public data (rather than releasing your own) in a medical-imaging ML venue, and it ships its code and data. Second, its headline number — 60.8% more citations per year for public-data users after stratifying by open access and code release — is a credible descriptive estimate, but the authors cannot control for author reputation, and they say so. Read it as association-level evidence, not as proof that public data causes citations.\n\nThe paper does real work. They manually classified 500 MICCAI CV/ML papers into public/released/private, recorded citation counts from Google Scholar, and computed a stratified ratio estimate with bootstrap CIs. The Winsorization robustness check is reported and it does not change the result. The 54.2% private-only prevalence, the 21.6% non-citation data references, and the 5% of datasets with no citable entity are all useful concrete numbers for the medical-imaging community. The policy recommendations (DAS requirement, reviewer instructions on data citations) follow from the data without overreach.\n\nSoft spots, in order of real weight. The manual coding has no inter-rater reliability estimate; one coder, one pass. Google Scholar counts are a single snapshot, so citation trajectories are treated as though they matured at the same rate regardless of publication year. The trend claims (64% down to 44%, code release 6% to 9%) have no significance tests. The 2017 private-data spike is labeled anomalous and left unexplained. And the author-reputation issue is the one that matters most: if senior labs are more likely to build on public benchmarks and to accrue citations regardless, the 60.8% ratio is partly a senior-author effect. The authors acknowledge this explicitly in §4.3, which is more than most papers in this genre do, but it is still the ceiling on how strongly the headline can be interpreted.\n\nThe stress-test note is right that reputation is the main threat, and it is not a fatal objection to the paper's stated scope — the paper claims an association, not a causal effect. The reader's conditional verdict is fair. I would send this to a serious referee. It is a transparent, reproducible measurement that the MICCAI community should see, and the limitations are mostly known and stated. For a revision, I would ask for a second coder on at least a subsample and a sensitivity analysis that includes author seniority or prior citation history if the authors can get it. The citation pattern in the related work is fair — they cite the obvious prior data-sharing citation-advantage studies and note the gap in the literature.\n\nWho gets value from this: anyone who writes or reviews policy about data sharing at biomedical conferences, and bibliometrics people who want a ready-made comparison point. I would bring it to reading group and I would cite it. It deserves peer review, not a desk reject.","headline":"A transparent, well-scoped measurement of public-data use and citation outcomes at MICCAI; the 60.8% association is credible as a descriptive estimate but should not be read as causal.","tokens_in":6503,"tokens_out":2231,"would_cite":true,"duration_ms":21067,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that using public datasets is associated with 60.8% more citations per year for medical image computing papers, after stratifying by code release and open access.","keywords":["public data","citation advantage","private data","medical image computing","data sharing","reproducibility","bibliometrics","code release"],"falsifier":"Have two independent, blinded coders re-classify the same 500 papers using the paper's own definitions, and retrieve citation counts at a later date; if inter-coder agreement on data status is low, or if re-measured citation counts shrink the public-data advantage toward zero, the central association fails.","tokens_in":5591,"feed_emoji":"📊","tokens_out":5902,"duration_ms":54853,"temperature":0.7,"pith_summary":"The paper sets out to measure how often medical image computing research is built on datasets that anyone can download, and whether that choice is associated with how often the work is cited. Reviewing a random sample of 500 accepted papers at the annual medical image computing conference between 2014 and 2018, it finds that more than half relied on private data alone, though that share fell from 64% to 44% over the period. After stratifying by code release and open-access status, papers using public data were cited about 60.8% more per year than those using only private data. The authors read this as evidence that reusable data is a major catalyst in this field and recommend policy changes, such as required data-availability statements, to make sharing the norm.","feed_headline":"Public data linked to 60% more citations in medical imaging papers","feed_subtitle":"A 500-paper review finds private-only data in over half of studies, while public benchmarks gain ground.","key_machinery":"The central mechanism is a stratified ratio-of-means analysis. Each of the 500 papers is manually coded for three binary attributes—data type (public versus private-only), code release, and open-access status—and papers are sorted into the four groups formed by the latter two attributes. Within each group the mean citations per year of public-data papers is divided by that of private-data papers, and the four ratios are combined into a weighted average by group prevalence. Winsorizing citation rates at 50 per year protects against a handful of extremely cited papers, and the bootstrap supplies the confidence interval around the final 60.8% estimate.","core_discovery":"On the paper's own terms, the discovery is a quantitative association: among accepted computer-vision and machine-learning papers at the annual medical image computing conference from 2014 to 2018, using a publicly available dataset (whether an existing benchmark or data released with the paper) is associated with 60.8% more citations per year than using only private data, with a 95% bootstrap confidence interval of 28.1% to 110.2%, after controlling for two known confounders: code release and open-access publication. The paper also establishes that 54.2% of these papers used only private data and that this proportion declined from 64.0% in 2014 to 44% in 2018. It further reports that 21.6% of papers using public datasets did not cite the dataset in a formal way, and that in 5.0% of such cases no citable entity existed at all.","pith_inferences":["By the authors' own reasoning, a similar citation advantage likely holds in other fields where data collection dominates method development, such as clinical natural-language processing; this is an extrapolation, not something the paper measures.","If the association reflects causality, the field's citation economy systematically undervalues private-data work regardless of quality, so evaluation and funding criteria that emphasize citations would add another incentive toward data sharing.","A direct test of the paper's mechanism would be a before/after study of a conference that introduces a required data-availability statement, tracking citation rates while holding methods otherwise constant.","The coding protocol could be applied to proceedings after 2018 to test whether the private-data share and the citation gap converge, a prediction that follows from the paper's narrative but is not tested here."],"forward_implications":["If the association holds, papers that reuse public datasets receive materially more attention than methodologically similar work built on private data, making dataset choice a de facto impact lever.","A required data-availability statement, along the lines recommended by the authors, could shift the field's private-data share quickly without requiring new infrastructure.","Dataset creators who provide a citable, indexed entity remove the main structural excuse for informal data references.","The decline in private-data-only papers over the five years can be read as early evidence that the norm is already moving toward openness, so policy changes may be reinforcing rather than reversing a trend."],"supporting_citations":[{"why":"Supplies the prior estimate that releasing data raises citations by about 10%, the baseline against which this study's 60.8% effect is compared.","marker":"[13]"},{"why":"Shows a 25-40% citation increase for linked data in astrophysics, providing a field-comparison effect size.","marker":"[6]"},{"why":"Gives large-scale evidence that data-availability statements linking to repositories are associated with higher citations.","marker":"[3]"},{"why":"Demonstrates that code sharing is associated with research impact, justifying its use as a control variable.","marker":"[17]"},{"why":"Demonstrates a citation advantage for open-access articles, justifying its use as a control variable.","marker":"[9]"},{"why":"Identifies author reputation as a known citation confound that this study cannot control for, marking a stated limitation.","marker":"[15]"},{"why":"Provides the Winsorization procedure used to trim extreme citation rates.","marker":"[5]"},{"why":"Provides the bootstrap method used to estimate the confidence interval for the citation ratio.","marker":"[7]"}],"fun_headline_variants":["Public data papers get 60% more citations in MICCAI","MICCAI: public data yields 60% more citations","Half of MICCAI papers use private data, but public data cited more","MICCAI: 60% citation advantage for public data papers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result rests on the manual labels that put each paper into 'public data' or 'private data' and on one-time citation counts; if those labels or counts are systematically wrong in a way that tracks citation success, the 60.8% gap and the 54.2% private-data share could be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Public data papers get 60% more citations in MICCAI","MICCAI: public data yields 60% more citations","Half of MICCAI papers use private data, but public data cited more","MICCAI: 60% citation advantage for public data papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001053,"raw_usage":{"total_tokens":4427,"prompt_tokens":957,"completion_tokens":3470,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":3394}},"tokens_in":573,"tokens_out":3470,"duration_ms":24548,"temperature":1.0,"reasoning_tokens":3394,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:46:59.174899+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent, blinded coders re-classify the same 500 papers using the paper's own definitions, and retrieve citation counts at a later date; if inter-coder agreement on data status is low, or if re-measured citation counts shrink the public-data advantage toward zero, the central association fails.","supporting_citations":[{"cited_title":"PeerJ 1, e175 (2013)","cited_arxiv_id":null,"evidence_quote":"Supplies the prior estimate that releasing data raises citations by about 10%, the baseline against which this study's 60.8% effect is compared."},{"cited_title":"Liber Quarterly 26(2) (2016)","cited_arxiv_id":null,"evidence_quote":"Shows a 25-40% citation increase for linked data in astrophysics, providing a field-comparison effect size."},{"cited_title":"The citation advantage of linking publications to research data","cited_arxiv_id":"1907.02565","evidence_quote":"Gives large-scale evidence that data-availability statements linking to repositories are associated with higher citations."},{"cited_title":"Computing in Science & Engineering 14(4), 42–47 (2012)","cited_arxiv_id":null,"evidence_quote":"Demonstrates that code sharing is associated with research impact, justifying its use as a control variable."},{"cited_title":"PLoS biology 4(5), e157 (2006)","cited_arxiv_id":null,"evidence_quote":"Demonstrates a citation advantage for open-access articles, justifying its use as a control variable."},{"cited_title":"Proceedings of the National Academy of Sciences 115(50), 12603–12607 (2018)","cited_arxiv_id":null,"evidence_quote":"Identifies author reputation as a known citation confound that this study cannot control for, marking a stated limitation."},{"cited_title":"Statistische Hefte 15(2-3), 157–170 (1974)","cited_arxiv_id":null,"evidence_quote":"Provides the Winsorization procedure used to trim extreme citation rates."},{"cited_title":"CRC press (1994)","cited_arxiv_id":null,"evidence_quote":"Provides the bootstrap method used to estimate the confidence interval for the citation ratio."}],"review_version":1}