REVIEW 4 major objections 5 minor 21 references
The Power of Data Communities
T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read MIMIC, a low-cost open clinical dataset, produces far more citation impact per dollar than far more expensive controlled-access repositories, the paper claims.
desk verdict A useful descriptive comparison of four health data repositories, but the per-dollar efficiency headline rests on an incomplete funding denominator and the 'data community' attribution is untested. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The comparison rests on three linked measures: cumulative citation counts gathered from Google Scholar for each dataset's original publication; a dataset h-index, defined as the largest h such that h papers citing the dataset each have at least h citations; and a funding-efficiency ratio that divides citations or h-index by every $1 million of compiled public funding. MIMIC's funding total of $14,427,192 is the pivotal denominator; the ratio turns small absolute impact advantages into large efficiency gaps.
What would settle it
Independently audit MIMIC's full cost base—including MIT Laboratory for Computational Physiology salaries, PhysioNet hosting, infrastructure, and administrative overhead—and recompute citations per $1 million. Given the reported 452 citations per $1 million for MIMIC versus 14 for UK Biobank, a true total cost above roughly $465 million would eliminate the advantage over UK Biobank.
Extended reading notes
Core claim
The central claim is that MIMIC, despite limited funding, delivers higher impact per dollar spent through accessible data communities. Using Google Scholar citations, a dataset-level h-index, and funding-adjusted citation counts, the authors report that MIMIC accumulates more citations per $1 million than UK Biobank, OpenSAFELY, or All of Us, and that its papers are more equitably distributed, with about 10.1% of publications from low- and middle-income countries. The paper argues the mechanism is not the data alone but the community around it—datathons, open code, transparent credentialing—which lowers barriers to entry and sustains a productive research ecosystem.
Load-bearing premise
The result assumes that MIMIC's complete cost is captured by the $14,427,192 in NIH grant funding the authors manually totaled; unreported institutional support or hosting costs could erase the efficiency gap.
Editorial extensions
If this is right
- If the cost-efficiency claim is right, funding agencies should treat community engagement as a core component of data infrastructure, not an optional add-on.
- Open, low-cost datasets with active communities could be a higher-return route to clinical AI research than exclusive biobanks, especially for institutions with limited resources.
- Policies that reward data sharing and community cultivation may widen participation: MIMIC-related work draws about 10.1% of its publications from low- and middle-income countries, versus 6.2% for UK Biobank.
- The dataset-level h-index and citations-per-dollar metrics offer a simple, reusable standard for future repository evaluations.
Reading between the lines
- The paper leaves implicit that its efficiency result is only as strong as the funding denominator: if MIMIC's true cost includes unreported MIT salaries, PhysioNet hosting, or in-kind support, the per-dollar advantage over larger repositories shrinks.
- A sharper test of the "community, not just openness" mechanism would compare MIMIC with a similarly cheap but passively shared ICU dataset, holding data type and cost constant.
- Citation efficiency measures scientific output, not clinical adoption; a community model could also be judged by downstream uses such as regulatory approvals, clinical deployments, or policy changes, which bibliometrics do not capture.
- The approach could be extended to other emerging open health datasets once enough publication years accumulate, converting this single comparison into a generalizable benchmark.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper conducts a bibliometric comparison of the MIMIC dataset with UK Biobank, OpenSAFELY, and All of Us, using manually collected Google Scholar citation counts, a dataset h-index, and a funding-adjusted 'citations per $1 million' efficiency metric. It reports that MIMIC has higher citation impact per dollar than the comparators and attributes this to an 'accessible data community' model with datathons and open infrastructure. The manuscript includes a GitHub repository with data and code and is framed as evidence for investing in open clinical data communities.
Significance. If the causal attribution were valid, the paper would provide a strong argument for community-based open data infrastructure in clinical research. The raw citation data are collected and shared transparently, and the idea of normalizing citation impact by funding is a useful descriptive exercise. However, the central claim—that MIMIC's higher per-dollar impact arises through data communities—is not identified by the study design. The four datasets differ in domain, access cost, age, and community engagement simultaneously; the funding denominator for MIMIC is incomplete; and MIMIC citations are summed across multiple versions while the comparators are single papers. These issues are load-bearing, not cosmetic.
major comments (4)
- [§2 Data Sources and Selection; Abstract; Discussion] The causal attribution to 'data communities' is unsupported by the comparison. MIMIC differs from UK Biobank, OpenSAFELY, and All of Us in access cost, data domain, dataset age, and community infrastructure at the same time. No dataset is both freely accessible and lacking an organized community, so the marginal effect of community engagement is unidentified. The paper explicitly excludes AmsterdamUMCdb and HiRID in §2, including AmsterdamUMCdb, which is a freely accessible ICU dataset without a comparable datathon/community program. Without a free-access/no-community control or an alternative identification strategy, the Abstract's claim that MIMIC achieved higher impact 'through accessible data communities' is a post hoc interpretation, not a finding.
- [§2 Data Retrieval; Table 1] The per-dollar efficiency result is directly determined by a manually assembled funding denominator. MIMIC's total funding is listed as $14,427,192 from NIH grants 2003–2023 only; this excludes institutional support, salaries, PhysioNet hosting costs, and non-NIH sources. Because the headline result is a ratio of citations to this figure, an incomplete denominator can change the ranking. The authors need a sensitivity analysis using alternative funding bounds—for example, adding PhysioNet operating costs, using total Laboratory for Computational Physiology budgets, or inflation-adjusting all grants—and reporting how the efficiency gap changes.
- [§2 Data Retrieval] MIMIC citations are summed across the original publications for MIMIC-I, II, III, and IV, while UK Biobank, OpenSAFELY, and All of Us are each represented by one original paper. This makes the cumulative citation counts non-comparable. The paper states this was done because the versions come from the same research group, but the result is that the MIMIC total can include citations to four separate papers. A consistent comparison should use a single MIMIC release (e.g., MIMIC-III from 2016) or sum all relevant comparator publications over the same window.
- [Results §3; Discussion §4 (Limitations)] The efficiency comparison ignores the very different time spans and funding periods: MIMIC has 27 years of availability versus 4–9 years for the comparators, and the funding streams are from different eras without inflation adjustment. The manuscript acknowledges dataset heterogeneity and version bias in the limitations but does not flag the missing free-access/no-community control or the uncertainty in the funding denominator. These omissions are not optional caveats; they are central to the validity of the paper's conclusion.
minor comments (5)
- [§2 Data Retrieval] The choice of an eight-year window for MIMIC because 'citations drop significantly since it does not include MIMIC-III anymore' is arbitrary and should be justified by a data-driven rule or replaced by a fixed same-length window for all datasets.
- [§2 H-index Calculation for datasets] The dataset h-index ranks papers that cite the dataset's original publication by their own citation counts. This measures the impact of the citing literature, not the dataset itself, and is not directly comparable across fields with different citation cultures. It should be de-emphasized or redefined.
- [Table 1] Funding figures are described only as 'manually compiled from publicly available sources.' The manuscript should provide a supplementary table with source URLs, access dates, and the exact grant entries for each dataset, including currency conversion rates for UK Biobank.
- [Discussion §4] The LMIC participation percentages (10.1% for MIMIC vs. 6.2% for UK Biobank) are cited from reference 4, not computed in this study. The text should clearly label them as prior published statistics rather than results of the present analysis.
- [Figures 1b, 2b, 3b] The ratio plots use a log-10 scale, which visually amplifies differences when the denominator is small. Since the MIMIC funding figure is the smallest and the least certain, the plots should include confidence intervals or a sensitivity band that reflects the uncertainty in the funding denominator.
Circularity Check
No significant circularity: the per-dollar efficiency numbers are arithmetic ratios of external citation counts and compiled funding totals; the causal community attribution is confounded but not circular.
full rationale
The paper's derivation chain is transparent: (1) citation counts are taken from Google Scholar for each dataset's original publication; (2) funding totals are manually compiled from public grant records; (3) efficiency metrics are computed by dividing citations (or dataset h-index) by every $1 million in funding; (4) community-engagement characteristics are assigned qualitatively in Table 1. Each step is a direct arithmetic transformation of external inputs. No parameter is fitted to a subset of the data and then reported as a prediction; the 'impact per dollar' figures are not forecasts but ratios. The central claim that MIMIC achieves higher impact per dollar is therefore a descriptive comparison of external citation data with the authors' compiled funding denominators, not a conclusion that reduces to its inputs by definition. The only same-group citation (ref. 4) provides LMIC-participation percentages and is a peer-reviewed external study; it is not used in the per-dollar calculation. Concerns about an incomplete MIMIC funding denominator or the absence of a free-access/no-community comparator are real threats to causal attribution and measurement validity, but they are not circularity: the result does not assume what it purports to show, and the paper does not rely on an unverified self-citation to derive its headline numbers. Hence no circular step is exhibited, and the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- MIMIC citation summing window =
8 years
- Funding normalization denominator =
$1 million per dataset
- Dataset age normalization =
years since release per dataset
assumptions (5)
- domain assumption Google Scholar 'Cited by' counts accurately measure scientific impact of a dataset.
- domain assumption Total funding for each dataset is complete and comparable as compiled from public sources.
- domain assumption Summing citations across MIMIC versions I-IV is comparable to single-publication citation counts of other datasets.
- domain assumption Citation counts can be normalized by dataset age to make datasets with different release years comparable.
- domain assumption Differences in citation rates reflect community engagement rather than dataset topic, size, or citation culture.
Cite this review
Pith. "Pith review of The Power of Data Communities." pith.science (2026). https://pith.science/paper/XKJH6UHZ
@misc{pith2026250820120,
author = {Pith},
title = {Pith review of: The Power of Data Communities},
year = {2026},
howpublished = {\url{https://pith.science/paper/XKJH6UHZ}},
note = {Machine review of arXiv:2508.20120}
}
read the original abstract
Datasets together with active scientific communities prepared to leverage them can contribute to scientific progress and facilitate making research more equitable. In this study we found that MIMIC, despite its limited amount of funding, managed to provide higher impact per dollar spent through accessible data communities. These findings support the notion that making clinical data available empowers innovation which directly addresses clinical concerns and can set new standards for inclusivity.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Ioannidis, J. P. A. Why Most Published Research Findings Are False. PLoS Med 2, e124 (2005)
work page 2005
-
[3]
Besançon, L., Peiffer-Smadja, N., Segalas, C., Jiang, H., Masuzzo, P., Smout, C., Billy, E., Deforet, M. & Leyrat, C. Open science saves lives: lessons from the COVID-19 pandemic. BMC Medical Research Methodology 21, 117 (2021)
work page 2021
-
[4]
A., Cobanaj, M., Eber, R., Fiske, A., Gallifant, J., Li, C., Lingamallu, G., Petushkov, A
Charpignon, M.-L., Celi, L. A., Cobanaj, M., Eber, R., Fiske, A., Gallifant, J., Li, C., Lingamallu, G., Petushkov, A. & Pierce, R. Diversity and inclusion: A hidden additional benefit of Open Data. PLOS Digit Health 3, e0000486 (2024)
work page 2024
-
[5]
Piwowar, H. A., Day, R. S. & Fridsma, D. B. Sharing Detailed Research Data Is Associated with Increased Citation Rate. PLoS ONE 2, e308 (2007)
work page 2007
-
[6]
Hicks, D., Wouters, P., Waltman, L., De Rijcke, S. & Rafols, I. Bibliometrics: The Leiden Manifesto for research metrics. Nature 520, 429–431 (2015)
work page 2015
-
[7]
Volk, C. J., Lucero, Y. & Barnas, K. Why is Data Sharing in Collaborative Natural Resource Efforts so Hard and What can We Do to Improve it? Environmental Management 53, 883–893 (2014)
work page 2014
- [8]
Show all 21 references
-
[9]
& Mark, R
Saeed, M., Lieu, C., Raber, G. & Mark, R. G. MIMIC II: a massive temporal ICU patient database to support research in intelligent patient monitoring. in Computers in Cardiology 641–644 (IEEE, 2002). doi:10.1109/CIC.2002.1166854
2002 arXiv
-
[10]
Johnson, A. E. W., Pollard, T. J., Shen, L., Lehman, L. H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Anthony Celi, L. & Mark, R. G. MIMIC-III, a freely accessible critical care database. Sci Data 3, 160035 (2016)
2016
-
[11]
Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L. H., Celi, L. A. & Mark, R. G. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data 10, 1 (2023)
2023
-
[12]
& Collins, R
Sudlow, C., Gallacher, J., Allen, N., Beral, V., Burton, P., Danesh, J., Downey, P., Elliott, P., Green, J., Landray, M., Liu, B., Matthews, P., Ong, G., Pell, J., Silman, A., Young, A., Sprosen, T., Peakman, T. & Collins, R. UK Biobank: An Open Access Resource for Identifying...
2015
-
[13]
J., Walker, A
Williamson, E. J., Walker, A. J., Bhaskaran, K., Bacon, S., Bates, C., Morton, C. E., Curtis, H. J., Mehrkar, A., Evans, D., Inglesby, P., Cockburn, J., McDonald, H. I., MacKenna, B., Tomlinson, L., Douglas, I. J., Rentsch, C. T., Mathur, R., Wong, A. Y. S., Grieve, R., Harris...
2020
-
[14]
C., Rutter, J
All of Us Research Program Investigators, Denny, J. C., Rutter, J. L., Goldstein, D. B., Philippakis, A., Smoller, J. W., Jenkins, G. & Dishman, E. The ‘All of Us’ Research Program. N Engl J Med 381, 668–676 (2019). 12
2019
-
[15]
J., Peppink, J
Thoral, P. J., Peppink, J. M., Driessen, R. H., Sijbrands, E. J. G., Kompanje, E. J. O., Kaplan, L., Bailey, H., Kesecioglu, J., Cecconi, M., Churpek, M., Clermont, G., van der Schaar, M., Ercole, A., Girbes, A. R. J. & Elbers, P. W. G. Sharing ICU Patient Data Responsibly Und...
2021
-
[16]
& Merz, T
Faltys, M., Zimmermann, M., Lyu, X., Hüser, M., Hyland, S., Rätsch, G. & Merz, T. HiRID, a high time-resolution ICU dataset. doi:10.13026/NKWC-JS72
-
[17]
Open data for AI: what now? (United Nations Educational, Scientific and Cultural Organization, 2023)
Ziesche, S. Open data for AI: what now? (United Nations Educational, Scientific and Cultural Organization, 2023)
2023
-
[18]
G., Gerosa, M
Steinmacher, I., Balali, S., Trinkenreich, B., Guizani, M., Izquierdo-Cortazar, D., Cuevas Zambrano, G. G., Gerosa, M. A. & Sarma, A. Being a Mentor in open source projects. J Internet Serv Appl 12, 7 (2021)
2021
-
[19]
& Pratt, B
Evertsz, N., Bull, S. & Pratt, B. What constitutes equitable data sharing in global health research? A scoping review of the literature on low-income and middle-income country stakeholders’ perspectives. BMJ Glob Health 8, e010157 (2023)
2023
-
[20]
A., Potter, R
Szomszor, M., Adams, J., Fry, R., Gebert, C., Pendlebury, D. A., Potter, R. W. K. & Rogers, G. Interpreting Bibliometric Data. Front. Res. Metr. Anal. 5, 628703 (2021)
2021
-
[21]
L., Hulme, W., DeVito, N
Nab, L., Schaffer, A. L., Hulme, W., DeVito, N. J., Dillingham, I., Wiedemann, M., Andrews, C. D., Curtis, H., Fisher, L., Green, A., Massey, J., Walters, C. E., Higgins, R., Cunningham, C., Morley, J., Mehrkar, A., Hart, L., Davy, S., Evans, D., Hickman, G., Inglesby, P., Mor...
2024
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.