Pith. sign in

REVIEW 4 major objections 5 minor 21 references

The Power of Data Communities

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read MIMIC, a low-cost open clinical dataset, produces far more citation impact per dollar than far more expensive controlled-access repositories, the paper claims.

desk verdict A useful descriptive comparison of four health data repositories, but the per-dollar efficiency headline rests on an incomplete funding denominator and the 'data community' attribution is untested. read the letter →

arxiv 2508.20120 v1 pith:XKJH6UHZ submitted 2025-08-22 cs.DL

classification cs.DL
keywords datacommunitiesMIMICopenclinicalbibliometricsdataseth-indexfundingefficiencysharinghealthequity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that the way a medical dataset is shared and cultivated matters as much as its size or cost. Comparing MIMIC—an open intensive-care dataset run through PhysioNet and the MIT Critical Data community—with UK Biobank, OpenSAFELY, and All of Us, the authors find that MIMIC, funded at about $14.4 million, generates more citations per dollar than comparators costing tens of millions to billions. They attribute the gap to an "accessible data community" model: free de-identified data, shared code, datathons, and deliberate outreach that draws in researchers from low- and middle-income countries. If the claim holds, modest investments in community-building around open clinical data could yield outsized research return and broader participation.

What carries the argument

The comparison rests on three linked measures: cumulative citation counts gathered from Google Scholar for each dataset's original publication; a dataset h-index, defined as the largest h such that h papers citing the dataset each have at least h citations; and a funding-efficiency ratio that divides citations or h-index by every $1 million of compiled public funding. MIMIC's funding total of $14,427,192 is the pivotal denominator; the ratio turns small absolute impact advantages into large efficiency gaps.

What would settle it

Independently audit MIMIC's full cost base—including MIT Laboratory for Computational Physiology salaries, PhysioNet hosting, infrastructure, and administrative overhead—and recompute citations per $1 million. Given the reported 452 citations per $1 million for MIMIC versus 14 for UK Biobank, a true total cost above roughly $465 million would eliminate the advantage over UK Biobank.

Watch

Extended reading notes

Core claim

The central claim is that MIMIC, despite limited funding, delivers higher impact per dollar spent through accessible data communities. Using Google Scholar citations, a dataset-level h-index, and funding-adjusted citation counts, the authors report that MIMIC accumulates more citations per $1 million than UK Biobank, OpenSAFELY, or All of Us, and that its papers are more equitably distributed, with about 10.1% of publications from low- and middle-income countries. The paper argues the mechanism is not the data alone but the community around it—datathons, open code, transparent credentialing—which lowers barriers to entry and sustains a productive research ecosystem.

Load-bearing premise

The result assumes that MIMIC's complete cost is captured by the $14,427,192 in NIH grant funding the authors manually totaled; unreported institutional support or hosting costs could erase the efficiency gap.

Editorial extensions

If this is right

  • If the cost-efficiency claim is right, funding agencies should treat community engagement as a core component of data infrastructure, not an optional add-on.
  • Open, low-cost datasets with active communities could be a higher-return route to clinical AI research than exclusive biobanks, especially for institutions with limited resources.
  • Policies that reward data sharing and community cultivation may widen participation: MIMIC-related work draws about 10.1% of its publications from low- and middle-income countries, versus 6.2% for UK Biobank.
  • The dataset-level h-index and citations-per-dollar metrics offer a simple, reusable standard for future repository evaluations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that its efficiency result is only as strong as the funding denominator: if MIMIC's true cost includes unreported MIT salaries, PhysioNet hosting, or in-kind support, the per-dollar advantage over larger repositories shrinks.
  • A sharper test of the "community, not just openness" mechanism would compare MIMIC with a similarly cheap but passively shared ICU dataset, holding data type and cost constant.
  • Citation efficiency measures scientific output, not clinical adoption; a community model could also be judged by downstream uses such as regulatory approvals, clinical deployments, or policy changes, which bibliometrics do not capture.
  • The approach could be extended to other emerging open health datasets once enough publication years accumulate, converting this single comparison into a generalizable benchmark.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper conducts a bibliometric comparison of the MIMIC dataset with UK Biobank, OpenSAFELY, and All of Us, using manually collected Google Scholar citation counts, a dataset h-index, and a funding-adjusted 'citations per $1 million' efficiency metric. It reports that MIMIC has higher citation impact per dollar than the comparators and attributes this to an 'accessible data community' model with datathons and open infrastructure. The manuscript includes a GitHub repository with data and code and is framed as evidence for investing in open clinical data communities.

Significance. If the causal attribution were valid, the paper would provide a strong argument for community-based open data infrastructure in clinical research. The raw citation data are collected and shared transparently, and the idea of normalizing citation impact by funding is a useful descriptive exercise. However, the central claim—that MIMIC's higher per-dollar impact arises through data communities—is not identified by the study design. The four datasets differ in domain, access cost, age, and community engagement simultaneously; the funding denominator for MIMIC is incomplete; and MIMIC citations are summed across multiple versions while the comparators are single papers. These issues are load-bearing, not cosmetic.

major comments (4)
  1. [§2 Data Sources and Selection; Abstract; Discussion] The causal attribution to 'data communities' is unsupported by the comparison. MIMIC differs from UK Biobank, OpenSAFELY, and All of Us in access cost, data domain, dataset age, and community infrastructure at the same time. No dataset is both freely accessible and lacking an organized community, so the marginal effect of community engagement is unidentified. The paper explicitly excludes AmsterdamUMCdb and HiRID in §2, including AmsterdamUMCdb, which is a freely accessible ICU dataset without a comparable datathon/community program. Without a free-access/no-community control or an alternative identification strategy, the Abstract's claim that MIMIC achieved higher impact 'through accessible data communities' is a post hoc interpretation, not a finding.
  2. [§2 Data Retrieval; Table 1] The per-dollar efficiency result is directly determined by a manually assembled funding denominator. MIMIC's total funding is listed as $14,427,192 from NIH grants 2003–2023 only; this excludes institutional support, salaries, PhysioNet hosting costs, and non-NIH sources. Because the headline result is a ratio of citations to this figure, an incomplete denominator can change the ranking. The authors need a sensitivity analysis using alternative funding bounds—for example, adding PhysioNet operating costs, using total Laboratory for Computational Physiology budgets, or inflation-adjusting all grants—and reporting how the efficiency gap changes.
  3. [§2 Data Retrieval] MIMIC citations are summed across the original publications for MIMIC-I, II, III, and IV, while UK Biobank, OpenSAFELY, and All of Us are each represented by one original paper. This makes the cumulative citation counts non-comparable. The paper states this was done because the versions come from the same research group, but the result is that the MIMIC total can include citations to four separate papers. A consistent comparison should use a single MIMIC release (e.g., MIMIC-III from 2016) or sum all relevant comparator publications over the same window.
  4. [Results §3; Discussion §4 (Limitations)] The efficiency comparison ignores the very different time spans and funding periods: MIMIC has 27 years of availability versus 4–9 years for the comparators, and the funding streams are from different eras without inflation adjustment. The manuscript acknowledges dataset heterogeneity and version bias in the limitations but does not flag the missing free-access/no-community control or the uncertainty in the funding denominator. These omissions are not optional caveats; they are central to the validity of the paper's conclusion.
minor comments (5)
  1. [§2 Data Retrieval] The choice of an eight-year window for MIMIC because 'citations drop significantly since it does not include MIMIC-III anymore' is arbitrary and should be justified by a data-driven rule or replaced by a fixed same-length window for all datasets.
  2. [§2 H-index Calculation for datasets] The dataset h-index ranks papers that cite the dataset's original publication by their own citation counts. This measures the impact of the citing literature, not the dataset itself, and is not directly comparable across fields with different citation cultures. It should be de-emphasized or redefined.
  3. [Table 1] Funding figures are described only as 'manually compiled from publicly available sources.' The manuscript should provide a supplementary table with source URLs, access dates, and the exact grant entries for each dataset, including currency conversion rates for UK Biobank.
  4. [Discussion §4] The LMIC participation percentages (10.1% for MIMIC vs. 6.2% for UK Biobank) are cited from reference 4, not computed in this study. The text should clearly label them as prior published statistics rather than results of the present analysis.
  5. [Figures 1b, 2b, 3b] The ratio plots use a log-10 scale, which visually amplifies differences when the denominator is small. Since the MIMIC funding figure is the smallest and the least certain, the plots should include confidence intervals or a sensitivity band that reflects the uncertainty in the funding denominator.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the per-dollar efficiency numbers are arithmetic ratios of external citation counts and compiled funding totals; the causal community attribution is confounded but not circular.

full rationale

The paper's derivation chain is transparent: (1) citation counts are taken from Google Scholar for each dataset's original publication; (2) funding totals are manually compiled from public grant records; (3) efficiency metrics are computed by dividing citations (or dataset h-index) by every $1 million in funding; (4) community-engagement characteristics are assigned qualitatively in Table 1. Each step is a direct arithmetic transformation of external inputs. No parameter is fitted to a subset of the data and then reported as a prediction; the 'impact per dollar' figures are not forecasts but ratios. The central claim that MIMIC achieves higher impact per dollar is therefore a descriptive comparison of external citation data with the authors' compiled funding denominators, not a conclusion that reduces to its inputs by definition. The only same-group citation (ref. 4) provides LMIC-participation percentages and is a peer-reviewed external study; it is not used in the per-dollar calculation. Concerns about an incomplete MIMIC funding denominator or the absence of a free-access/no-community comparator are real threats to causal attribution and measurement validity, but they are not circularity: the result does not assume what it purports to show, and the paper does not rely on an unverified self-citation to derive its headline numbers. Hence no circular step is exhibited, and the appropriate score is 0.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on publication and funding variables and on normalization choices, not on fitted constants or new entities. The main burden is the completeness and comparability of funding data and the validity of Google Scholar citation counts.

free parameters (3)
  • MIMIC citation summing window = 8 years
    Chosen because citations drop after this point when MIMIC-III is excluded; this post hoc choice affects cumulative citation comparisons.
  • Funding normalization denominator = $1 million per dataset
    Citations are divided by every $1 million of funding; this defines the efficiency metric and determines the headline ranking.
  • Dataset age normalization = years since release per dataset
    H-index and citation counts are normalized by dataset age; this choice changes relative rankings across datasets with different release years.
assumptions (5)
  • domain assumption Google Scholar 'Cited by' counts accurately measure scientific impact of a dataset.
    All citation outcomes are built from manual Google Scholar counts; no validation against Scopus or Web of Science is provided.
  • domain assumption Total funding for each dataset is complete and comparable as compiled from public sources.
    MIMIC funding counts only NIH grants 2003-2023, while other repositories have different accounting; if MIMIC's full cost is higher, the per-dollar result weakens.
  • domain assumption Summing citations across MIMIC versions I-IV is comparable to single-publication citation counts of other datasets.
    The comparison treats MIMIC as one continuous dataset even though it has multiple version papers; the paper itself flags this as 'version related bias'.
  • domain assumption Citation counts can be normalized by dataset age to make datasets with different release years comparable.
    The paper normalizes by years since release and by total funding, but does not justify that these normalizations remove age and scale effects.
  • domain assumption Differences in citation rates reflect community engagement rather than dataset topic, size, or citation culture.
    The Discussion attributes differences to the MIT Critical Data community; the analysis does not control for COVID-19 focus (OpenSAFELY) or disease coverage.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Power of Data Communities." pith.science (2026). https://pith.science/paper/XKJH6UHZ

@misc{pith2026250820120,
  author       = {Pith},
  title        = {Pith review of: The Power of Data Communities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XKJH6UHZ}},
  note         = {Machine review of arXiv:2508.20120}
}
read the original abstract

Datasets together with active scientific communities prepared to leverage them can contribute to scientific progress and facilitate making research more equitable. In this study we found that MIMIC, despite its limited amount of funding, managed to provide higher impact per dollar spent through accessible data communities. These findings support the notion that making clinical data available empowers innovation which directly addresses clinical concerns and can set new standards for inclusivity.

Figures

Figures reproduced from arXiv: 2508.20120 by the authors.

Figure 1
Figure 1. The number of cumulative citations for each dataset for each year since publication. a) unadjusted, b) number of citations adjusted per $1 million funding, shown on the log-10 scale. The data in the figures start at 1-year post-publication since no citations are available at year 0 and cannot be represented on the y-axis log plot. The dataset h-index, defined as the largest h where h publications citing the dataset … view at source ↗
Figure 2
Figure 2. Number of papers referencing each dataset ranked by number of citations within each paper. a) unadjusted, b) adjusted per $1 million funding received. The distribution of citations was analyzed by ranking papers that cited each dataset from highest to lowest citation count. In order to compare datasets with differing numbers of publications, the x-axis was normalized as a percentage of the total number of papers for… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 20 canonical work pages

  1. [1]

    & Sepp, T

    Tedersoo, L., Küngas, R., Oras, E., Köster, K., Eenmaa, H., Leijen, Ä., Pedaste, M., Raju, M., Astapova, A., Lukner, H., Kogermann, K. & Sepp, T. Data sharing practices and data availability upon request differ across scientific disciplines. Sci Data 8, 192 (2021)

  2. [2]

    Ioannidis, J. P. A. Why Most Published Research Findings Are False. PLoS Med 2, e124 (2005)

  3. [3]

    & Leyrat, C

    Besançon, L., Peiffer-Smadja, N., Segalas, C., Jiang, H., Masuzzo, P., Smout, C., Billy, E., Deforet, M. & Leyrat, C. Open science saves lives: lessons from the COVID-19 pandemic. BMC Medical Research Methodology 21, 117 (2021)

  4. [4]

    A., Cobanaj, M., Eber, R., Fiske, A., Gallifant, J., Li, C., Lingamallu, G., Petushkov, A

    Charpignon, M.-L., Celi, L. A., Cobanaj, M., Eber, R., Fiske, A., Gallifant, J., Li, C., Lingamallu, G., Petushkov, A. & Pierce, R. Diversity and inclusion: A hidden additional benefit of Open Data. PLOS Digit Health 3, e0000486 (2024)

  5. [5]

    A., Day, R

    Piwowar, H. A., Day, R. S. & Fridsma, D. B. Sharing Detailed Research Data Is Associated with Increased Citation Rate. PLoS ONE 2, e308 (2007)

  6. [6]

    & Rafols, I

    Hicks, D., Wouters, P., Waltman, L., De Rijcke, S. & Rafols, I. Bibliometrics: The Leiden Manifesto for research metrics. Nature 520, 429–431 (2015)

  7. [7]

    J., Lucero, Y

    Volk, C. J., Lucero, Y. & Barnas, K. Why is Data Sharing in Collaborative Natural Resource Efforts so Hard and What can We Do to Improve it? Environmental Management 53, 883–893 (2014)

  8. [8]

    Moody, G. B. & Mark, R. G. A database to support development and evaluation of intelligent intensive care monitoring. in Computers in Cardiology 1996 657–660 (IEEE, 1996). doi:10.1109/CIC.1996.542622

Show all 21 references
  1. [9]

    & Mark, R

    Saeed, M., Lieu, C., Raber, G. & Mark, R. G. MIMIC II: a massive temporal ICU patient database to support research in intelligent patient monitoring. in Computers in Cardiology 641–644 (IEEE, 2002). doi:10.1109/CIC.2002.1166854

  2. [10]

    Johnson, A. E. W., Pollard, T. J., Shen, L., Lehman, L. H., Feng, M., Ghassemi, M., Moody, B., Szolovits, P., Anthony Celi, L. & Mark, R. G. MIMIC-III, a freely accessible critical care database. Sci Data 3, 160035 (2016)

  3. [11]

    Johnson, A. E. W., Bulgarelli, L., Shen, L., Gayles, A., Shammout, A., Horng, S., Pollard, T. J., Hao, S., Moody, B., Gow, B., Lehman, L. H., Celi, L. A. & Mark, R. G. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data 10, 1 (2023)

  4. [12]

    & Collins, R

    Sudlow, C., Gallacher, J., Allen, N., Beral, V., Burton, P., Danesh, J., Downey, P., Elliott, P., Green, J., Landray, M., Liu, B., Matthews, P., Ong, G., Pell, J., Silman, A., Young, A., Sprosen, T., Peakman, T. & Collins, R. UK Biobank: An Open Access Resource for Identifying...

  5. [13]

    J., Walker, A

    Williamson, E. J., Walker, A. J., Bhaskaran, K., Bacon, S., Bates, C., Morton, C. E., Curtis, H. J., Mehrkar, A., Evans, D., Inglesby, P., Cockburn, J., McDonald, H. I., MacKenna, B., Tomlinson, L., Douglas, I. J., Rentsch, C. T., Mathur, R., Wong, A. Y. S., Grieve, R., Harris...

  6. [14]

    C., Rutter, J

    All of Us Research Program Investigators, Denny, J. C., Rutter, J. L., Goldstein, D. B., Philippakis, A., Smoller, J. W., Jenkins, G. & Dishman, E. The ‘All of Us’ Research Program. N Engl J Med 381, 668–676 (2019). 12

  7. [15]

    J., Peppink, J

    Thoral, P. J., Peppink, J. M., Driessen, R. H., Sijbrands, E. J. G., Kompanje, E. J. O., Kaplan, L., Bailey, H., Kesecioglu, J., Cecconi, M., Churpek, M., Clermont, G., van der Schaar, M., Ercole, A., Girbes, A. R. J. & Elbers, P. W. G. Sharing ICU Patient Data Responsibly Und...

  8. [16]

    & Merz, T

    Faltys, M., Zimmermann, M., Lyu, X., Hüser, M., Hyland, S., Rätsch, G. & Merz, T. HiRID, a high time-resolution ICU dataset. doi:10.13026/NKWC-JS72

  9. [17]

    Open data for AI: what now? (United Nations Educational, Scientific and Cultural Organization, 2023)

    Ziesche, S. Open data for AI: what now? (United Nations Educational, Scientific and Cultural Organization, 2023)

  10. [18]

    G., Gerosa, M

    Steinmacher, I., Balali, S., Trinkenreich, B., Guizani, M., Izquierdo-Cortazar, D., Cuevas Zambrano, G. G., Gerosa, M. A. & Sarma, A. Being a Mentor in open source projects. J Internet Serv Appl 12, 7 (2021)

  11. [19]

    & Pratt, B

    Evertsz, N., Bull, S. & Pratt, B. What constitutes equitable data sharing in global health research? A scoping review of the literature on low-income and middle-income country stakeholders’ perspectives. BMJ Glob Health 8, e010157 (2023)

  12. [20]

    A., Potter, R

    Szomszor, M., Adams, J., Fry, R., Gebert, C., Pendlebury, D. A., Potter, R. W. K. & Rogers, G. Interpreting Bibliometric Data. Front. Res. Metr. Anal. 5, 628703 (2021)

  13. [21]

    L., Hulme, W., DeVito, N

    Nab, L., Schaffer, A. L., Hulme, W., DeVito, N. J., Dillingham, I., Wiedemann, M., Andrews, C. D., Curtis, H., Fisher, L., Green, A., Massey, J., Walters, C. E., Higgins, R., Cunningham, C., Morley, J., Mehrkar, A., Hart, L., Davy, S., Evans, D., Hickman, G., Inglesby, P., Mor...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.