REVIEW 3 major objections 5 minor 1 cited by
Fostering Data Communities -- perspective from a Data Archive Service Provider
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Data communities and data archives co-evolve, each creating the conditions for the other's existence, this study argues.
desk verdict A transparent, useful provider-side case study whose central co-evolution claim outruns its evidence; worth engaging with, but the authors should narrow the claim or bring in external triangulation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the 'data station': a discipline-customised archive instance, built on a shared platform, each assigned a dedicated manager who maintains ongoing contact with the relevant research community. The manager acts as a boundary worker, translating the community's data practices into service requirements such as formats, metadata standards, and submission pipelines, and in turn translating archival expertise back to the community. The paper's evidence for this mechanism comes from an institutional autoethnography: interviews with the station manager, analysis of strategy and policy documents, and regular informal consultations inside the archive.
What would settle it
Observe a data archive that operates as a passive storage service, with no community manager and no customisation to community needs, and check whether the data community around it develops the same reciprocal dynamic. If the community's standards, identity, and service expectations evolve just as strongly without active mediation by the archive, the paper's claim that archives help create the conditions for community existence would be undercut.
Extended reading notes
Core claim
The paper's central claim is that a data community and a data archive evolve reciprocally, with the archive both responding to and shaping the community's needs and features. The case of the life sciences data station is used to show how this happens in practice: the station's manager identified long-tail research data as the niche, established project-based collaborations, and fed domain-specific metadata standards into the repository design, so the infrastructure itself became part of the community's identity and capabilities. The authors state that, despite being a single case, similar dynamics occur for other data communities and archives because the underlying topics—technology, data requirements, and community formation—are general.
Load-bearing premise
The load-bearing premise is that the experience of one national archive, described by its own staff and one community manager, is representative of how other archives and data communities co-evolve.
Editorial extensions
If this is right
- Archives seeking to serve data communities should invest in community managers, not only in storage and preservation technology.
- Repository services will need to be customised per discipline, since each community carries its own metadata schemes, formats, and interfaces.
- Long-tail research data can become part of larger 'big science' infrastructures when an archive supplies standardisation and machine-readable formats.
- Archive strategy should become responsive and reflexive, anticipating community evolution rather than waiting for deposits.
- The boundary between archive and community blurs: the archive's service design partly determines what the community is and can become.
Reading between the lines
- If co-evolution is real, the success of a data archive should be measured by community formation and reuse, not just by dataset counts.
- The autoethnographic method means the paper's generalisation to other archives is a hypothesis; a cross-archive comparison would test it.
- One implicit consequence is that data communities are partly infrastructure effects: change the repository's metadata standards and you may change who can participate.
- A testable extension would be to see whether communities served by passive repositories develop weaker shared standards than communities served by actively mediating archives.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a qualitative case study of the DANS Data Station Life Sciences, drawing on autoethnography, two unstructured interviews with the station manager (Cees Hof), informal weekly consultations with DANS staff, and analysis of internal policy documents. It argues that data communities and data archives co-evolve, with the archive acting not as a passive repository but as an active stakeholder that both supports and shapes research communities. The paper situates this claim in a historical account of DANS's development, its designated-community and long-tail-data strategies, and the establishment of the Life Sciences Data Station. It concludes that the observed dynamics are generalizable to other data communities and data archives.
Significance. If accepted as an empirically grounded account, the paper would make a useful contribution to the growing literature on data communities by foregrounding the institutional and organisational context that earlier scholarship (e.g., Borgman, Gregory, Leonelli) has tended to background. Its strengths are transparency: the authors explicitly acknowledge insider bias, describe their data collection procedures in detail, and credit the informant's co-authorship of the case study. The historical documentation of DANS's evolution from EASY to the Data Station model is valuable and well referenced. However, the central co-evolution claim is only partially supported by the reported evidence: the archive-to-community direction is richly documented, while the community-to-archive direction rests on the testimony of a single actor who also helped write the paper. The generalizability conclusion is asserted rather than argued. These limitations are acknowledged by the authors in places, but the conclusions do not adequately temper the claims.
major comments (3)
- [Methodological approach: Data Collected (pp. 6–8); Conclusions (p. 23)] The central claim of bidirectional co-evolution is not fully evidenced. The paper reports two unstructured interviews with the Data Station manager, informal consultations with DANS staff, and internal policy documents, but no interviews, surveys, or observational data from community members. The community-to-archive direction therefore rests on the manager's paraphrase of community needs, and the manager is also acknowledged as having 'actively contributed to the text, in particular to the Case Study Section.' The Conclusions state that 'a data community and a data archive [are] evolving, reciprocally creating the conditions for their existence,' yet the data as presented can support only the archive's own account of its interactions. To make the bidirectional claim persuasive, the authors should either add independent evidence from community members (e.g., interviews, usage data, or documented community requests) or explicitly reframe the paper as an archive-side perspective rather than a demonstration of mutual co-evolution.
- [The role of data communities and mediating network structures in shaping the Data Station (pp. 19–20); Conclusions (p.] The generalization that 'similar dynamics occur for other data communities and data archives' is unsupported by the single-case insider design. The paper itself notes that its informal notes 'do not follow the rigour of participant observations during fieldwork,' yet the conclusion moves from one self-analysed case to a cross-context claim. The case examples of influence—LTER-LIFE and Health-RI—are large research infrastructures rather than the long-tail data communities the station is intended to serve, so they do not provide a comparative or even a representative basis for the generalization. I would recommend softening the conclusion to a hypothesis or research agenda, or adding at least one comparative case from a different institutional setting.
- [Case Study: The current content of the DANS Data Station Life Sciences (p.] The paper's own case material suggests that the archive's most significant relationships are with institutional partners and research infrastructures, not with the long-tail data communities emphasized in the framing. The butterfly dataset example is hypothetical, and the documented collaborations are with LTER-LIFE, Health-RI, and QUANTUM. If the paper wants to claim co-evolution between a 'data community' and a 'data archive,' it should clarify whether the community in question is the long-tail community, the set of institutional intermediaries, or both. As written, the term 'community' shifts between these referents, which weakens the coherence of the central claim.
minor comments (5)
- [Section heading, p. 15] The heading 'DANS Data Station for the Live Sciences' contains a typo; it should be 'Life Sciences.'
- [Figure 3 caption, p. 31] The caption reads 'Part of the USCD Map'; this should be 'UCSD Map' (University of California San Diego), and the text on p. 16 similarly refers to the 'UCSD Map of Science.'
- [Figure 1 caption, p. 29] The caption contains an unresolved placeholder 'reproduction from (XXX et al., 2012)'; this needs to be completed with the actual citation (likely Scharnhorst et al., 2012, which appears in the reference list).
- [Reference list, p. 27] The reference 'Glissen, V. (2014)' does not match the in-text citation 'Gilissen, 2014'; the spelling should be consistent.
- [Reference list, p. 27] The reference 'Latour, B., & WooIgar, S. (1986)' contains a typographical error: 'Woolgar' is misspelled as 'WooIgar' (with a capital I).
Circularity Check
No significant circularity: the co-evolution claim is an interpretive autoethnographic finding, not a derivation from its own inputs.
full rationale
This paper is a qualitative autoethnographic case study, not a formal derivation with equations or fitted parameters, so the circularity patterns of definitional identity or prediction-by-construction do not apply. The central claim that the DANS Data Station Life Sciences and its data communities co-evolve is presented as an interpretation of two unstructured interviews with the station manager, informal consultations, policy documents, and institutional experience. The paper explicitly discloses the insider character of the evidence: the authors state, 'We are aware of the potential bias that may occur since all authors were employed by the institute at some point in time, which now serves as the object of study,' and the acknowledgement notes that the interviewee 'actively contributed to the text, in particular to the Case Study Section.' That transparency identifies a methodological limitation (single-case insider perspective), but not a circular reduction: the paper's general definition of 'data community' is not defined in terms of the archive, and the archive-to-community direction is not assumed by construction but illustrated with concrete examples such as metadata standards, the LTER-LIFE collaboration, and Health-RI. The self-citations (e.g., Borgman et al. 2019; Gregory et al. 2020; Scharnhorst et al. 2012) supply conceptual framing and prior context; the empirical conclusions rest on the reported case materials rather than on the authority of those citations. No step in the paper makes the conclusion equivalent to its inputs by definition. The generalizability claim ('similar dynamics occur for other data communities and data archives') is an external-validity concern, not a circularity concern, and the under-determination of the community-to-archive direction is an internal-evidence weakness, not a logical reduction. Accordingly, the score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Autoethnography is an appropriate method to derive general insights about institutional dynamics.
- domain assumption The concept of 'data community' is sufficiently coherent to study even though the paper concedes there is no univocal definition.
- domain assumption The views of a single data station manager, supplemented by author observations, are representative of the broader life sciences data community.
- domain assumption Findings from one archive can be transferred to other archives.
Cite this review
Pith. "Pith review of Fostering Data Communities -- perspective from a Data Archive Service Provider." pith.science (2026). https://pith.science/paper/WXHPB6D2
@misc{pith2026250202321,
author = {Pith},
title = {Pith review of: Fostering Data Communities -- perspective from a Data Archive Service Provider},
year = {2026},
howpublished = {\url{https://pith.science/paper/WXHPB6D2}},
note = {Machine review of arXiv:2502.02321}
}
read the original abstract
This paper aims to bridge between the current scientific discourse about the dynamics of data communities in research infrastructures and practical experiences at a data archive which provides services for such data communities. We describe and analyse policies and practices within DANS-KNAW, the Dutch national centre of expertise and repository for research data concerning the interaction with communities in general. We take the case of the emerging DANS Data Station Life Sciences to study how a data archive navigates between observation of data research needs and anticipation of research data archival solutions. This paper offers a unique view of the complex dynamics between data communities (including lay experts) and data service providers. It adds nuances to understanding the emergence of a data community and the role of data service providers, both supporting and shaping, in this process.
Figures
Forward citations
Cited by 1 Pith paper
-
Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
The paper presents the design of ReSearch_SSH, a GraphRAG-based LLM adaptation for multilingual SSH literature discovery within the ISIDORE platform, currently in a preparatory phase with no experimental results yet.
Reference graph
Works this paper leans on
-
[1]
De Boelelaan 1105; 1081HV; HG-0E
Francesca Morselli (corresponding author), VU Vrije University Amsterdam email: f.morselli@vu.nl Adress: Vrije Universiteit Amsterdam. De Boelelaan 1105; 1081HV; HG-0E
-
[2]
Jetze Touber, DANS-KNAW
-
[3]
Long-term archiving of digital research data and the role of communities
Andrea Scharnhorst, DANS-KNAW 4. Fostering Data Communities - perspective from a Data Archive Service Provider Abstract This paper aims to bridge between the current scientific discourse about the dynamics of data communities in research infrastructures and practical experiences at a data archive which provides services for such data communities. We descr...
work page 2023
-
[4]
How do data communities and data archival services co-evolve? 5 More specifically, we ask:
-
[5]
how does the interaction between communities with data needs and the provision of a service by an institution take place?
-
[6]
What is the influence of the data communities on the design of such a data archiving service, and
-
[7]
How does the data service create a reference framework for such a community? We approach these questions based on the understanding that digital research archives are much more than mere providers of services required by specific scientific communities. As unfolded in Borgman (Borgman et al., 2019), digital research archives are mediators between communit...
work page 2019
-
[164]
https://doi.org/10.1126/science.134.3473.161 Weisz, D. (1982). Studies of Scientific Disciplines: An Annotated Bibliography. Division of Planning and Policy Analysis, Office of Planning and Resources Management, National Science Foundation. Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W....
Show all 11 references
-
[2018]
Long Tail of Science
and the Workshop for FAIR Data for the “Long Tail of Science” (Lorentz Center, 2021). The DANS Data Station Life Sciences aims to help the long tail of research data to be findable and reusable. The assumption is that by using domain-specific (meta)data standards, the long tai...
-
[2020]
Long tail data
and shows a substantial growth in the number of datasets and their size. The last self-evaluation of DANS lists more than 250k datasets (DANS Self-Evaluation-Report, 2024) compared to the 20+k datasets about 10 years prior. This growth was triggered by further digitisation of ...
2024
-
[2021]
way” to report auto-ethnographic methodology. Chang speaks of reporting “styles
and scientific laboratories (Latour & WooIgar, 1986). In particular, institutional ethnography has studied how people interact with one another in the context of social institutions, such as at school or in the office (Smith, 2005). To our knowledge, there is no proper “way” t...
2022
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.