REVIEW 3 major objections 7 minor 24 references
Challenges in designing research infrastructure software in multi-stakeholder contexts
T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Research infrastructure software built for automated publication faces two structural challenges: stakeholder groups want different things, and each group is internally heterogeneous.
desk verdict Useful survey of HERMES stakeholder requirements, but the headline cross-group differences outrun the evidence because the two surveys used different instruments. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's central object is HERMES, a configurable workflow that runs inside continuous integration systems to automatically harvest software metadata, let it be curated and signed off, deposit the software and metadata in a publication repository with a persistent identifier, and optionally feed metadata back into the source repository. HERMES is the concrete case that makes the multi-stakeholder problem visible, and the two surveys are the instruments that turn that problem into evidence. The multi-stakeholder context is the analytic mechanism: it treats users and operators of research infrastructure as occupying reciprocal provider-user roles, so that any design decision for one group is simultaneously a constraint for the other.
What would settle it
A replication study that gives both stakeholder groups the same feature list on an identical ranking scale and finds no consistent between-group differences in priorities, and no meaningful within-group variation across experience levels, would undermine the claim that multiple stakeholder groups and internal heterogeneity are the central design challenges.
Extended reading notes
Core claim
The paper's central claim is that the HERMES workflow for automated software publication encounters two design challenges that are structural, not accidental: different stakeholder groups are reciprocally linked as providers and users and have different priorities, and each group is heterogeneous along dimensions like technical experience and discipline. The survey evidence shows that research software engineers most value compatibility with existing infrastructure, out-of-the-box usability, metadata standards, and automation of metadata updates, while infrastructure facility staff most value usability for researchers, documentation, open-source licensing, and metadata standards. On organizational aspects, infrastructure staff rate responsibility structures, guidelines, community management, and quality assurance as more important than engineers do. The paper concludes that requirements engineering for such systems should deprioritize features that only one group ranks highly and prioritize features both groups rate as at least important, and that capacity building and interfaces should accommodate very different experience levels.
Load-bearing premise
The load-bearing premise is that the two surveys, which used different question sets and different Likert scales, measure the same priorities, so that ratings from engineers and infrastructure staff can be directly compared.
Editorial extensions
If this is right
- Requirements engineering for automated software publication should prioritize features that both stakeholder groups rate as at least important and deprioritize features that only one group ranks highly.
- The HERMES design must accommodate both technical and non-technical users, with online tutorials and introductory courses as the most broadly requested capacity-building formats.
- Because only half of the surveyed engineers currently publish their software, adoption of publication automation may depend on lowering barriers or on cultural change as much as on technical features.
- Infrastructure facility staff may be misjudging what their users want, since they rate usability highest while engineers rate system compatibility highest.
- Organizational aspects such as responsibility structures and quality assurance need explicit design attention because infrastructure staff consistently weight them much more heavily than engineers do.
Reading between the lines
- A natural extension the paper leaves implicit is to model this as a multi-objective design problem, where a feature is kept only if it does not strongly reduce acceptance in either stakeholder group.
- The gap between engineers' and operators' priorities suggests that adoption decisions may be made by infrastructure operators, so an operator-facing evaluation of HERMES could predict uptake better than a user-facing one.
- A testable follow-up would measure whether institutions with a dedicated research software engineering support group have higher software publication rates than institutions without one, which would speak to whether the 'only half publish' result is cultural or structural.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports two online surveys of research software engineers (N=83) and infrastructure facility staff (N=39) conducted to elicit requirements for HERMES, an automated software publication workflow. The authors present descriptive statistics on technical, organizational, and social requirements and claim to find significant differences between the two stakeholder groups, identifying two main design challenges: multiple stakeholder groups with differing requirements and internal heterogeneity within each group. The paper concludes with recommendations for user-centered design and explicitly acknowledges several limitations of the survey design.
Significance. If the comparative finding were valid, the paper would be a useful first step in requirements elicitation for research infrastructure software, highlighting tensions between infrastructure providers and end-user developers. The study is also valuable for its detailed description of the HERMES use case and its transparent discussion of limitations. However, because the cross-group comparisons rest on non-comparable survey instruments, the paper's central claim about differing requirements is not currently established; the result holds only as a tentative, descriptive hypothesis.
major comments (3)
- [§4.2, §6, Tables 2–3] The headline cross-group differences that motivate the first main challenge are based on non-comparable survey instruments. The RSE questionnaire asked about 'compatibility with infrastructure systems' and 'out-of-the-box usability' on a 'very important' to 'not important' scale, whereas the IF questionnaire asked about 'compatibility with existing infrastructure' and 'high usability for researchers' on a 'significant' to 'unimportant' scale, and the IF items capture a different perspective (what IFs believe researchers need rather than what RSEs themselves prefer). The percentage gaps cited in §6 (83.1% vs 48.7% for compatibility; 47% vs 84.6% for usability) therefore conflate construct, wording, and scale-anchor differences with genuine stakeholder differences. The authors acknowledge this in §6 ('not optimal for a comparative analysis'), but the Abstract and conclusions still assert significant differences. The comparative analysis should be re-presented as descriptive within each group, or a common item set with ranking should be used if the comparison is to be retained.
- [Abstract; §4.3; §6] The phrase 'significant differences' is not supported by the statistical analysis. §4.3 explicitly states that the analysis was descriptive (frequencies, percentages, means, standard deviations); no significance tests, confidence intervals, or effect sizes are reported. With an IF sample of N=39, the observed gaps could easily arise from sampling variation. The authors should replace 'significant' with 'descriptive' or provide appropriate inferential statistics.
- [§6] The inference that 'IFs misjudge their users' requirements to some extent' goes beyond the data. The IF item 'high usability for researchers' measures IFs' priorities for researchers, not IFs' own priorities as users, so a gap between IF ratings and RSE self-ratings does not establish misjudgment. This statement should be removed or recast as a hypothesis for future work.
minor comments (7)
- [§4.1] The IF sample description uses a comma as decimal separator ('43,6%'); use a period for consistency with the rest of the manuscript.
- [§5.2 (IF paragraph)] 'With 59.0%, open-source licensing is another critical factor, identifying it as very important or with 30.8% as "important."' is awkwardly phrased; consider rewriting as 'Open-source licensing was rated very important by 59.0% and important by 30.8%.'
- [§5.2] 'combined importance score of 81.%' is missing a digit; Table 2 gives 45.8% + 36.1% = 81.9%.
- [§5.3] The sentence 'The participants indicated that they understand the researchers' role slightly better, with 34.9% rating it as "clear and defined" as their role as RSE by 65.1% rating the RSE role as "unclear and vague"' is grammatically broken and should be reworded.
- [§5.3] 'Participants even suggested that RSEs and researchers should not clearly define deliverables and expectations' appears to contain a typo (likely 'should clearly define'); as written it contradicts the preceding recommendations.
- [Table 1] The 'Total' row's '% Cases' value (208.3%) may confuse readers because multiple responses were allowed; add a note explaining the percentage-of-cases convention.
- [General] The survey instruments are not provided as an appendix or supplementary material, limiting reproducibility; consider adding them.
Circularity Check
No circularity: the findings are descriptive survey results, not derived from their own assumptions.
full rationale
The paper's derivation chain consists of two surveys analyzed descriptively (frequencies, percentages, means, thematic coding), with no fitted parameters, equations, or formal models whose outputs could equal their inputs by construction. The load-bearing findings—differing RSE/IF priorities, internal RSE heterogeneity, and the 50.6% software-publication rate—are direct observations from newly collected data. The skeptic's concern that the two questionnaires used different item wording and Likert anchors is a genuine measurement-validity limitation and is explicitly conceded in Section 6 ('these scales... are not optimal for a comparative analysis... could have been asked to rank rather than rate them'), but a validity threat is not circularity: the conclusions are not encoded in the survey definitions, and the authors present the cross-group gaps as observed percentages rather than as outputs of a model fitted to those same percentages. Self-citations (the HERMES concept and the Hasselbring et al. categorization) provide context and definitions; even where a cited categorization includes the present authors, it is not invoked as the evidence for the empirical claims. No step reduces to its own inputs, so the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Survey respondents are representative of RSE and IF stakeholder populations.
- domain assumption Likert-scale ratings are comparable across the two stakeholder groups despite different question sets.
Cite this review
Pith. "Pith review of Challenges in designing research infrastructure software in multi-stakeholder contexts." pith.science (2026). https://pith.science/paper/4MYE64GK
@misc{pith2026250601492,
author = {Pith},
title = {Pith review of: Challenges in designing research infrastructure software in multi-stakeholder contexts},
year = {2026},
howpublished = {\url{https://pith.science/paper/4MYE64GK}},
note = {Machine review of arXiv:2506.01492}
}
read the original abstract
This study investigates the challenges in designing research infrastructure software for automated software publication in multi-stakeholder environments, focusing specifically on the HERMES system. Through two quantitative surveys of research software engineers (RSEs) and infrastructure facility staff (IFs), it examines technical, organizational, and social requirements across these stakeholder groups. The study reveals significant differences in how RSEs and IFs prioritize various system features. While RSEs highly value compatibility with existing infrastructure, IFs prioritize user-focused aspects like system usability and documentation. The research identifies two main challenges in designing research infrastructure software: (1) the existence of multiple stakeholder groups with differing requirements, and (2) the internal heterogeneity within each stakeholder group across dimensions such as technical experience. The study also highlights that only half of RSE respondents actively practice software publication, pointing to potential cultural or technical barriers. Additionally, the research reveals discrepancies in how stakeholders view organizational aspects, with IFs consistently rating factors like responsibility structures and quality assurance as more important than RSEs do. These findings contribute to a better understanding of the complexities involved in designing research infrastructure software and emphasize the need for systems that can accommodate diverse user groups while maintaining usability across different technical expertise levels.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Scientific Data9(1), 622 (Oct 2022)
Barker, M., Chue Hong, N.P., Katz, D.S., Lamprecht, A.L., Martinez-Ortiz, C., Psomopoulos, F., Har- row, J., Castro, L.J., Gruenpeter, M., Martinez, P.A., Honeyman, T.: Introducing the FAIR Principles for research software. Scientific Data9(1), 622 (Oct 2022). https://doi.org/10.1038/s41597-022-0 1710-x
-
[2]
Research Data Alliance (2021).https://doi.org/10.15497/RDA00065
Chue Hong, N.P., Katz, D.S., Barker, M., Lamprecht, A.L., Martinez, C., Psomopoulos, F.E., Harrow, J., Castro, L.J., Gruenpeter, M., Martinez, P.A., al., e.: FAIR principles for research software (FAIR4RS principles). Research Data Alliance (2021).https://doi.org/10.15497/RDA00065
-
[3]
Code of Con- duct (Apr 2022).https://doi.org/10.5281/zenodo.6472827
Deutsche Forschungsgemeinschaft: Guidelines for Safeguarding Good Research Practice. Code of Con- duct (Apr 2022).https://doi.org/10.5281/zenodo.6472827
-
[4]
Deutsche Forschungsgemeinschaft: Handling of Research Software in the DFG’s Funding Activities [Umgang mit Forschungssoftware im Förderhandeln der DFG] (Oct 2024).https://doi.org/10.528 1/zenodo.13919790, publisher: Zenodo
work page 2024
-
[5]
Software publications with rich metadata: state of the art, automated workflows and HERMES concept
Druskat, S., Bertuch, O., Juckeland, G., Knodel, O., Schlauch, T.: Software publications with rich metadata: State of the art, automated workflows and HERMES concept. arXiv (Jan 2022).https: //doi.org/10.48550/arXiv.2201.09015
work page Pith review arXiv doi:10.48550/arxiv.2201.09015 2022
-
[6]
https://doi.org/10.1515/abitech-2023-0031 Research software in multi-stakeholder contexts 19
Druskat, S., Bertuch, O., Struck, A.: Towards Research Software-ready Libraries: Forschungssoftware in Bibliotheken.ABITechnik 43(3),168–178(Aug2023). https://doi.org/10.1515/abitech-2023-0031 Research software in multi-stakeholder contexts 19
-
[7]
Softwaretechnik-Trends44(1), 12–15 (2024).https://dl.gi.de/handle/20.500.12116/44043
Druskat, S., Felderer, M., Haupt, C.: Software engineering, research software and requirements engi- neering. Softwaretechnik-Trends44(1), 12–15 (2024).https://dl.gi.de/handle/20.500.12116/44043
work page 2024
-
[8]
Dworatzyk, K., Dekorsy, V., Theis, S.: Decoding the Diversity of the German Software Developer Community: Insights from an Exploratory Cluster Analysis. In: Mori, H., Asahi, Y. (eds.) Human Interface and the Management of Information. pp. 275–295. Springer Nature Switzerland, Cham (2024). https://doi.org/10.1007/978-3-031-60125-5_19
Show all 24 references
-
[9]
Directorate General for Research and Innovation.: Technology readiness level: guidance principles for renewable energy technologies : final report
European Commission. Directorate General for Research and Innovation.: Technology readiness level: guidance principles for renewable energy technologies : final report. Publications Office, LU (2017). https://data.europa.eu/doi/10.2777/577767
2017 doi
-
[10]
Gruenpeter, M., Katz, D.S., Lamprecht, A.L., Honeyman, T., Garijo, D., Struck, A., Niehues, A., Martinez, P.A., Castro, L.J., Rabemanantsoa, T., Chue Hong, N.P., Martinez-Ortiz, C., Sesink, L., Liffers, M., Fouilloux, A.C., Erdmann, C., Peroni, S., Martinez Lavanchy, P., Todor...
2021 doi
-
[11]
Computing in Science & Engineering pp
Hasselbring, W., Druskat, S., Bernoth, J., Betker, P., Felderer, M., Ferenz, S., Hermann, B., Lamprecht, A.L.,Linxweiler,J.,Prat,A.,Rumpe,B.,Schoening-Stierand,K.,Yang,S.:Multi-DimensionalResearch Software Categorization. Computing in Science & Engineering pp. 1–10 (2025).http...
2025
-
[12]
Hettrick, S., Bast, R., Crouch, S., Wyatt, C., Philippe, O., Botzki, A., Carver, J., Cosden, I., D’Andrea, F., Dasgupta, A., Godoy, W., Gonzalez-Beltran, A., Hamster, U., Henwood, S., Holmvall, P., Janosch, S., Lestang, T., May, N., Philips, J., Poonawala-Lohani, N., Richmond,...
2022 doi
-
[13]
Computing in Science & Engineering9(3), 90–95 (2007)
Hunter, J.D.: Matplotlib: A 2D Graphics Environment. Computing in Science & Engineering9(3), 90–95 (2007). https://doi.org/10.1109/MCSE.2007.55
2007 doi
-
[14]
International Journal of Digital Curation16(1), 6 (Apr 2021)
Jay, C., Haines, R., Katz, D.S.: Software Must be Recognised as an Important Output of Scholarly Research. International Journal of Digital Curation16(1), 6 (Apr 2021). https://doi.org/10.2218/ ijdc.v16i1.745
2021
-
[15]
Electronic Communications of the EASST 83(Electronic Communications of the EASST, Vol
Kernchen, S., Meinel, M., Druskat, S., Fritzsche, M., Pape, D., Bertuch, O.: Extending and apply- ing automated HERMES software publication workflows. Electronic Communications of the EASST 83(Electronic Communications of the EASST, Vol. 83 (2025): deRSE24 - Selected Contribut...
2025
-
[16]
LimeSurvey GmbH: LimeSurvey: An open source survey tool.https://www.limesurvey.org/
-
[17]
https://doi.org/10.5281/zenodo.13221383
Meinel, M., Druskat, S., Kelling, J., Bertuch, O., Knodel, O., Pape, D., Kernchen, S.: hermes (Aug 2024). https://doi.org/10.5281/zenodo.13221383
2024 doi
-
[18]
Physics World32(3), 40 (Mar 2019)
Skuse, B.: The third pillar. Physics World32(3), 40 (Mar 2019). https://doi.org/10.1088/2058-7 058/32/3/33
2019 doi
-
[19]
PeerJ Computer Science2(e86) (2016)
Smith, A.M., Katz, D.S., Niemeyer, K.E., FORCE11 Software Citation Working Group: Software cita- tion principles. PeerJ Computer Science2(e86) (2016). https://doi.org/10.7717/peerj-cs.86
2016 doi
-
[20]
https://doi.org/10.5281/ZENODO.14464227
The Matplotlib Development Team: Matplotlib: Visualization with Python (Version v3.10.0) (Dec 2024). https://doi.org/10.5281/ZENODO.14464227
2024 doi
-
[21]
The pandas development team: Pandas (Version v2.2.3) (Sep 2024).https://doi.org/10.5281/ZENO DO.13819579
2024 doi
-
[22]
National Science Foundation: Proposal & Award Policies & Procedures Guide (PAPPG) (NSF 24-1) (2024)
U.S. National Science Foundation: Proposal & Award Policies & Procedures Guide (PAPPG) (NSF 24-1) (2024). https://new.nsf.gov/policies/pappg/24-1
2024
-
[23]
(ed.): Guide to the Software Engineering Body of Knowledge (SWEBOK Guide), Version 4.0
Washizaki, H. (ed.): Guide to the Software Engineering Body of Knowledge (SWEBOK Guide), Version 4.0. IEEE Computer Society (2024).https://www.swebok.org
2024
-
[24]
Scientific Data 3, 160018 (Mar 2016).https://doi.org/10.1038/sdata.2016.18
Wilkinson, M.D., et al.: The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018 (Mar 2016).https://doi.org/10.1038/sdata.2016.18
2016 doi
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.