Pith. sign in

REVIEW 3 major objections 7 minor 24 references

Challenges in designing research infrastructure software in multi-stakeholder contexts

T0 review · 3 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Research infrastructure software built for automated publication faces two structural challenges: stakeholder groups want different things, and each group is internally heterogeneous.

desk verdict Useful survey of HERMES stakeholder requirements, but the headline cross-group differences outrun the evidence because the two surveys used different instruments. read the letter →

arxiv 2506.01492 v1 pith:4MYE64GK submitted 2025-06-02 cs.SE cs.HC

classification cs.SEcs.HC
keywords researchinfrastructuresoftwarerequirementselicitationmulti-stakeholdercontextspublicationHERMESFAIR4RSengineersautomatedworkflows
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that research infrastructure software for automated software publication is hard to design because it sits between stakeholder groups that want different things and because each group is internally varied. Using two surveys, one of research software engineers and one of infrastructure facility staff, it argues that engineers rank compatibility with existing infrastructure highest, while facility staff rank usability and documentation highest, and that the two groups diverge still more on organizational questions such as responsibility structures and quality assurance. The study names two general challenges: multiple stakeholder groups with differing requirements, and internal heterogeneity within each group. If this is right, a system like HERMES cannot be built around a single typical user; its design and validation must handle conflicts across groups and across levels of technical experience.

What carries the argument

The paper's central object is HERMES, a configurable workflow that runs inside continuous integration systems to automatically harvest software metadata, let it be curated and signed off, deposit the software and metadata in a publication repository with a persistent identifier, and optionally feed metadata back into the source repository. HERMES is the concrete case that makes the multi-stakeholder problem visible, and the two surveys are the instruments that turn that problem into evidence. The multi-stakeholder context is the analytic mechanism: it treats users and operators of research infrastructure as occupying reciprocal provider-user roles, so that any design decision for one group is simultaneously a constraint for the other.

What would settle it

A replication study that gives both stakeholder groups the same feature list on an identical ranking scale and finds no consistent between-group differences in priorities, and no meaningful within-group variation across experience levels, would undermine the claim that multiple stakeholder groups and internal heterogeneity are the central design challenges.

Watch

Extended reading notes

Core claim

The paper's central claim is that the HERMES workflow for automated software publication encounters two design challenges that are structural, not accidental: different stakeholder groups are reciprocally linked as providers and users and have different priorities, and each group is heterogeneous along dimensions like technical experience and discipline. The survey evidence shows that research software engineers most value compatibility with existing infrastructure, out-of-the-box usability, metadata standards, and automation of metadata updates, while infrastructure facility staff most value usability for researchers, documentation, open-source licensing, and metadata standards. On organizational aspects, infrastructure staff rate responsibility structures, guidelines, community management, and quality assurance as more important than engineers do. The paper concludes that requirements engineering for such systems should deprioritize features that only one group ranks highly and prioritize features both groups rate as at least important, and that capacity building and interfaces should accommodate very different experience levels.

Load-bearing premise

The load-bearing premise is that the two surveys, which used different question sets and different Likert scales, measure the same priorities, so that ratings from engineers and infrastructure staff can be directly compared.

Editorial extensions

If this is right

  • Requirements engineering for automated software publication should prioritize features that both stakeholder groups rate as at least important and deprioritize features that only one group ranks highly.
  • The HERMES design must accommodate both technical and non-technical users, with online tutorials and introductory courses as the most broadly requested capacity-building formats.
  • Because only half of the surveyed engineers currently publish their software, adoption of publication automation may depend on lowering barriers or on cultural change as much as on technical features.
  • Infrastructure facility staff may be misjudging what their users want, since they rate usability highest while engineers rate system compatibility highest.
  • Organizational aspects such as responsibility structures and quality assurance need explicit design attention because infrastructure staff consistently weight them much more heavily than engineers do.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is to model this as a multi-objective design problem, where a feature is kept only if it does not strongly reduce acceptance in either stakeholder group.
  • The gap between engineers' and operators' priorities suggests that adoption decisions may be made by infrastructure operators, so an operator-facing evaluation of HERMES could predict uptake better than a user-facing one.
  • A testable follow-up would measure whether institutions with a dedicated research software engineering support group have higher software publication rates than institutions without one, which would speak to whether the 'only half publish' result is cultural or structural.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper reports two online surveys of research software engineers (N=83) and infrastructure facility staff (N=39) conducted to elicit requirements for HERMES, an automated software publication workflow. The authors present descriptive statistics on technical, organizational, and social requirements and claim to find significant differences between the two stakeholder groups, identifying two main design challenges: multiple stakeholder groups with differing requirements and internal heterogeneity within each group. The paper concludes with recommendations for user-centered design and explicitly acknowledges several limitations of the survey design.

Significance. If the comparative finding were valid, the paper would be a useful first step in requirements elicitation for research infrastructure software, highlighting tensions between infrastructure providers and end-user developers. The study is also valuable for its detailed description of the HERMES use case and its transparent discussion of limitations. However, because the cross-group comparisons rest on non-comparable survey instruments, the paper's central claim about differing requirements is not currently established; the result holds only as a tentative, descriptive hypothesis.

major comments (3)
  1. [§4.2, §6, Tables 2–3] The headline cross-group differences that motivate the first main challenge are based on non-comparable survey instruments. The RSE questionnaire asked about 'compatibility with infrastructure systems' and 'out-of-the-box usability' on a 'very important' to 'not important' scale, whereas the IF questionnaire asked about 'compatibility with existing infrastructure' and 'high usability for researchers' on a 'significant' to 'unimportant' scale, and the IF items capture a different perspective (what IFs believe researchers need rather than what RSEs themselves prefer). The percentage gaps cited in §6 (83.1% vs 48.7% for compatibility; 47% vs 84.6% for usability) therefore conflate construct, wording, and scale-anchor differences with genuine stakeholder differences. The authors acknowledge this in §6 ('not optimal for a comparative analysis'), but the Abstract and conclusions still assert significant differences. The comparative analysis should be re-presented as descriptive within each group, or a common item set with ranking should be used if the comparison is to be retained.
  2. [Abstract; §4.3; §6] The phrase 'significant differences' is not supported by the statistical analysis. §4.3 explicitly states that the analysis was descriptive (frequencies, percentages, means, standard deviations); no significance tests, confidence intervals, or effect sizes are reported. With an IF sample of N=39, the observed gaps could easily arise from sampling variation. The authors should replace 'significant' with 'descriptive' or provide appropriate inferential statistics.
  3. [§6] The inference that 'IFs misjudge their users' requirements to some extent' goes beyond the data. The IF item 'high usability for researchers' measures IFs' priorities for researchers, not IFs' own priorities as users, so a gap between IF ratings and RSE self-ratings does not establish misjudgment. This statement should be removed or recast as a hypothesis for future work.
minor comments (7)
  1. [§4.1] The IF sample description uses a comma as decimal separator ('43,6%'); use a period for consistency with the rest of the manuscript.
  2. [§5.2 (IF paragraph)] 'With 59.0%, open-source licensing is another critical factor, identifying it as very important or with 30.8% as "important."' is awkwardly phrased; consider rewriting as 'Open-source licensing was rated very important by 59.0% and important by 30.8%.'
  3. [§5.2] 'combined importance score of 81.%' is missing a digit; Table 2 gives 45.8% + 36.1% = 81.9%.
  4. [§5.3] The sentence 'The participants indicated that they understand the researchers' role slightly better, with 34.9% rating it as "clear and defined" as their role as RSE by 65.1% rating the RSE role as "unclear and vague"' is grammatically broken and should be reworded.
  5. [§5.3] 'Participants even suggested that RSEs and researchers should not clearly define deliverables and expectations' appears to contain a typo (likely 'should clearly define'); as written it contradicts the preceding recommendations.
  6. [Table 1] The 'Total' row's '% Cases' value (208.3%) may confuse readers because multiple responses were allowed; add a note explaining the percentage-of-cases convention.
  7. [General] The survey instruments are not provided as an appendix or supplementary material, limiting reproducibility; consider adding them.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the findings are descriptive survey results, not derived from their own assumptions.

full rationale

The paper's derivation chain consists of two surveys analyzed descriptively (frequencies, percentages, means, thematic coding), with no fitted parameters, equations, or formal models whose outputs could equal their inputs by construction. The load-bearing findings—differing RSE/IF priorities, internal RSE heterogeneity, and the 50.6% software-publication rate—are direct observations from newly collected data. The skeptic's concern that the two questionnaires used different item wording and Likert anchors is a genuine measurement-validity limitation and is explicitly conceded in Section 6 ('these scales... are not optimal for a comparative analysis... could have been asked to rank rather than rate them'), but a validity threat is not circularity: the conclusions are not encoded in the survey definitions, and the authors present the cross-group gaps as observed percentages rather than as outputs of a model fitted to those same percentages. Self-citations (the HERMES concept and the Hasselbring et al. categorization) provide context and definitions; even where a cited categorization includes the present authors, it is not invoked as the evidence for the empirical claims. No step reduces to its own inputs, so the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The paper introduces no free parameters or invented entities. Its central claims rest on two domain assumptions about sample representativeness and cross-group comparability of different Likert instruments, both partly acknowledged in the limitations.

assumptions (2)
  • domain assumption Survey respondents are representative of RSE and IF stakeholder populations.
    Generalizability of the descriptive findings relies on this, but the sample is self-selected and small (N=83 and N=39), and the authors note underrepresentation of some disciplines in Section 6.
  • domain assumption Likert-scale ratings are comparable across the two stakeholder groups despite different question sets.
    The authors acknowledge in Section 6 that the surveys did not use the same options and that ranking would have been better; cross-group comparisons assume comparability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Challenges in designing research infrastructure software in multi-stakeholder contexts." pith.science (2026). https://pith.science/paper/4MYE64GK

@misc{pith2026250601492,
  author       = {Pith},
  title        = {Pith review of: Challenges in designing research infrastructure software in multi-stakeholder contexts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4MYE64GK}},
  note         = {Machine review of arXiv:2506.01492}
}
read the original abstract

This study investigates the challenges in designing research infrastructure software for automated software publication in multi-stakeholder environments, focusing specifically on the HERMES system. Through two quantitative surveys of research software engineers (RSEs) and infrastructure facility staff (IFs), it examines technical, organizational, and social requirements across these stakeholder groups. The study reveals significant differences in how RSEs and IFs prioritize various system features. While RSEs highly value compatibility with existing infrastructure, IFs prioritize user-focused aspects like system usability and documentation. The research identifies two main challenges in designing research infrastructure software: (1) the existence of multiple stakeholder groups with differing requirements, and (2) the internal heterogeneity within each stakeholder group across dimensions such as technical experience. The study also highlights that only half of RSE respondents actively practice software publication, pointing to potential cultural or technical barriers. Additionally, the research reveals discrepancies in how stakeholders view organizational aspects, with IFs consistently rating factors like responsibility structures and quality assurance as more important than RSEs do. These findings contribute to a better understanding of the complexities involved in designing research infrastructure software and emphasize the need for systems that can accommodate diverse user groups while maintaining usability across different technical expertise levels.

Figures

Figures reproduced from arXiv: 2506.01492 by the authors.

Figure 1
Figure 1. Dimensions of diversity in the research software engineer (RSE) role (adapted from [7]) [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. HERMES workflow phases Currently, HERMES has successfully been validated in lab conditions [15]. Developing the HER￾MES prototype into research infrastructure software requires its validation and demonstration in a relevant environment (see [9]), which in turn requires careful system design. Relevant environments [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Institutional support for software publication for IFs [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: RSE’s preferences for technical features in automated software publication systems [PITH_FULL_IMAGE:figures/full_fig_p013_4.png]
Figure 5
Figure 5. Figure 5: RSE’s relevance ratings of organizational aspects for automated research software publication [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]
Figure 6
Figure 6. Figure 6: IF’s importance ratings of organizational aspects [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

24 extracted references · 18 canonical work pages

  1. [1]

    Scientific Data9(1), 622 (Oct 2022)

    Barker, M., Chue Hong, N.P., Katz, D.S., Lamprecht, A.L., Martinez-Ortiz, C., Psomopoulos, F., Har- row, J., Castro, L.J., Gruenpeter, M., Martinez, P.A., Honeyman, T.: Introducing the FAIR Principles for research software. Scientific Data9(1), 622 (Oct 2022). https://doi.org/10.1038/s41597-022-0 1710-x

  2. [2]

    Research Data Alliance (2021).https://doi.org/10.15497/RDA00065

    Chue Hong, N.P., Katz, D.S., Barker, M., Lamprecht, A.L., Martinez, C., Psomopoulos, F.E., Harrow, J., Castro, L.J., Gruenpeter, M., Martinez, P.A., al., e.: FAIR principles for research software (FAIR4RS principles). Research Data Alliance (2021).https://doi.org/10.15497/RDA00065

  3. [3]

    Code of Con- duct (Apr 2022).https://doi.org/10.5281/zenodo.6472827

    Deutsche Forschungsgemeinschaft: Guidelines for Safeguarding Good Research Practice. Code of Con- duct (Apr 2022).https://doi.org/10.5281/zenodo.6472827

  4. [4]

    Deutsche Forschungsgemeinschaft: Handling of Research Software in the DFG’s Funding Activities [Umgang mit Forschungssoftware im Förderhandeln der DFG] (Oct 2024).https://doi.org/10.528 1/zenodo.13919790, publisher: Zenodo

  5. [5]

    Software publications with rich metadata: state of the art, automated workflows and HERMES concept

    Druskat, S., Bertuch, O., Juckeland, G., Knodel, O., Schlauch, T.: Software publications with rich metadata: State of the art, automated workflows and HERMES concept. arXiv (Jan 2022).https: //doi.org/10.48550/arXiv.2201.09015

  6. [6]

    https://doi.org/10.1515/abitech-2023-0031 Research software in multi-stakeholder contexts 19

    Druskat, S., Bertuch, O., Struck, A.: Towards Research Software-ready Libraries: Forschungssoftware in Bibliotheken.ABITechnik 43(3),168–178(Aug2023). https://doi.org/10.1515/abitech-2023-0031 Research software in multi-stakeholder contexts 19

  7. [7]

    Softwaretechnik-Trends44(1), 12–15 (2024).https://dl.gi.de/handle/20.500.12116/44043

    Druskat, S., Felderer, M., Haupt, C.: Software engineering, research software and requirements engi- neering. Softwaretechnik-Trends44(1), 12–15 (2024).https://dl.gi.de/handle/20.500.12116/44043

  8. [8]

    In: Mori, H., Asahi, Y

    Dworatzyk, K., Dekorsy, V., Theis, S.: Decoding the Diversity of the German Software Developer Community: Insights from an Exploratory Cluster Analysis. In: Mori, H., Asahi, Y. (eds.) Human Interface and the Management of Information. pp. 275–295. Springer Nature Switzerland, Cham (2024). https://doi.org/10.1007/978-3-031-60125-5_19

Show all 24 references
  1. [9]

    Directorate General for Research and Innovation.: Technology readiness level: guidance principles for renewable energy technologies : final report

    European Commission. Directorate General for Research and Innovation.: Technology readiness level: guidance principles for renewable energy technologies : final report. Publications Office, LU (2017). https://data.europa.eu/doi/10.2777/577767

  2. [10]

    Gruenpeter, M., Katz, D.S., Lamprecht, A.L., Honeyman, T., Garijo, D., Struck, A., Niehues, A., Martinez, P.A., Castro, L.J., Rabemanantsoa, T., Chue Hong, N.P., Martinez-Ortiz, C., Sesink, L., Liffers, M., Fouilloux, A.C., Erdmann, C., Peroni, S., Martinez Lavanchy, P., Todor...

  3. [11]

    Computing in Science & Engineering pp

    Hasselbring, W., Druskat, S., Bernoth, J., Betker, P., Felderer, M., Ferenz, S., Hermann, B., Lamprecht, A.L.,Linxweiler,J.,Prat,A.,Rumpe,B.,Schoening-Stierand,K.,Yang,S.:Multi-DimensionalResearch Software Categorization. Computing in Science & Engineering pp. 1–10 (2025).http...

  4. [12]

    Hettrick, S., Bast, R., Crouch, S., Wyatt, C., Philippe, O., Botzki, A., Carver, J., Cosden, I., D’Andrea, F., Dasgupta, A., Godoy, W., Gonzalez-Beltran, A., Hamster, U., Henwood, S., Holmvall, P., Janosch, S., Lestang, T., May, N., Philips, J., Poonawala-Lohani, N., Richmond,...

  5. [13]

    Computing in Science & Engineering9(3), 90–95 (2007)

    Hunter, J.D.: Matplotlib: A 2D Graphics Environment. Computing in Science & Engineering9(3), 90–95 (2007). https://doi.org/10.1109/MCSE.2007.55

  6. [14]

    International Journal of Digital Curation16(1), 6 (Apr 2021)

    Jay, C., Haines, R., Katz, D.S.: Software Must be Recognised as an Important Output of Scholarly Research. International Journal of Digital Curation16(1), 6 (Apr 2021). https://doi.org/10.2218/ ijdc.v16i1.745

  7. [15]

    Electronic Communications of the EASST 83(Electronic Communications of the EASST, Vol

    Kernchen, S., Meinel, M., Druskat, S., Fritzsche, M., Pape, D., Bertuch, O.: Extending and apply- ing automated HERMES software publication workflows. Electronic Communications of the EASST 83(Electronic Communications of the EASST, Vol. 83 (2025): deRSE24 - Selected Contribut...

  8. [16]

    LimeSurvey GmbH: LimeSurvey: An open source survey tool.https://www.limesurvey.org/

  9. [17]

    https://doi.org/10.5281/zenodo.13221383

    Meinel, M., Druskat, S., Kelling, J., Bertuch, O., Knodel, O., Pape, D., Kernchen, S.: hermes (Aug 2024). https://doi.org/10.5281/zenodo.13221383

  10. [18]

    Physics World32(3), 40 (Mar 2019)

    Skuse, B.: The third pillar. Physics World32(3), 40 (Mar 2019). https://doi.org/10.1088/2058-7 058/32/3/33

  11. [19]

    PeerJ Computer Science2(e86) (2016)

    Smith, A.M., Katz, D.S., Niemeyer, K.E., FORCE11 Software Citation Working Group: Software cita- tion principles. PeerJ Computer Science2(e86) (2016). https://doi.org/10.7717/peerj-cs.86

  12. [20]

    https://doi.org/10.5281/ZENODO.14464227

    The Matplotlib Development Team: Matplotlib: Visualization with Python (Version v3.10.0) (Dec 2024). https://doi.org/10.5281/ZENODO.14464227

  13. [21]

    The pandas development team: Pandas (Version v2.2.3) (Sep 2024).https://doi.org/10.5281/ZENO DO.13819579

  14. [22]

    National Science Foundation: Proposal & Award Policies & Procedures Guide (PAPPG) (NSF 24-1) (2024)

    U.S. National Science Foundation: Proposal & Award Policies & Procedures Guide (PAPPG) (NSF 24-1) (2024). https://new.nsf.gov/policies/pappg/24-1

  15. [23]

    (ed.): Guide to the Software Engineering Body of Knowledge (SWEBOK Guide), Version 4.0

    Washizaki, H. (ed.): Guide to the Software Engineering Body of Knowledge (SWEBOK Guide), Version 4.0. IEEE Computer Society (2024).https://www.swebok.org

  16. [24]

    Scientific Data 3, 160018 (Mar 2016).https://doi.org/10.1038/sdata.2016.18

    Wilkinson, M.D., et al.: The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018 (Mar 2016).https://doi.org/10.1038/sdata.2016.18

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.