{"id":"53ed9f23-aad5-43a0-a467-11784644297b","arxiv_id":"2506.01492","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Surveys of research software engineers and infrastructure staff show divergent priorities for automated software publication systems, with RSEs valuing infrastructure compatibility and IFs valuing usability and documentation.","lead":"This paper reports two surveys of research software engineers and infrastructure staff about what they want from automated software publication tools. It finds the two groups prioritize different features and that neither group is uniform.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline cross-group comparisons rely on different survey instruments, so the central 'differing requirements' finding may be an artifact of question wording and scale anchors rather than real stakeholder differences.","rationale":"The reader identified the same load-bearing weakness: the two surveys used different question sets and Likert scales, undermining direct cross-group comparisons. This is the single most consequential flaw because the paper's headline contribution—that RSEs and IFs have differing priorities—depends entirely on comparing percentages across instruments. The authors' own limitation statement in Section 6 confirms the issue. A conditional verdict is appropriate: the descriptive account of each group separately is useful, and the internal-heterogeneity finding within the RSE sample is largely independent of the cross-survey comparability problem. But the strong comparative claim, and the abstract's 'significant differences,' should be reworded or re-analyzed before acceptance. I would keep the reader's CONDITIONAL verdict, with the condition being the reanalysis or explicit downgrading of the cross-group comparison.","tokens_in":15048,"tokens_out":1623,"duration_ms":19267,"concrete_test":"Re-analyze the raw survey data using only items that are semantically matched across the two questionnaires (e.g., compatibility with existing infrastructure, usability, documentation, compliance with metadata standards), and convert each group's ratings to within-group ranks or standardized scores before comparing. If the between-group gaps in shared items shrink to non-significant or reverse direction, the claim that RSEs and IFs fundamentally differ in their requirements is not supported. Additionally, compute a simple chi-square or Fisher exact test on the top-box ('very important' vs. all other) counts for each matched item; if no differences reach significance at N=83 and N=39, the abstract's wording of 'significant differences' must be downgraded to descriptive observation.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim—that multiple stakeholder groups have differing requirements—rests on direct percentage comparisons between the RSE and IF surveys (e.g., Section 6: 'compatibility with existing infrastructure' 83.1% vs 48.7%; 'usability' 47% vs 84.6%; documentation 36.1% vs 64.1%). However, the two questionnaires used different item sets and different Likert anchors. Section 4.2 states the RSE survey asked about 'compatibility with infrastructure systems' and 'out-of-the-box usability' on a four-point scale from 'very important' to 'not important,' while the IF survey asked about 'high usability for researchers,' 'compatibility with existing infrastructure,' etc., on a scale from 'significant' to 'unimportant.' These are not the same constructs measured with the same instrument, so the observed percentage gaps conflate genuine stakeholder differences with item wording, response anchor, and context effects. The authors themselves acknowledge in Section 6 that the Likert scales were 'not optimal for a comparative analysis' and that a common item set with ranking 'would have made differences in attitudes clearer.' No inferential statistics are provided to test whether the gaps exceed sampling variability. The second claimed challenge, internal heterogeneity within RSEs, is better supported because it is based on descriptive diversity within one sample, but the first challenge—the load-bearing pillar of the paper's contribution—is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports two online surveys of research software engineers (N=83) and infrastructure facility staff (N=39) conducted to elicit requirements for HERMES, an automated software publication workflow. The authors present descriptive statistics on technical, organizational, and social requirements and claim to find significant differences between the two stakeholder groups, identifying two main design challenges: multiple stakeholder groups with differing requirements and internal heterogeneity within each group. The paper concludes with recommendations for user-centered design and explicitly acknowledges several limitations of the survey design.","tokens_in":15395,"tokens_out":7023,"duration_ms":63965,"significance":"If the comparative finding were valid, the paper would be a useful first step in requirements elicitation for research infrastructure software, highlighting tensions between infrastructure providers and end-user developers. The study is also valuable for its detailed description of the HERMES use case and its transparent discussion of limitations. However, because the cross-group comparisons rest on non-comparable survey instruments, the paper's central claim about differing requirements is not currently established; the result holds only as a tentative, descriptive hypothesis.","major_comments":[{"comment":"The headline cross-group differences that motivate the first main challenge are based on non-comparable survey instruments. The RSE questionnaire asked about 'compatibility with infrastructure systems' and 'out-of-the-box usability' on a 'very important' to 'not important' scale, whereas the IF questionnaire asked about 'compatibility with existing infrastructure' and 'high usability for researchers' on a 'significant' to 'unimportant' scale, and the IF items capture a different perspective (what IFs believe researchers need rather than what RSEs themselves prefer). The percentage gaps cited in §6 (83.1% vs 48.7% for compatibility; 47% vs 84.6% for usability) therefore conflate construct, wording, and scale-anchor differences with genuine stakeholder differences. The authors acknowledge this in §6 ('not optimal for a comparative analysis'), but the Abstract and conclusions still assert significant differences. The comparative analysis should be re-presented as descriptive within each group, or a common item set with ranking should be used if the comparison is to be retained.","section":"§4.2, §6, Tables 2–3"},{"comment":"The phrase 'significant differences' is not supported by the statistical analysis. §4.3 explicitly states that the analysis was descriptive (frequencies, percentages, means, standard deviations); no significance tests, confidence intervals, or effect sizes are reported. With an IF sample of N=39, the observed gaps could easily arise from sampling variation. The authors should replace 'significant' with 'descriptive' or provide appropriate inferential statistics.","section":"Abstract; §4.3; §6"},{"comment":"The inference that 'IFs misjudge their users' requirements to some extent' goes beyond the data. The IF item 'high usability for researchers' measures IFs' priorities for researchers, not IFs' own priorities as users, so a gap between IF ratings and RSE self-ratings does not establish misjudgment. This statement should be removed or recast as a hypothesis for future work.","section":"§6"}],"minor_comments":[{"comment":"The IF sample description uses a comma as decimal separator ('43,6%'); use a period for consistency with the rest of the manuscript.","section":"§4.1"},{"comment":"'With 59.0%, open-source licensing is another critical factor, identifying it as very important or with 30.8% as \"important.\"' is awkwardly phrased; consider rewriting as 'Open-source licensing was rated very important by 59.0% and important by 30.8%.'","section":"§5.2 (IF paragraph)"},{"comment":"'combined importance score of 81.%' is missing a digit; Table 2 gives 45.8% + 36.1% = 81.9%.","section":"§5.2"},{"comment":"The sentence 'The participants indicated that they understand the researchers' role slightly better, with 34.9% rating it as \"clear and defined\" as their role as RSE by 65.1% rating the RSE role as \"unclear and vague\"' is grammatically broken and should be reworded.","section":"§5.3"},{"comment":"'Participants even suggested that RSEs and researchers should not clearly define deliverables and expectations' appears to contain a typo (likely 'should clearly define'); as written it contradicts the preceding recommendations.","section":"§5.3"},{"comment":"The 'Total' row's '% Cases' value (208.3%) may confuse readers because multiple responses were allowed; add a note explaining the percentage-of-cases convention.","section":"Table 1"},{"comment":"The survey instruments are not provided as an appendix or supplementary material, limiting reproducibility; consider adding them.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest descriptive account with useful context, but the comparative claims are the core of the contribution and they are not supported by the data collection design. In my view the paper could be publishable as a case study/experience report if the authors reframe the findings as within-group descriptive results and present the cross-group differences as hypotheses. The authors' own Section 6 limitations are commendably candid, which makes the revision feasible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you are designing software publication infrastructure. The paper reports two surveys (83 RSEs, 39 infrastructure staff) about requirements for HERMES, an automated software publication workflow. The data are original, and the paper is transparent about what it did and did not find. The strongest parts are the descriptive findings: RSEs prioritize compatibility with existing infrastructure (83% \"very important\"), IFs prioritize usability for researchers (85%) and documentation (64% users, 59% admins). The observation that only half of RSE respondents practice software publication is a real signal.\n\nThe main problem is the abstract's claim of \"significant differences\" between stakeholder groups. No inferential statistics are reported, and the two surveys used different question items and Likert anchors (Section 4.2), so the percentage gaps in Section 6 are not clean comparisons. The authors themselves concede in Section 6 that ranking instead of rating and a common item set would have made differences clearer. That does not sink the paper, because the qualitative pattern is plausible and the organizational-aspect comparisons (IFs consistently rating responsibility, QA, etc. higher) are less susceptible to wording differences. But the specific numbers should be treated as descriptive, and the overclaim in the abstract should be fixed.\n\nThe stress-test note worried that the central claim might be an artifact of question wording. I think that is too harsh. The second claimed challenge, internal heterogeneity of RSEs, rests on within-sample diversity and is well supported. And the cross-group direction is consistent with role differences. Still, the burden is on the authors to share the instruments and data, or temper the language.\n\nNot a groundbreaking paper, but a solid requirements-elicitation case study. The authors know the limitations and say so. If you want to know what actual RSEs and library staff want from a publication tool, this is useful. I would send it to peer review, with the expectation of minor to moderate revision: soften \"significant,\" add a methods caveat, and offer supplementary materials.","headline":"Useful survey of HERMES stakeholder requirements, but the headline cross-group differences outrun the evidence because the two surveys used different instruments.","tokens_in":15757,"tokens_out":2533,"would_cite":true,"duration_ms":25404,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Research infrastructure software built for automated publication faces two structural challenges: stakeholder groups want different things, and each group is internally heterogeneous.","keywords":["research infrastructure software","requirements elicitation","multi-stakeholder contexts","software publication","HERMES","FAIR4RS","research software engineers","automated publication workflows"],"falsifier":"A replication study that gives both stakeholder groups the same feature list on an identical ranking scale and finds no consistent between-group differences in priorities, and no meaningful within-group variation across experience levels, would undermine the claim that multiple stakeholder groups and internal heterogeneity are the central design challenges.","tokens_in":14848,"feed_emoji":"⚙️","tokens_out":5673,"duration_ms":57997,"temperature":0.7,"pith_summary":"The paper sets out to show that research infrastructure software for automated software publication is hard to design because it sits between stakeholder groups that want different things and because each group is internally varied. Using two surveys, one of research software engineers and one of infrastructure facility staff, it argues that engineers rank compatibility with existing infrastructure highest, while facility staff rank usability and documentation highest, and that the two groups diverge still more on organizational questions such as responsibility structures and quality assurance. The study names two general challenges: multiple stakeholder groups with differing requirements, and internal heterogeneity within each group. If this is right, a system like HERMES cannot be built around a single typical user; its design and validation must handle conflicts across groups and across levels of technical experience.","feed_headline":"Survey: engineers and infrastructure staff want different tools","feed_subtitle":"A study of the HERMES publication workflow shows the real design problem is divergent priorities across user groups.","key_machinery":"The paper's central object is HERMES, a configurable workflow that runs inside continuous integration systems to automatically harvest software metadata, let it be curated and signed off, deposit the software and metadata in a publication repository with a persistent identifier, and optionally feed metadata back into the source repository. HERMES is the concrete case that makes the multi-stakeholder problem visible, and the two surveys are the instruments that turn that problem into evidence. The multi-stakeholder context is the analytic mechanism: it treats users and operators of research infrastructure as occupying reciprocal provider-user roles, so that any design decision for one group is simultaneously a constraint for the other.","core_discovery":"The paper's central claim is that the HERMES workflow for automated software publication encounters two design challenges that are structural, not accidental: different stakeholder groups are reciprocally linked as providers and users and have different priorities, and each group is heterogeneous along dimensions like technical experience and discipline. The survey evidence shows that research software engineers most value compatibility with existing infrastructure, out-of-the-box usability, metadata standards, and automation of metadata updates, while infrastructure facility staff most value usability for researchers, documentation, open-source licensing, and metadata standards. On organizational aspects, infrastructure staff rate responsibility structures, guidelines, community management, and quality assurance as more important than engineers do. The paper concludes that requirements engineering for such systems should deprioritize features that only one group ranks highly and prioritize features both groups rate as at least important, and that capacity building and interfaces should accommodate very different experience levels.","pith_inferences":["A natural extension the paper leaves implicit is to model this as a multi-objective design problem, where a feature is kept only if it does not strongly reduce acceptance in either stakeholder group.","The gap between engineers' and operators' priorities suggests that adoption decisions may be made by infrastructure operators, so an operator-facing evaluation of HERMES could predict uptake better than a user-facing one.","A testable follow-up would measure whether institutions with a dedicated research software engineering support group have higher software publication rates than institutions without one, which would speak to whether the 'only half publish' result is cultural or structural."],"forward_implications":["Requirements engineering for automated software publication should prioritize features that both stakeholder groups rate as at least important and deprioritize features that only one group ranks highly.","The HERMES design must accommodate both technical and non-technical users, with online tutorials and introductory courses as the most broadly requested capacity-building formats.","Because only half of the surveyed engineers currently publish their software, adoption of publication automation may depend on lowering barriers or on cultural change as much as on technical features.","Infrastructure facility staff may be misjudging what their users want, since they rate usability highest while engineers rate system compatibility highest.","Organizational aspects such as responsibility structures and quality assurance need explicit design attention because infrastructure staff consistently weight them much more heavily than engineers do."],"supporting_citations":[{"why":"Defines the HERMES concept and workflow that the whole study is about.","marker":"[5]"},{"why":"Supplies the definition of software publication that the paper uses to frame the surveys.","marker":"[6]"},{"why":"Introduces the research infrastructure software category and the multi-stakeholder view the paper adopts.","marker":"[11]"},{"why":"Provides the FAIR4RS principles that motivate why software publication and metadata standards matter.","marker":"[2]"},{"why":"Reports the lab validation of HERMES that the paper builds on to argue for real-environment design.","marker":"[15]"},{"why":"Identifies the HERMES Python package in which the workflow is prototyped.","marker":"[17]"}],"fun_headline_variants":["Engineers want compatibility, staff want usability","Research infra design: stakeholder priorities split","Designing research software: who's the real user?","Multi-stakeholder tools: why one-size-fits-all is hard","For research infra, user groups value different features"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the two surveys, which used different question sets and different Likert scales, measure the same priorities, so that ratings from engineers and infrastructure staff can be directly compared.","fun_headline_variants_meta":{"raw":{"variants":["Engineers want compatibility, staff want usability","Research infra design: stakeholder priorities split","Designing research software: who's the real user?","Multi-stakeholder tools: why one-size-fits-all is hard","For research infra, user groups value different features"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3185,"prompt_tokens":927,"completion_tokens":2258,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":2183}},"tokens_in":543,"tokens_out":2258,"duration_ms":17502,"temperature":1.0,"reasoning_tokens":2183,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:38:24.605272+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication study that gives both stakeholder groups the same feature list on an identical ranking scale and finds no consistent between-group differences in priorities, and no meaningful within-group variation across experience levels, would undermine the claim that multiple stakeholder groups and internal heterogeneity are the central design challenges.","supporting_citations":[{"cited_title":"Software publications with rich metadata: state of the art, automated workflows and HERMES concept","cited_arxiv_id":"2201.09015","evidence_quote":"Defines the HERMES concept and workflow that the whole study is about."},{"cited_title":"https://doi.org/10.1515/abitech-2023-0031 Research software in multi-stakeholder contexts 19","cited_arxiv_id":null,"evidence_quote":"Supplies the definition of software publication that the paper uses to frame the surveys."},{"cited_title":"Computing in Science & Engineering pp","cited_arxiv_id":null,"evidence_quote":"Introduces the research infrastructure software category and the multi-stakeholder view the paper adopts."},{"cited_title":"Research Data Alliance (2021).https://doi.org/10.15497/RDA00065","cited_arxiv_id":null,"evidence_quote":"Provides the FAIR4RS principles that motivate why software publication and metadata standards matter."},{"cited_title":"Electronic Communications of the EASST 83(Electronic Communications of the EASST, Vol","cited_arxiv_id":null,"evidence_quote":"Reports the lab validation of HERMES that the paper builds on to argue for real-environment design."},{"cited_title":"https://doi.org/10.5281/zenodo.13221383","cited_arxiv_id":null,"evidence_quote":"Identifies the HERMES Python package in which the workflow is prototyped."}],"review_version":1}