REVIEW 3 major objections 4 minor 28 references
Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This registered report proposes a controlled experiment to determine whether LLM-generated evidence briefings can match human-made briefings in perceived content fidelity, ease of understanding, and usefulness; the experiment has not yet…
desk verdict A careful registered protocol for comparing LLM vs human evidence briefings, held back by a content-fidelity instrument that cannot detect omitted findings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The experimental factor is generated by a pipeline that combines a commercial LLM with retrieval-augmented generation: for each briefing section, examples are retrieved from a store of 54 human-made evidence briefings and fed to the model as grounded context, with fixed generation parameters and instruction-based prompting. The measurement machinery is a 7-point Likert survey whose items operationalize the three constructs: content fidelity via contradiction, certainty illusion, and fabricated content; ease of understanding via clarity, structure, and conciseness; and usefulness via relevance and actionability. The design machinery is a completely randomized one-factor, two-treatment crossover in which each participant rates two different topics, pair programming and definition of done, so order and carryover effects are controlled.
What would settle it
Collect the two generated briefings and the two source papers, have coders independently count contradictions, certainty mismatches, and unsupported claims, then compare those counts with the participants' perceived-fidelity Likert ratings; a weak or absent correlation would undercut the construct-validity assumption on which the experiment's conclusions depend.
Extended reading notes
Core claim
The central claim is not yet an empirical result; it is that the suitability of LLM-generated evidence briefings can be settled by a crossover experiment in which researchers rate content fidelity against the full source paper and practitioners rate ease of understanding and usefulness, with everyone blinded to which briefing is automatic and which is human-made. The paper asserts that this design, with its three null hypotheses, provides an appropriate test. It also demonstrates that the material conditions for the test exist: an automated pipeline can produce briefings for two previously hand-made cases, and the prompts, tool configuration, outputs, and surveys are archived for reproduction.
Load-bearing premise
The load-bearing premise is that participants' ratings measure the constructs they are asked about: that Likert answers about contradiction, certainty, and fabrication track real fidelity to the source text, and that self-reported understanding and usefulness track how well practitioners can actually use the briefing.
Editorial extensions
If this is right
- If the null hypotheses are not rejected, automation could produce briefings perceived as equivalent to human ones, removing the main scalability barrier to evidence briefings.
- If practitioners rate the LLM briefings lower on understanding or usefulness, the result would show that current generation settings are not yet ready to replace manual production.
- The open-text responses embedded in the questionnaire let the authors explain quantitative differences with participant reasoning, strengthening the interpretation.
- Because all generation parameters, prompts, outputs, and instruments are archived, other teams can replicate or extend the comparison on different secondary studies.
Reading between the lines
- The comparison may partly measure briefing-quality differences rather than generation method alone, since one human-made briefing was previously validated and the other was not; a null result could reflect that asymmetry as much as LLM competence.
- The experiment as designed tests a hybrid pipeline, human-written briefings as retrieval examples plus LLM rewriting, rather than pure LLM generation; varying the retrieval corpus would reveal how much of any measured quality depends on those human examples.
- The Likert self-reports leave room for an objective companion check: counting contradictions, certainty mismatches, or fabricated claims in the generated briefings and comparing those counts with perceived-fidelity ratings.
- A cost-adjusted reading would be that LLM briefings do not need to be better, only comparable, since the motivation is labor savings; the current protocol is well suited to establish equivalence, not superiority.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This registered report describes a protocol for two controlled experiments comparing LLM-generated evidence briefings against human-made ones. In the first experiment, researchers rate content fidelity after reading full research papers and briefings; in the second, practitioners rate ease of understanding and usefulness from the briefings alone. The briefings cover pair programming and definition of done, each with both a human and an LLM-generated version produced by an RAG-based tool. The paper defines null hypotheses for each research question, a counterbalanced crossover design, survey instruments with Likert items, and a threat analysis. No results are reported; the central claim is that the described protocol can test the hypotheses.
Significance. The paper is a protocol paper, so its contribution is a falsifiable experimental design. If executed, it would provide empirical evidence on whether LLM-generated evidence briefings can match human-produced ones in perceived fidelity, understandability, and usefulness, which is relevant to evidence-based software engineering and the adoption of LLM synthesis tools. The authors openly provide their materials, follow LLM reporting guidelines, and sensibly separate researcher and practitioner evaluation roles. The main weaknesses are the incomplete operationalization of content fidelity, the absence of a pre-specified analysis plan and sample-size justification, and the use of an unvalidated human comparator for one of the two briefings; these are fixable in revision.
major comments (3)
- [III-D / III-G] The content fidelity instrument in Section III-D uses three Likert dimensions adapted from Tang et al. (contradiction, certainty illusion, fabricated content). All three detect the presence of incorrect or unsupported content; none detects the absence of important content. Since the paper defines content fidelity as faithfully representing the content of the original research article and the evidence briefing template includes a 'Main Findings' section, an LLM briefing that omits a key finding could still receive maximal fidelity ratings, making the test for H0_1 insensitive to omission, a central failure mode in automatic summarization. The pilot described in Section III-G (two researchers, two practitioners) does not establish construct coverage. I recommend adding a completeness dimension (e.g., items on whether all key findings are present) or a verification task against a predefined list of key findings, and reporting pilot evidence on content coverage.
- [III-H / overall] The protocol does not pre-specify an analysis plan or sample-size justification. Section III-H states that sample size will be 'carefully considered throughout the experiment,' which is insufficient for a confirmatory registered report. The null hypotheses H0_1–H0_3 can only be tested if the statistical procedures, effect size of interest, significance level, power, and handling of the repeated-measures crossover structure are fixed before data collection. Please specify, for example, the planned mixed-effects model (with random effects for participants and period/sequence), the treatment of ordinal Likert data, and a power analysis based on a planned effect size.
- [III-G / III-H] The human-made comparator for the definition-of-done briefing was not validated to the same degree as the pair-programming briefing (Section III-G, External Validity). Since this briefing is one of only two baselines in the crossover, any result concerning equivalence or difference could be confounded by the quality of the human comparator. The acknowledgment in the threats section is appropriate, but the protocol should either include two validated human briefings or add a preliminary validation step for the unvalidated briefing (e.g., review by SLR authors) to ensure the baseline represents the intended condition.
minor comments (4)
- [Table I] The model identifier 'GPT-4-o-mini (gpt-4-0125-preview)' is inconsistent; gpt-4-0125-preview is a GPT-4 Turbo model, not GPT-4o mini. Specify the exact model string used in the API calls to ensure reproducibility.
- [III-G] The pilot study is described as having two researchers and two practitioners with 'no further adjustments deemed necessary.' Please report what was assessed (e.g., item clarity, instrument length) and any quantitative pilot results, to support the claim that the instruments are ready for deployment.
- [III-D] In the Likert scale description, '3- Slightly Disagree' should be formatted as '3 - Slightly Disagree' to match the other anchors; also add a space after '2 - Disagree' for consistency.
- [III-D] The description of the RAG mechanism does not specify retrieval details (e.g., number of snippets per section, similarity threshold, chunk size). Providing these parameters would improve reproducibility, consistent with the open-science claim.
Circularity Check
Registered protocol for comparing LLM- and human-generated evidence briefings; no fitted parameters, no derivation by construction, and no load-bearing self-citation.
full rationale
This is a registered report describing a planned controlled experiment, not a paper claiming derived predictions. No parameter is fitted to data and then renamed as a prediction; the LLM tool is configured with fixed, documented parameters and its outputs are archived rather than tuned against the dependent variables. The hypotheses H0_1 through H0_3 (Section III-C) are testable comparisons of perceived content fidelity, ease of understanding, and usefulness, and the measurement instruments are adapted from an external taxonomy (Tang et al. [21]) and external survey guidelines, not defined circularly in terms of the hypotheses. The human-made briefings used as comparators come from prior work [7], [8] with some author overlap, but this is a baseline artifact rather than a load-bearing theorem: the experiment does not assume human briefings are correct, and the outcome is an empirical comparison by third-party participants. The RAG database contains 54 human-made briefings as style exemplars, but it explicitly excludes the two briefings under evaluation, so the generated outputs are not copied from the objects of comparison. The skeptic concern about content-fidelity items omitting completeness is a construct-validity threat, not a circular step, because the paper does not define fidelity solely in terms of those items; it operationalizes fidelity through a subset of error types while also collecting open-text responses. No equation, derivation, or cited uniqueness theorem reduces the protocol's claims to its own inputs.
Assumptions & free parameters
free parameters (3)
- LLM temperature =
0.5
- Top-p =
1.0
- Max tokens =
1024
assumptions (4)
- domain assumption Evidence briefings, as defined by Cartaxo et al., are a valid and useful knowledge-transfer medium for SE practice.
- domain assumption The Tang et al. error taxonomy can be operationalized as Likert items to measure content fidelity.
- domain assumption Convenience-sampled volunteers, randomly assigned, yield unbiased estimates of treatment differences.
- domain assumption Researchers and practitioners are distinct populations with appropriate evaluation roles.
Cite this review
Pith. "Pith review of Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering." pith.science (2026). https://pith.science/paper/IFWERXI4
@misc{pith2026250715828,
author = {Pith},
title = {Pith review of: Investigating the Use of LLMs for Evidence Briefings Generation in Software Engineering},
year = {2026},
howpublished = {\url{https://pith.science/paper/IFWERXI4}},
note = {Machine review of arXiv:2507.15828}
}
read the original abstract
[Context] An evidence briefing is a concise and objective transfer medium that can present the main findings of a study to software engineers in the industry. Although practitioners and researchers have deemed Evidence Briefings useful, their production requires manual labor, which may be a significant challenge to their broad adoption. [Goal] The goal of this registered report is to describe an experimental protocol for evaluating LLM-generated evidence briefings for secondary studies in terms of content fidelity, ease of understanding, and usefulness, as perceived by researchers and practitioners, compared to human-made briefings. [Method] We developed an RAG-based LLM tool to generate evidence briefings. We used the tool to automatically generate two evidence briefings that had been manually generated in previous research efforts. We designed a controlled experiment to evaluate how the LLM-generated briefings compare to the human-made ones regarding perceived content fidelity, ease of understanding, and usefulness. [Results] To be reported after the experimental trials. [Conclusion] Depending on the experiment results.
Figures
Reference graph
Works this paper leans on
-
[1]
The design of a survey on bridging the gap between software industry expectations and academia
Deniz Akdur. The design of a survey on bridging the gap between software industry expectations and academia. In2019 8th Mediterranean Conference on Embedded Computing (MECO), pages 1–5. IEEE, 2019
work page 2019
-
[2]
Deepika Badampudi, Farnaz Fotrousi, Bruno Cartaxo, and Muhammad Usman. Reporting consent, anonymity and confidentiality procedures adopted in empirical studies using human participants.e-Informatica Software Engineering Journal, 16(1):220109, July 2022. Available online: 22 Jul. 2022
work page 2022
-
[3]
Rouge metric evaluation for text summarization techniques.Available at SSRN 4120317, 2022
Marcello Barbella and Genoveffa Tortora. Rouge metric evaluation for text summarization techniques.Available at SSRN 4120317, 2022
work page 2022
-
[4]
V .R. Basili and H.D. Rombach. The tame project: towards improvement- oriented software environments.IEEE Transactions on Software Engi- neering, 14(6):758–773, 1988
work page 1988
-
[5]
Making software engineering research relevant.Computer, 47(4):80–83, 2014
Sarah Beecham, P ´adraig O’Leary, Sean Baker, Ita Richardson, and John Noll. Making software engineering research relevant.Computer, 47(4):80–83, 2014
work page 2014
-
[6]
David Budgen, Pearl Brereton, Nikki Williams, and Sarah Drummond. What support do systematic reviews provide for evidence-informed teaching about software engineering practice?e-informatica software engineering journal, 14(1):7–60, 2020
work page 2020
-
[7]
Towards a model to transfer knowledge from software engineering research to practice
Bruno Cartaxo, Gustavo Pinto, and Sergio Soares. Towards a model to transfer knowledge from software engineering research to practice. Information and Software Technology, 97:80–82, 2018
work page 2018
-
[8]
Bruno Cartaxo, Gustavo Pinto, Elton Vieira, and S ´ergio Soares. Ev- idence briefings: Towards a medium to transfer knowledge from sys- tematic reviews to practitioners. InProceedings of the 10th ACM/IEEE international symposium on empirical software engineering and mea- surement, ESEM ’16, pages 1–10, New York, NY , USA, 2016. Associ- ation for Computing...
work page 2016
Show all 28 references
-
[9]
Sage publications, 2017
John W Creswell and J David Creswell.Research design: Qualitative, quantitative, and mixed methods approaches. Sage publications, 2017
2017
-
[10]
Fabbri, Wojciech Kry ´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev
Alexander R. Fabbri, Wojciech Kry ´sci´nski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. Summeval: Re-evaluating summarization evaluation, 2021
2021
-
[11]
Practical relevance of software engineering research: synthesizing the community’s voice
Vahid Garousi, Markus Borg, and Markku Oivo. Practical relevance of software engineering research: synthesizing the community’s voice. Empirical Software Engineering, 25:1687–1754, 2020
2020
-
[12]
Aligning software engineering education with industrial needs: A meta-analysis.Journal of Systems and Software, 156:65–83, 2019
Vahid Garousi, G ¨orkem Giray, Eray T ¨uz¨un, Cagatay Catal, and Michael Felderer. Aligning software engineering education with industrial needs: A meta-analysis.Journal of Systems and Software, 156:65–83, 2019
2019
-
[13]
Key challenges in prompt engineering
Vladimir Geroimenko. Key challenges in prompt engineering. InThe Essential Guide to Prompt Engineering: Key Principles, Techniques, Challenges, and Security Risks, pages 85–102. Springer, 2025
2025
-
[14]
Hannay, Tore Dyb ˚a, Erik Arisholm, and Dag I.K
Jo E. Hannay, Tore Dyb ˚a, Erik Arisholm, and Dag I.K. Sjøberg. The effectiveness of pair programming: A meta-analysis.Information and Software Technology, 51(7):1110–1122, 2009. Special Section: Software Engineering for Secure Systems
2009
-
[15]
Springer Nature Switzerland, Cham, 2024
Marcos Kalinowski, Allysson Allex Ara ´ujo, and Daniel Mendez.Teach- ing Survey Research in Software Engineering, pages 501–527. Springer Nature Switzerland, Cham, 2024
2024
-
[16]
Jorgensen
Barbara Kitchenham, Tore Dyb ˚a, and M. Jorgensen. Evidence-based software engineering. InProceedings. 26th International Conference on Software Engineering, pages 273– 281, 06 2004
2004
-
[17]
What makes agile software development agile?IEEE Transactions on Software Engineering, 48(9):3523–3539, 2022
Marco Kuhrmann, Paolo Tell, Regina Hebig, et al. What makes agile software development agile?IEEE Transactions on Software Engineering, 48(9):3523–3539, 2022
2022
-
[18]
A systematic review on the use of definition of done on agile software development projects
Mirko Perkusich, Ana Silva, Thalles Ara ´ujo, Ednaldo Dilorenzo, Jo ˜ao Nunes, Hyggo Almeida, and Angelo Perkusich. A systematic review on the use of definition of done on agile software development projects. InProceedings of the 21st international conference on evaluation and...
2017
-
[19]
Building understandable messaging for policy and evidence review (bumper) with ai.arXiv preprint arXiv:2407.12812, 2024
Katherine A Rosenfeld, Maike Sonnewald, Sonia J Jindal, Kevin A McCarthy, and Joshua L Proctor. Building understandable messaging for policy and evidence review (bumper) with ai.arXiv preprint arXiv:2407.12812, 2024
2024 arXiv
-
[20]
Motivation to perform systematic reviews and their impact on software engineering practice
Ronnie ES Santos and Fabio QB Da Silva. Motivation to perform systematic reviews and their impact on software engineering practice. In2013 ACM/IEEE international symposium on empirical software engineering and measurement, pages 292–295. IEEE, 2013
2013
-
[21]
Evaluating large language models on medical evidence summa- rization.NPJ digital medicine, 6(1):158, 2023
Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G Nestor, Ali Soroush, Pierre A Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F Rousseau, et al. Evaluating large language models on medical evidence summa- rization.NPJ digital medicine, 6(1):158, 2023
2023
-
[22]
Gomez, Łukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. InProceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, page 6000–6010, Red Ho...
2017
-
[23]
Towards evaluation guidelines for empirical studies involving llms
Stefan Wagner, Marvin Mu ˜noz Bar ´on, Davide Falessi, and Sebastian Baltes. Towards evaluation guidelines for empirical studies involving llms. InProceedings of the 2nd International Workshop on Method- ological Issues with Empirical Studies in Software Engineering (WSESE 202...
2025 arXiv
-
[24]
Thoughts on applicability.Journal of Systems and Software, 215:112086, 2024
Titus Winters. Thoughts on applicability.Journal of Systems and Software, 215:112086, 2024
2024
-
[25]
Springer Nature, 2024
Claes Wohlin, Per Runeson, Martin H ¨ost, Magnus C Ohlsson, Bj ¨o Regnell, and Anders Wessl´en.Experimentation in Software Engineering. Springer Nature, 2024
2024
-
[26]
Closing the gap between open source and commercial large language models for medical evidence summarization.NPJ digital medicine, 7(1):239, 2024
Gongbo Zhang, Qiao Jin, Yiliang Zhou, Song Wang, Betina Idnay, Yiming Luo, Elizabeth Park, Jordan G Nestor, Matthew E Spotnitz, Ali Soroush, et al. Closing the gap between open source and commercial large language models for medical evidence summarization.NPJ digital medicine,...
2024
-
[27]
Yu, and Jiawei Zhang
Haopeng Zhang, Philip S. Yu, and Jiawei Zhang. A systematic survey of text summarization: From statistical methods to large language models, 2024
2024
-
[28]
Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019
1904 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.