Pith. sign in

REVIEW 3 major objections 2 minor 68 references

Quality defects in use case requirements can either improve or impair the performance of automated traceability recovery methods.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Empirical analysis of 189 annotated use cases shows that some requirements quality defects reduce TLR performance while others improve it, with effects varying by approach type.

T0 review reviewed 2026-06-27 challenge →

load-bearing objection The paper measures how 28 specific defects affect five TLR approaches on two datasets, but annotation details are missing from the abstract. the 3 major comments →

arxiv 2606.11834 v1 pith:4RDFAJTG submitted 2026-06-10 cs.SE

How Requirements Quality Makes (or Breaks) Traceability Link Recovery

classification cs.SE
keywords requirements qualitytraceability link recoveryuse casesempirical studysoftware traceabilityquality defectsTLR approachesrequirements engineering
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper investigates how defects in requirements quality influence automated traceability link recovery (TLR). Researchers annotated 28 types of quality defects across 189 use cases from two datasets. They then applied five different TLR approaches and used statistical tests to measure the impact of each defect type on performance. Results indicate that some defects, such as sentences not starting with noun phrases, reduce recovery accuracy, whereas others, like use cases containing implementation details, enhance it. Different TLR approaches show varying sensitivities to these defects, implying that optimal method selection depends on the specific quality profile of the requirements.

Core claim

Annotating 28 quality defect types in 189 use cases and evaluating five TLR approaches reveals that certain defects harm traceability performance while others benefit it, with approach-specific responses.

What carries the argument

Statistical analysis of the effect of manually annotated quality defects on TLR performance metrics across multiple approaches.

Load-bearing premise

The manual annotation of the 28 quality defect types in the 189 use cases is accurate, consistent, and free of annotator bias that could systematically affect the measured performance differences.

What would settle it

Re-annotating the same use cases with a different team of annotators and observing whether the performance effect sizes and directions remain the same.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper empirically examines how 28 types of quality defects in 189 use-case descriptions from two datasets affect the performance of five traceability link recovery (TLR) approaches. After manual annotation of the defects, the authors execute the TLR methods, measure link-recovery performance, and apply statistical tests to quantify effect strengths. Key findings are that certain defects (e.g., sentences not beginning with noun phrases) degrade TLR performance while others (e.g., use cases containing implementation details) improve it, and that different families of TLR approaches exhibit distinct sensitivities to the same defects.

Significance. If the annotation and statistical results prove reliable, the work supplies concrete, approach-specific evidence on the relationship between requirements quality and automated traceability. This could inform both requirements-writing guidelines and the selection of TLR techniques for a given dataset, moving beyond the generic assumption that higher quality always improves downstream automation.

major comments (3)
  1. [Annotation procedure (methods)] The description of the annotation procedure (abstract and methods) provides no information on the number of annotators, their training, conflict-resolution process, or inter-rater reliability metrics (e.g., Cohen’s or Fleiss’ kappa). Because the central claims rest on the labeling of 28 defect types across 189 use cases, the absence of these controls leaves open the possibility that measured performance deltas reflect annotation artifacts rather than the defects themselves.
  2. [Statistical analysis (results)] The manuscript does not specify the exact statistical tests employed, the handling of multiple comparisons across 28 defects and five approaches, or any correction for family-wise error rate. Without these details it is impossible to assess whether the reported effect strengths and significance levels are robust.
  3. [Results (tables/figures)] Table or figure presenting the per-defect performance deltas should include effect-size measures (e.g., Cohen’s d or odds ratios) in addition to p-values; the current reporting of “effect strength” is insufficient to judge practical significance.
minor comments (2)
  1. [Abstract] The abstract states that “different types of approaches respond differently” but does not name the five TLR approaches or the taxonomy used to group them; this information should appear in the abstract or be cross-referenced to a table in the introduction.
  2. [Datasets (methods)] Dataset provenance and licensing for the two use-case collections should be stated explicitly to allow replication.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive feedback. We address each major comment below and will revise the manuscript accordingly to improve methodological transparency and reporting.

read point-by-point responses
  1. Referee: [Annotation procedure (methods)] The description of the annotation procedure (abstract and methods) provides no information on the number of annotators, their training, conflict-resolution process, or inter-rater reliability metrics (e.g., Cohen’s or Fleiss’ kappa). Because the central claims rest on the labeling of 28 defect types across 189 use cases, the absence of these controls leaves open the possibility that measured performance deltas reflect annotation artifacts rather than the defects themselves.

    Authors: We agree that these details are critical. The annotation was performed by two authors experienced in requirements engineering. They first completed a training phase by independently annotating a pilot set of 20 use cases, resolved disagreements via discussion until consensus, and then annotated the remaining cases with periodic joint reviews. Inter-rater reliability was measured with Cohen’s kappa on an overlapping sample of 30 use cases. We will add a dedicated methods subsection describing the full procedure, annotator count, training, conflict resolution, and the resulting kappa value. revision: yes

  2. Referee: [Statistical analysis (results)] The manuscript does not specify the exact statistical tests employed, the handling of multiple comparisons across 28 defects and five approaches, or any correction for family-wise error rate. Without these details it is impossible to assess whether the reported effect strengths and significance levels are robust.

    Authors: We will clarify the analysis in the revised manuscript. We applied the Mann-Whitney U test to compare performance metrics between defect-present and defect-absent groups. Multiple comparisons (28 defects × 5 approaches) were addressed with Bonferroni correction (adjusted α = 0.05/140). A new paragraph in the statistical analysis section will explicitly state the tests chosen, their rationale, and the correction applied. revision: yes

  3. Referee: [Results (tables/figures)] Table or figure presenting the per-defect performance deltas should include effect-size measures (e.g., Cohen’s d or odds ratios) in addition to p-values; the current reporting of “effect strength” is insufficient to judge practical significance.

    Authors: We agree that effect sizes are needed to interpret practical significance. We will update all relevant tables and figures to report Cohen’s d alongside p-values and the existing effect-strength descriptions. This addition will be made in the results section of the revised manuscript. revision: yes

Circularity Check

0 steps flagged

No circularity; purely empirical measurement study

full rationale

The paper performs an empirical investigation: annotating 28 defect types across 189 use cases, running five TLR approaches, measuring performance, and applying statistical tests to quantify effects. No equations, derivations, predictions, or self-referential definitions appear that reduce results to inputs by construction. Claims rest on observed data and tests rather than fitted parameters, self-citation chains, or ansatzes smuggled via prior work. The study is self-contained against external benchmarks with no load-bearing circular steps.

Axiom & Free-Parameter Ledger

0 free parameters · 2 axioms · 0 invented entities

The central claim rests on the assumption that manual defect annotation can be performed reliably and that the chosen statistical tests isolate causal effects of individual defect types. No free parameters or invented entities are introduced; the work uses standard empirical methods.

axioms (2)
  • domain assumption Manual annotation of 28 quality defect types can be performed consistently across 189 use cases without substantial annotator disagreement or bias.
    The study depends on the quality of the defect labels to attribute performance differences to specific defects.
  • domain assumption Statistical tests applied after running the five TLR approaches correctly quantify the effect strength of each defect type.
    The reported harm/benefit conclusions rest on the validity of these tests.

reviewed 2026-06-27 · how reviews work

0 comments
Cite this review

Pith. "Pith review of How Requirements Quality Makes (or Breaks) Traceability Link Recovery." pith.science (2026). https://pith.science/paper/4RDFAJTG

@misc{pith2026260611834,
  author       = {Pith},
  title        = {Pith review of: How Requirements Quality Makes (or Breaks) Traceability Link Recovery},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4RDFAJTG}},
  note         = {Machine review of arXiv:2606.11834}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Traceability information between requirements and source code greatly benefits the maintenance of a software system. Since manually establishing trace links is cumbersome and error-prone, previous research explored automated traceability link recovery (TLR) approaches to support this task. However, quality defects in requirements impact subsequent activities such as TLR, yet evidence about this remains scarce. Our objective is to contribute empirical evidence on this impact. At the same time, we aim to understand how the performance of TLR approaches varies given these quality defects. To this end, we annotated 28 types of quality defect in 189 use case descriptions from two datasets. Then, we executed five distinct TLR approaches on the dataset and measured their performance in recovering trace links. Finally, we performed statistical tests to quantify the defects' effect strength on this performance. Our results show that some quality defects harm TLR performance, e.g., sentences that do not start with noun phrases, while others actually benefit performance, e.g., use cases that include implementation details. Moreover, different types of approaches respond differently to these defects. As a consequence, the performance-optimizing choice of a TLR approach depends on the quality of the dataset.

Figures

Figures reproduced from arXiv: 2606.11834 by Julian Frattini, Tobias Hey.

Figure 1
Figure 1. Figure 1: Use case description adapted from the eTour dataset with quality [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of the phases of this study TABLE I OVERVIEW OF THE USED TRACEABILITY LINK RECOVERY DATASETS ILLUSTRATING THE NUMBER OF REQUIREMENTS (REQ.), EXTRACTED SENTENCES (SENT.), CODE FILES AND TRACE LINKS (REQ. TO CODE) Number of Artifacts Dataset Domain SLOC Req. Sent. Code TLs eTour Tourism 12.4k 58 266 116 308 iTrust Healthcare 14.6k 131 333 226 286 Total: 189 599 342 594 [PITH_FULL_IMAGE:figures/full… view at source ↗
Figure 3
Figure 3. Figure 3: Posterior coefficient distribution of (at least weakly) significant coefficients from the analysis [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Conditional effects between the approach and significant quality factors [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Marginal effects on the outcome metrics While the x-axes of both plots show the different levels of the two factors, the two y-axes show the expected average value (as a line) and the 95% CI around it. Both plots show that recall, while highest, is affected the least and, therefore, shows the flattest slope, which is consistent with the coefficient values in [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

68 extracted references · 1 canonical work pages

  1. [1]

    Cleland-Huang, O

    J. Cleland-Huang, O. Gotel, A. Zismanet al.,Software and Systems Traceability. Springer, 2012, vol. 2

  2. [2]

    Empirical studies on software traceability: A mapping study,

    S. Charalampidou, A. Ampatzoglou, E. Karountzos, and P. Avgeriou, “Empirical studies on software traceability: A mapping study,”Journal of Software: Evolution and Process, vol. 33, no. 2, p. e2294, Feb. 2021

  3. [3]

    Requirements classification for trace- ability link recovery,

    T. Hey, J. Keim, and S. Corallo, “Requirements classification for trace- ability link recovery,” in2024 IEEE 32nd International Requirements Engineering Conference (RE’24), 2024

  4. [4]

    On the Impact of Requirements Smells in Prompts: The Case of Automated Traceability,

    A. V ogelsang, A. Korn, G. Broccia, A. Ferrari, J. Fischbach, and C. Arora, “On the Impact of Requirements Smells in Prompts: The Case of Automated Traceability,” inProceedings of the 2025 ACM/IEEE 45th International Conference on Software Engineering: New Ideas and Emerging Results, 2025

  5. [5]

    Assessing the quality of use case descriptions,

    K. T. Phalp, J. Vincent, and K. Cox, “Assessing the quality of use case descriptions,”Software Quality Journal, vol. 15, pp. 69–97, 2007

  6. [6]

    Requirements quality is quality in use,

    H. Femmer and A. V ogelsang, “Requirements quality is quality in use,” IEEE Software, vol. 36, no. 3, pp. 83–91, 2018

  7. [7]

    It’s the activities, stupid! a new perspective on re quality,

    H. Femmer, J. Mund, and D. M. Fernández, “It’s the activities, stupid! a new perspective on re quality,” in2015 IEEE/ACM 2nd International Workshop on Requirements Engineering and Testing. IEEE, 2015, pp. 13–19

  8. [8]

    Requirements quality research: a harmonized theory, evaluation, and roadmap,

    J. Frattini, L. Montgomery, J. Fischbach, D. Mendez, D. Fucci, and M. Unterkalmsteiner, “Requirements quality research: a harmonized theory, evaluation, and roadmap,”Requirements Engineering, pp. 1–14, 2023

  9. [9]

    Measuring the fitness-for-purpose of requirements: An initial model of activities and attributes,

    J. Frattini, J. Fischbach, D. Fucci, M. Unterkalmsteiner, and D. Mendez, “Measuring the fitness-for-purpose of requirements: An initial model of activities and attributes,” in2024 IEEE 30th International Requirements Engineering Conference (RE). IEEE, 2024

  10. [10]

    On the impact of passive voice requirements on domain modelling,

    H. Femmer, J. Ku ˇcera, and A. Vetrò, “On the impact of passive voice requirements on domain modelling,” inProceedings of the 8th ACM/IEEE international symposium on empirical software engineering and measurement, 2014, pp. 1–4

  11. [11]

    A live extensible ontology of quality factors for textual requirements,

    J. Frattini, L. Montgomery, J. Fischbach, M. Unterkalmsteiner, D. Mendez, and D. Fucci, “A live extensible ontology of quality factors for textual requirements,” in2022 IEEE 30th International Requirements Engineering Conference (RE). IEEE, 2022, pp. 274–280

  12. [12]

    Recovering traceability links between code and documentation,

    G. Antoniol, G. Canfora, G. Casazza, A. D. Lucia, and E. Merlo, “Recovering traceability links between code and documentation,”IEEE Transactions on Software Engineering, vol. 28, no. 10, pp. 970–983, Oct. 2002

  13. [13]

    Recovering Documentation-to-source-code Traceability Links Using Latent Semantic Indexing,

    A. Marcus and J. I. Maletic, “Recovering Documentation-to-source-code Traceability Links Using Latent Semantic Indexing,” inProceedings of the 25th International Conference on Software Engineering, ser. ICSE ’03. Washington, DC, USA: IEEE Computer Society, 2003, pp. 125– 135

  14. [14]

    Software trace- ability with topic modeling,

    H. U. Asuncion, A. U. Asuncion, and R. N. Taylor, “Software trace- ability with topic modeling,” in2010 ACM/IEEE 32nd International Conference on Software Engineering, vol. 1, May 2010, pp. 95–104

  15. [15]

    On inte- grating orthogonal information retrieval methods to improve traceability recovery,

    M. Gethers, R. Oliveto, D. Poshyvanyk, and A. D. Lucia, “On inte- grating orthogonal information retrieval methods to improve traceability recovery,” in2011 27th IEEE International Conference on Software Maintenance (ICSM), Sep. 2011, pp. 133–142

  16. [16]

    Improving the effectiveness of traceability link recovery using hierarchical bayesian networks,

    K. Moran, D. N. Palacio, C. Bernal-Cárdenas, D. McCrystal, D. Poshy- vanyk, C. Shenefiel, and J. Johnson, “Improving the effectiveness of traceability link recovery using hierarchical bayesian networks,” inPro- ceedings of the ACM/IEEE 42nd International Conference on Software Engineering, ser. ICSE ’20. New York, NY , USA: Association for Computing Machi...

  17. [17]

    When and How Using Structural Information to Improve IR-Based Traceability Recovery,

    A. Panichella, C. McMillan, E. Moritz, D. Palmieri, R. Oliveto, D. Poshyvanyk, and A. D. Lucia, “When and How Using Structural Information to Improve IR-Based Traceability Recovery,” in2013 17th European Conference on Software Maintenance and Reengineering, Mar. 2013, pp. 199–208

  18. [18]

    Can Method Data Dependencies Support the Assessment of Traceabil- ity Between Requirements and Source Code?

    H. Kuang, P. Mäder, H. Hu, A. Ghabi, L. Huang, J. Lü, and A. Egyed, “Can Method Data Dependencies Support the Assessment of Traceabil- ity Between Requirements and Source Code?”J. Softw. Evol. Process, vol. 27, no. 11, pp. 838–866, Nov. 2015

  19. [19]

    Using Consensual Biterms from Text Structures of Requirements and Code to Improve IR-Based Traceability Recovery,

    H. Gao, H. Kuang, K. Sun, X. Ma, A. Egyed, P. Mäder, G. Rong, D. Shao, and H. Zhang, “Using Consensual Biterms from Text Structures of Requirements and Code to Improve IR-Based Traceability Recovery,” inProceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, ser. ASE ’22. New York, NY , USA: Association for Computing M...

  20. [20]

    TRIAD: Automated Traceability Recovery based on Biterm-enhanced Deduction of Transitive Links among Artifacts,

    H. Gao, H. Kuang, W. K. G. Assunção, C. Mayr-Dorn, G. Rong, H. Zhang, X. Ma, and A. Egyed, “TRIAD: Automated Traceability Recovery based on Biterm-enhanced Deduction of Transitive Links among Artifacts,” inProceedings of the IEEE/ACM 46th International Conference on Software Engineering, ser. ICSE ’24. New York, NY , USA: Association for Computing Machine...

  21. [21]

    Improving Traceability Link Recovery Using Fine-grained Requirements-to-Code Relations,

    T. Hey, F. Chen, S. Weigelt, and W. F. Tichy, “Improving Traceability Link Recovery Using Fine-grained Requirements-to-Code Relations,” in2021 IEEE International Conference on Software Maintenance and Evolution (ICSME), Sep. 2021, pp. 12–22

  22. [22]

    Semantically Enhanced Software Traceability Using Deep Learning Techniques,

    J. Guo, J. Cheng, and J. Cleland-Huang, “Semantically Enhanced Software Traceability Using Deep Learning Techniques,” inProceedings of the 39th International Conference on Software Engineering, ser. ICSE ’17. Piscataway, NJ, USA: IEEE Press, 2017, pp. 3–14

  23. [23]

    Enhancing Automated Re- quirements Traceability by Resolving Polysemy,

    W. Wang, N. Niu, H. Liu, and Z. Niu, “Enhancing Automated Re- quirements Traceability by Resolving Polysemy,” in2018 IEEE 26th International Requirements Engineering Conference (RE), Aug. 2018, pp. 40–51

  24. [24]

    Tracing with Less Data: Active Learning for Classification-Based Traceability Link Recovery,

    C. Mills, J. Escobar-Avila, A. Bhattacharya, G. Kondyukov, S. Chakraborty, and S. Haiduc, “Tracing with Less Data: Active Learning for Classification-Based Traceability Link Recovery,” in2019 IEEE International Conference on Software Maintenance and Evolution (ICSME), Sep. 2019, pp. 103–113

  25. [25]

    Recovering Semantic Traceability between Requirements and Source Code Using Feature Rep- resentation Techniques,

    M. Zhang, C. Tao, H. Guo, and Z. Huang, “Recovering Semantic Traceability between Requirements and Source Code Using Feature Rep- resentation Techniques,” in2021 IEEE 21st International Conference on Software Quality, Reliability and Security (QRS), Dec. 2021, pp. 873– 882

  26. [26]

    Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT Models,

    J. Lin, Y . Liu, Q. Zeng, M. Jiang, and J. Cleland-Huang, “Traceability Transformed: Generating more Accurate Links with Pre-Trained BERT Models,” inProceedings of the 43rd International Conference on Soft- ware Engineering, ser. ICSE ’21. Madrid, Spain: IEEE Press, Nov. 2021, pp. 324–335

  27. [27]

    HGNNLink: Recover- ing requirements-code traceability links with text and dependency-aware heterogeneous graph neural networks,

    B. Wang, Z. Zou, X. Liang, H. Jin, and P. Liang, “HGNNLink: Recover- ing requirements-code traceability links with text and dependency-aware heterogeneous graph neural networks,”Autom Softw Eng, vol. 32, no. 2, p. 55, May 2025

  28. [28]

    Prompts Matter: Insights and Strategies for Prompt Engineering in Automated Software Traceability,

    A. D. Rodriguez, K. R. Dearstyne, and J. Cleland-Huang, “Prompts Matter: Insights and Strategies for Prompt Engineering in Automated Software Traceability,” in2023 IEEE 31st International Requirements Engineering Conference Workshops (REW), Sep. 2023, pp. 455–464

  29. [29]

    LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented Generation,

    D. Fuchß, T. Hey, J. Keim, H. Liu, N. Ewald, T. Thirolf, and A. Koziolek, “LiSSA: Toward Generic Traceability Link Recovery Through Retrieval- Augmented Generation,” in2025 IEEE/ACM 47th International Confer- ence on Software Engineering (ICSE), 2025, pp. 1396–1408

  30. [30]

    Requirements Traceability Link Recovery via Retrieval-Augmented Generation,

    T. Hey, D. Fuchß, J. Keim, and A. Koziolek, “Requirements Traceability Link Recovery via Retrieval-Augmented Generation,” inRequirements Engineering: Foundation for Software Quality, A. Hess and A. Susi, Eds. Cham: Springer Nature Switzerland, 2025, pp. 381–397

  31. [31]

    Beyond Retrieval: A Study of Using LLM Ensembles for Candidate Filtering in Requirements Traceability,

    D. Fuchß, S. Schwedt, J. Keim, and T. Hey, “Beyond Retrieval: A Study of Using LLM Ensembles for Candidate Filtering in Requirements Traceability,” in2025 IEEE 33rd International Requirements Engineer- ing Conference Workshops (REW), 2025, pp. 5–12

  32. [32]

    An Approach for Automating Use Case Refactoring,

    A. Rago, P. Frade, M. Ruiva, and C. A. Marcos, “An Approach for Automating Use Case Refactoring,”Electronic Journal of SADIO, vol. vol. 13, Jun. 2014

  33. [33]

    Quality improvement for use case model,

    R. Ramos, J. Castro, F. Alencar, J. Araújo, A. Moreira, C. d. E. da Computacao, and R. Penteado, “Quality improvement for use case model,” in2009 XXIII Brazilian Symposium on Software Engineering. IEEE, 2009, pp. 187–195

  34. [34]

    A method- ology for the classification of quality of requirements using machine learning techniques,

    E. Parra, C. Dimou, J. Llorens, V . Moreno, and A. Fraga, “A method- ology for the classification of quality of requirements using machine learning techniques,”Inf. Softw. Technol., vol. 67, no. C, pp. 180–195, Nov. 2015

  35. [35]

    An Empirical Study on Assessing the Quality of Use Case Metrics,

    C. Usdadiya, S. Tiwari, and A. Banerjee, “An Empirical Study on Assessing the Quality of Use Case Metrics,” inProceedings of the 12th Innovations in Software Engineering Conference (Formerly Known as India Software Engineering Conference), ser. ISEC ’19. New York, NY , USA: Association for Computing Machinery, Feb. 2019, pp. 1–11

  36. [36]

    Detecting requirements defects with NLP patterns: An industrial experience in the railway domain,

    A. Ferrari, G. Gori, B. Rosadini, I. Trotta, S. Bacherini, A. Fantechi, and S. Gnesi, “Detecting requirements defects with NLP patterns: An industrial experience in the railway domain,”Empirical Softw. Engg., vol. 23, no. 6, pp. 3684–3733, Dec. 2018

  37. [37]

    Which Require- ments Artifact Quality Defects are Automatically Detectable? A Case Study,

    H. Femmer, M. Unterkalmsteiner, and T. Gorschek, “Which Require- ments Artifact Quality Defects are Automatically Detectable? A Case Study,” in2017 IEEE 25th International Requirements Engineering Conference Workshops (REW), Sep. 2017, pp. 400–406

  38. [38]

    An Automatic Tool for the Analysis of Natural Language Requirements,

    G. Lami, S. Gnesi, F. Fabbrini, M. Fusani, and G. Trentanni, “An Automatic Tool for the Analysis of Natural Language Requirements,” 2004

  39. [39]

    Rapid quality assurance with Requirements Smells,

    H. Femmer, D. Méndez Fernández, S. Wagner, and S. Eder, “Rapid quality assurance with Requirements Smells,”Journal of Systems and Software, vol. 123, pp. 190–213, Jan. 2017

  40. [40]

    C. Y . Din and D. Rine,Requirements content goodness and complexity measurement based on NP chunks. VDM Publishing Saarbrücken, 2008

  41. [41]

    Detection of defective requirements using rule-based scripts,

    H. Hasso, H. Geppert, M. Dembach, and D. Toews, “Detection of defective requirements using rule-based scripts,” inInternational Con- ference on Requirements Engineering - Foundation for Software Quality (REFSQ) 2019, 2019

  42. [42]

    Supporting requirements update during software evolution,

    E. Ben Charrada, A. Koziolek, and M. Glinz, “Supporting requirements update during software evolution,”J. Softw. Evol. Process, vol. 27, no. 3, pp. 166–194, Mar. 2015

  43. [43]

    Adopting use case descriptions for re- quirements specification: an industrial case study,

    J. Frattini and A. Frattini, “Adopting use case descriptions for re- quirements specification: an industrial case study,” in2025 IEEE 31st International Requirements Engineering Conference (RE). IEEE, 2025

  44. [44]

    Can clone detection support quality assessments of requirements specifications?

    E. Juergens, F. Deissenboeck, M. Feilkas, B. Hummel, B. Schaetz, S. Wagner, C. Domann, and J. Streit, “Can clone detection support quality assessments of requirements specifications?” inProceedings of the 32nd ACM/IEEE International Conference on Software Engineering- Volume 2, 2010, pp. 79–88

  45. [45]

    How Requirements Quality Makes (or Breaks) Traceability Link Recovery - Replication Package,

    T. Hey and J. Frattini, “How Requirements Quality Makes (or Breaks) Traceability Link Recovery - Replication Package,” https://doi.org/10. 5281/zenodo.20448214, 2026

  46. [46]

    A coefficient of agreement for nominal scales,

    J. Cohen, “A coefficient of agreement for nominal scales,”Educational and psychological measurement, vol. 20, no. 1, pp. 37–46, 1960

  47. [47]

    Guidelines for Benchmark- ing Automated Software Traceability Techniques,

    Y . Shin, J. H. Hayes, and J. Cleland-Huang, “Guidelines for Benchmark- ing Automated Software Traceability Techniques,” in2015 IEEE/ACM 8th International Symposium on Software and Systems Traceability, May 2015, pp. 61–67

  48. [48]

    Pearl,Causality

    J. Pearl,Causality. Cambridge university press, 2009

  49. [49]

    Causal inference in statistics: An overview,

    ——, “Causal inference in statistics: An overview,”Statistical Surveys, 2009

  50. [50]

    Applications of statistical causal inference in software engineering,

    J. Siebert, “Applications of statistical causal inference in software engineering,”Information and Software Technology, vol. 159, p. 107198, 2023

  51. [51]

    McElreath,Statistical rethinking: A Bayesian course with examples in R and Stan

    R. McElreath,Statistical rethinking: A Bayesian course with examples in R and Stan. Chapman and Hall/CRC, 2018

  52. [52]

    Bayesian data analysis in em- pirical software engineering research,

    C. A. Furia, R. Feldt, and R. Torkar, “Bayesian data analysis in em- pirical software engineering research,”IEEE Transactions on Software Engineering, vol. 47, no. 9, pp. 1786–1810, 2019

  53. [53]

    Applying bayesian analysis guide- lines to empirical software engineering data: The case of programming languages and code quality,

    C. A. Furia, R. Torkar, and R. Feldt, “Applying bayesian analysis guide- lines to empirical software engineering data: The case of programming languages and code quality,”ACM Transactions on Software Engineering and Methodology (TOSEM), vol. 31, no. 3, pp. 1–38, 2022

  54. [54]

    Bayesian data analysis in em- pirical software engineering: The case of missing data,

    R. Torkar, R. Feldt, and C. A. Furia, “Bayesian data analysis in em- pirical software engineering: The case of missing data,”Contemporary Empirical Methods in Software Engineering, pp. 289–324, 2020

  55. [55]

    Applying bayesian data analysis for causal inference about requirements quality: a controlled experiment,

    J. Frattini, D. Fucci, R. Torkar, L. Montgomery, M. Unterkalmsteiner, J. Fischbach, and D. Mendez, “Applying bayesian data analysis for causal inference about requirements quality: a controlled experiment,” Empirical Software Engineering, vol. 30, no. 1, p. 29, 2025

  56. [56]

    Graphical causal models,

    F. Elwert, “Graphical causal models,” inHandbook of causal analysis for social research. Springer, 2013, pp. 245–273

  57. [57]

    A crash course in good and bad controls,

    C. Cinelli, A. Forney, and J. Pearl, “A crash course in good and bad controls,”Sociological Methods & Research, vol. 53, no. 3, pp. 1071– 1104, 2024

  58. [58]

    Bayesian workflow,

    A. Gelman, A. Vehtari, D. Simpson, C. C. Margossian, B. Carpenter, Y . Yao, L. Kennedy, J. Gabry, P.-C. Bürkner, and M. Modrák, “Bayesian workflow,”arXiv preprint arXiv:2011.01808, 2020

  59. [59]

    E. T. Jaynes,Probability theory: The logic of science. Cambridge: Cambridge University Press, 2003

  60. [60]

    Maximum likelihood estimation of models with beta- distributed dependent variables,

    P. Paolino, “Maximum likelihood estimation of models with beta- distributed dependent variables,”Political Analysis, vol. 9, no. 4, pp. 325–346, 2001

  61. [61]

    A general class of zero-or-one inflated beta regression models,

    R. Ospina and S. L. Ferrari, “A general class of zero-or-one inflated beta regression models,”Computational Statistics & Data Analysis, vol. 56, no. 6, pp. 1609–1623, 2012

  62. [62]

    Choosing priors in Bayesian ecological models by simulating from the prior predictive distribution,

    J. S. Wesner and J. P. Pomeranz, “Choosing priors in Bayesian ecological models by simulating from the prior predictive distribution,”Ecosphere, vol. 12, no. 9, p. e03739, 2021

  63. [63]

    Brooks, A

    S. Brooks, A. Gelman, G. Jones, and X.-L. Meng,Handbook of Markov Chain Monte Carlo. CRC press, 2011

  64. [64]

    A general framework for comparing predictions and marginal effects across models,

    T. D. Mize, L. Doan, and J. S. Long, “A general framework for comparing predictions and marginal effects across models,”Sociological Methodology, vol. 49, no. 1, pp. 152–189, 2019

  65. [65]

    Conditional regression analysis: Problems, solutions and an application,

    B. Denters and R. A. Van Puijenbroek, “Conditional regression analysis: Problems, solutions and an application,”Quality and Quantity, vol. 23, no. 1, pp. 83–108, 1989

  66. [66]

    Field study on requirements engineering: Investigation of artefacts, project parameters, and execution strategies,

    D. M. Fernandez, S. Wagner, K. Lochmann, A. Baumann, and H. de Carne, “Field study on requirements engineering: Investigation of artefacts, project parameters, and execution strategies,”Information and Software Technology, vol. 54, no. 2, pp. 162–178, 2012

  67. [67]

    Wohlin, P

    C. Wohlin, P. Runeson, M. Höst, M. C. Ohlsson, B. Regnell, A. Wesslén et al.,Experimentation in software engineering. Springer, 2012, vol. 236

  68. [68]

    You need 16 times the sample size to estimate an interaction than to estimate a main effect,

    A. Gelman, “You need 16 times the sample size to estimate an interaction than to estimate a main effect,” https://statmodeling.stat.columbia.edu/ 2018/03/15/need16/, accessed: 2023-11-24

This paper was first reviewed by grok-4.3 on June 27, 2026.