REVIEW 3 major objections 4 minor 37 references
Methodological Issues in Observational Studies
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read An early-stage PhD plan would bring epidemiology's observational studies to software engineering to test causality without controlled experiments.
desk verdict Early-stage PhD plan with a real gap and a sound structure, but no deliverable yet and the causal claims are over-stated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrier of the argument is the trio of observational study designs borrowed from medicine: cohort studies compare exposed and unexposed groups over time and analyze from exposure to outcome; case-control studies start with cases that have the outcome and controls that do not, then look backward at exposures; cross-sectional studies take a simultaneous snapshot of exposure and outcome. To these the author adds bias-control practices from epidemiology (selection, information, and recall bias) and the proposal that dependent software engineering data—commits in a repository, for example—should be analyzed with techniques such as time-series analysis or Markov chains rather than the methods used in typical case studies and experiments. The cohort design's clear temporal ordering is what would allow the methodology to claim causal insight.
What would settle it
An exhaustive search of the empirical software engineering literature would falsify the feasibility premise if it found fewer than two studies with complete public datasets and fully reported analysis techniques; alternatively, a head-to-head comparison in which a cohort analysis of a development practice and a randomized controlled trial of the same practice yield conflicting causal estimates would refute the claim that observational studies can substitute for controlled experiments.
Extended reading notes
Core claim
The central claim is that observational studies can be imported from epidemiology into empirical software engineering and used to test causal hypotheses when controlled experiments are impractical or not feasible. Cohort studies, which follow exposed and unexposed groups from exposure to outcome, are presented as the most immediately applicable—two recent software engineering studies have already used them—and case-control studies are suggested for rare outcomes, while cross-sectional studies are described as only able to establish prevalence, not causality. The proposed validation strategy is replication: re-run published software engineering studies that have public data and clearly reported analysis techniques, but conducted as epidemiological designs with analysis techniques suited to dependent data, and compare the results with the original findings. As released, the paper sets out this program for a doctoral thesis rather than reporting completed results.
Load-bearing premise
The plan depends on finding already-published software engineering studies that have complete, publicly available datasets and clearly reported analysis techniques to replicate; the paper itself concedes that suitable studies may not exist because data is often not public.
Editorial extensions
If this is right
- Researchers could investigate causal questions—such as whether code smells lead to faults or whether test-driven development improves retention—without assigning treatments or waiting for controlled experiments.
- Repository mining and effort-estimation studies could be recast as cohort or case-control studies, making their causal assumptions explicit and their evidence levels comparable to medical observational studies.
- A reporting standard for observational studies in software engineering would reduce the methodological divergence seen in the first two published cohort studies.
- Validated analysis techniques for dependent data would improve how commit histories and other time-ordered data are analyzed across empirical software engineering.
Reading between the lines
- Editorial inference: If this methodology matures, much retrospective software repository mining could be redescribed as observational epidemiology, with causal language justified by exposure-outcome timing and confounding control.
- Editorial inference: The same approach could transfer beyond software engineering to any discipline with dense timestamped trace data, such as ML operations or hardware telemetry.
- Editorial inference: A natural testable extension would be a case-control study linking static-analysis warnings to production failures, then comparing its causal estimate with a randomized experiment on the same intervention.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript, labeled an early-stage PhD plan, proposes to adapt epidemiological observational study designs (cohort, case-control, and cross-sectional) to empirical software engineering so that causality can be investigated without controlled experiments. It poses three research questions: which observational designs are applicable, which analysis techniques should be used, and how such studies should be reported. The proposed approach has four steps: a literature analysis, identification of observational methodologies, validation by replicating existing studies, and the production of reporting guidelines. The paper describes current progress including a dataset of Apache Java projects and two related publications, but it does not yet present a methodology, analysis techniques, or guidelines. The central causal claim, the validation strategy, and the plan for selecting analysis techniques are the main points of concern.
Significance. If the promised methodology were delivered and validated, it could provide empirical software engineering with a principled way to draw causal conclusions from observational repository data, with potential applications to mining software repositories, effort estimation, and technical debt research. The manuscript is transparent about several threats to validity, and the plan to build on existing datasets such as the Technical Debt Dataset and on medical reporting standards such as STROBE is sensible. However, as submitted, the paper is a research proposal rather than a substantive contribution: no methodology, analysis technique, or guideline is actually presented, and the central causal promise is not backed by a causal identification strategy or a validation design capable of supporting causal claims. The paper also contains an internal inconsistency about whether observational studies can 'prove causality.' These issues need to be resolved before the contribution can be evaluated.
major comments (3)
- [Abstract, Section 1, Section 2.1.3, Section 3.1.4] The manuscript uses 'prove causality' in incompatible ways. The abstract states that 'Other fields use observational studies for proving causality,' and Section 3.1.4 says the guidelines will enable researchers to conduct studies 'which can prove causality.' In contrast, Section 2.1.3 states that cross-sectional studies 'cannot be used to prove causality as there is no information about when events took place,' and Section 1 states that 'causality cannot be proven without controlled experiments.' Because the central promised contribution is a methodology for causal inference, the paper must state precisely which observational designs are claimed to support causal conclusions and under which assumptions (for example, exchangeability, positivity, no unmeasured confounding, and correct handling of time-dependent confounding). The text as written is internally inconsistent and does not provide this.
- [Section 3.1.3, Section 3.2] The validation plan is not sufficient to support the causal promise. Replicating existing studies and comparing results (Section 3.1.3) can only check whether the new methodology reproduces the original findings; if the original studies are correlational or flawed, agreement does not demonstrate that the estimates are causal. The paper does not describe any identification strategy for confounding, selection, measurement error, or time-varying exposures, nor does it mention validation against known causal effects, negative controls, or simulations. Section 3.2's concession that results may not change when applying observational designs underscores that the plan as stated cannot validate causal claims.
- [Section 3.1.2, Section 3.2] The plan for selecting analysis techniques is underspecified. Step 2 says the author will 'investigate appropriate data analysis techniques' and suggests Markov chains or time series as examples, but it does not specify criteria for choosing among techniques, how dependent-data methods will be evaluated, or how the analysis will address confounding rather than just autocorrelation. Without this, RQ2 ('Which analysis techniques should be applied?') remains unanswered even as a research plan, and the promised 'validated analysis techniques for handling dependent data' are not concretely defined.
minor comments (4)
- [Section 2.1.2] The text says 'odd-ratio' but should read 'odds ratio.'
- [Section 2.1.2] The sentence 'A major concern with case-studies is that is that if are not done properly, they can suffer from biases stemming from several sources' is grammatically incomplete and appears to refer to case-control studies rather than case studies; it should be rewritten.
- [Section 3.1.4] The sentence 'For example, case studies cannot achieve this and thus guidelines for such studies are not enough' conflates the 'case study' methodology, which is a separate research design, with the 'case-control' observational design discussed earlier; this should be clarified.
- [Throughout] Because the manuscript is explicitly an early-stage plan, it would help to state at the start, in both the abstract and the introduction, that the paper describes a research proposal and that no methodology or guidelines have yet been delivered. Some formulations, such as 'we will propose a set of methodologies' and 'the guidelines will be assessed,' are clear, but the abstract's conclusion in particular reads as if the methods already exist.
Circularity Check
No circularity: the paper is a doctoral plan with no implemented derivation; planned validation by replication is not a fit-as-prediction, and self-citations concern dataset availability only.
full rationale
This is an early-stage PhD plan, not a derivation. The abstract promises future methodology, but no equations, model, or fitted parameters are presented, so there is no claimed derivation chain whose conclusion could reduce to its inputs. Section 3.1.2 and Section 3.1.3 describe a plan to identify methodologies and validate them by replicating existing studies and comparing results; a planned comparison is not a prediction from fitted input because no parameter is fit and no result is claimed. Section 3.2 concedes that suitable studies may not exist and cites the authors' own Technical Debt Dataset [25]-[29] only as a possible replication source. These self-citations support feasibility of finding datasets, not any load-bearing conclusion, and therefore do not constitute circularity. The internal tension between the abstract's claim that observational studies can 'prove causality' and Section 2.1.3's statement that cross-sectional studies 'cannot be used to prove causality' is a substantive correctness and consistency concern, not a derivation reducing to its inputs. Accordingly, no circular steps are identified, and the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Observational study designs can be transferred from medicine and epidemiology to software engineering.
- domain assumption Software engineering data, such as commit histories, are dependent and can be analyzed with time series or Markov chain techniques.
- domain assumption Suitable published studies with complete, public datasets exist for replication.
- domain assumption Observational studies can support causal inference without controlled experiments.
Cite this review
Pith. "Pith review of Methodological Issues in Observational Studies." pith.science (2026). https://pith.science/paper/JBPZLUVG
@misc{pith2026190804366,
author = {Pith},
title = {Pith review of: Methodological Issues in Observational Studies},
year = {2026},
howpublished = {\url{https://pith.science/paper/JBPZLUVG}},
note = {Machine review of arXiv:1908.04366}
}
read the original abstract
Background. Starting from the 1960s, practitioners and researchers have looked for ways to empirically investigate new technologies such as inspecting the effectiveness of new methods, tools, or practices. With this purpose, the empirical software engineering domain started to identify different empirical methods, borrowing them from various domains such as medicine, biology, and psychology. Nowadays, a variety of empirical methods are commonly applied in software engineering, ranging from controlled and quasi-controlled experiments to case studies, from systematic literature reviews to the newly introduced multivocal literature reviews. However, to date, the only available method for proving any cause-effect relationship are controlled experiments. Objectives. The goal of the thesis is introducing new methodologies for studying causality in empirical software engineering. Methods. Other fields use observational studies for proving causality. They allow observing the effect of a risk factor and testing this without trying to change who is or is not exposed to it. As an example, with an observational study it is possible to observe the effect of pollution on the growth of a forest or the effect of different factors on development productivity without the need of waiting years for the forest to grow or exposing developers to a specific treatment. Conclusion. In this thesis, we aim at defining a methodology for applying observational studies in empirical software engineering, providing guidelines on how to conduct such studies, how to analyze the data, and how to report the studies themselves.
Figures
Reference graph
Works this paper leans on
-
[11]
Experimentation in software engineering,
V. R. Basili, R. W. Selby, and D. H. Hutchens, “Experimentation in software engineering,” IEEE Trans. Softw. Eng., vol. 12, no. 7, pp. 733–743, Jul. 1986. [Online]. Available: http://dl.acm.org/citation.cfm?id=9775.9777
-
[12]
Reporting experiments in software engineering,
A. Jedlitschka, M. Ciolkowski, and D. Pfahl, “Reporting experiments in software engineering,” in In Guide to Advanced Empirical Software Engineering . Springer, 2008, pp. 201–228
work page 2008
-
[25]
A history of epidemiologic methods and concepts. vol. 1,
A. Morabia, “A history of epidemiologic methods and concepts. vol. 1,” 2004
work page 2004
-
[29]
V. Gallo, M. Egger, V. McCormack, P. B. Farmer, J. P. A. Ioannidis, M. Kirsch-Volders, G. Matullo, D. H. Phillips, B. Schoket, U. Stromberg, R. Vermeulen, C. Wild, M. Porta, and P. Vineis, “STrengthening the Reporting of OBservational studies in Epidemiology – Molecular Epidemiology (STROBE-ME): An extension of the STROBE statement,” Mutagenesis, vol. 27,...
work page 2011
-
[31]
A dynamical quality model to continuously monitor software maintenance,
V. Lenarduzzi, C. Stan, D. Taibi, D. Tosi, and G. Venters, “A dynamical quality model to continuously monitor software maintenance,” 2017
work page 2017
-
[32]
Towards surgically-precise technical debt estimation: Early results and research roadmap,
V. Lenarduzzi, A. Martini, D. Taibi, and D. A. Tamburri, “Towards surgically-precise technical debt estimation: Early results and research roadmap,” in 2019 IEEE Workshop on Machine Learning Techniques for Software Quality Evaluation (MaLTeSQuE), August 2019
work page 2019
-
[1]
Methodological Issues in Observational Studies
INTRODUCTION Software engineering is a relatively new field of research compared to other engineering disciplines such as mechanical engineering or civil engineering. It was formed in the 1960s when developers realized that understanding the code is not enough when creat- ing a piece of software. From that time on, software engineering research started to ...
work page Pith review arXiv 1908
-
[2]
BACKGROUND AND RELA TED WORK Different study methodologies have been proposed in the field of empirical software engineering. The most common ones are: Controlled (and quasi-) experiments [6][7], which originated from medicine and are used when researchers want to con- trol the behavior of different factors. Traditionally, there is one group that gets a trea...
Show all 37 references
-
[3]
The second RQ, concerning the applicability of analysis techniques, is answered in steps 2 and 3
THE PROPOSED APPROACH The PhD plan consists of four main steps: Step 1 : Analysis of the literature Step 1.1 Analysis of the literature on observational studies Step 1.2 Analysis of the literature on data analysis in SW Step 2 : Identification of methodologies for applying diffe...
2018
-
[4]
CURRENT STA TUS We started this PhD in July 2018. We have already investigated Step 1.2, performing a systematic literature review on the meth- ods and analysis techniques adopted in software maintenance and evolution models, and are currently performing a mapping study of the...
2018
-
[5]
EXPECTED CONTRIBUTION Currently, researchers in empirical software engineering have ac- cess to large quantities of data. However, many studies concen- trate on correlational findings, as there is no commonly accepted research methodology for performing causal studies without r...
-
[6]
A spiral model of software development and enhancement,
B. W. Boehm, “A spiral model of software development and enhancement,” Computer, vol. 21, no. 5, pp. 61–72, May 1988
1988
-
[7]
Manifesto for agile software development,
K. Beck, M. Beedle, A. van Bennekum, A. Cockburn, W. Cunningham, M. Fowler, J. Grenning, J. Highsmith, A. Hunt, R. Jeffries, J. Kern, B. Marick, R. C. Martin, S. Mellor, K. Schwaber, J. Sutherland, and D. Thomas, “Manifesto for agile software development,” 2001. [Online]. Avail...
2001
-
[8]
Lean software development,
M. Poppendieck, “Lean software development,” in Companion to the Proceedings of the 29th International Conference on Software Engineering , ser. ICSE COMPANION ’07, 2007, pp. 165–166
2007
-
[9]
Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 2002
Beck, Test Driven Development: By Example . Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 2002
2002
-
[10]
Dynamic fault tree models: Techniques for analysis of advanced fault tolerant computer systems,
M. A. Boyd, “Dynamic fault tree models: Techniques for analysis of advanced fault tolerant computer systems,” Ph.D. dissertation, Durham, NC, USA, 1992, uMI Order No. GAX92-02503
1992
-
[13]
Guidelines for conducting and reporting case study research in software engineering,
P. Runeson and M. H ¨ost, “Guidelines for conducting and reporting case study research in software engineering,” Empirical Software Engineering , vol. 14, pp. 131–164, April 2009
2009
-
[14]
Guidelines for performing systematic literature reviews in software engineering,
B. Kitchenham and S. Charters, “Guidelines for performing systematic literature reviews in software engineering,” 2007
2007
-
[15]
Guidelines for including grey literature and conducting multivocal literature reviews in software engineering,
V. Garousi, M. Felderer, and M. V. M ¨antyl¨a, “Guidelines for including grey literature and conducting multivocal literature reviews in software engineering,” Information and Software Technology, vol. 106, pp. 101 – 121, 2019
2019
-
[16]
A longitudinal cohort study on the retainment of test-driven development,
D. Fucci, S. Romano, M. T. Baldassarre, D. Caivano, G. Scanniello, B. Turhan, and N. Juristo, “A longitudinal cohort study on the retainment of test-driven development,” in 12th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement , ser. ESEM ’18, 2018
2018
-
[17]
A reflection on diversity and inclusivity efforts in a software engineering program,
D. S. Janzen, S. Bahrami, B. C. D. Silva, and D. Falessi, “A reflection on diversity and inclusivity efforts in a software engineering program,” in Proceedings - Frontiers in Education Conference, FIE , vol. 2018-October, 2019
2018
-
[18]
Statistics for epidemiology. nicholas p jewell. boca raton: Chapman & hall/crc, 2004, pp. 352,
R. Harbord, “Statistics for epidemiology. nicholas p jewell. boca raton: Chapman & hall/crc, 2004, pp. 352,” International Journal of Epidemiology - INT J EPIDEMIOL, vol. 33, pp. 1158–1159, 05 2004
2004
-
[19]
Introducing evidence-based medicine to plastic and reconstructive surgery,
K. C. Chung, J. A. Swanson, D. Schmitz, D. Sullivan, and R. J. Rohrich, “Introducing evidence-based medicine to plastic and reconstructive surgery,” Plastic and reconstructive surgery, vol. 123, no. 4, p. 1385, 2009
2009
-
[20]
Observational studies: cohort and case-control studies,
J. W. Song and K. C. Chung, “Observational studies: cohort and case-control studies,” Plastic and reconstructive surgery, vol. 126, no. 6, p. 2234, 2010
2010
-
[21]
A comparison of observational studies and randomized, controlled trials,
K. Benson and A. J. Hartz, “A comparison of observational studies and randomized, controlled trials,” New England Journal of Medicine , vol. 342, no. 25, pp. 1878–1886, 2000
2000
-
[22]
Randomized, controlled trials, observational studies, and the hierarchy of research designs,
J. Concato, N. Shah, and R. I. Horwitz, “Randomized, controlled trials, observational studies, and the hierarchy of research designs,” New England journal of medicine , vol. 342, no. 25, pp. 1887–1892, 2000
2000
-
[23]
Introduction to epidemiology: Ray m. merrill, thomas c. timmreck,
R. Merril, “Introduction to epidemiology: Ray m. merrill, thomas c. timmreck,” 2006
2006
-
[24]
Cohort studies: marching towards outcomes,
D. A. Grimes and K. F. Schulz, “Cohort studies: marching towards outcomes,” The Lancet, vol. 359, no. 9303, pp. 341–345, 2002
2002
-
[26]
Case-control studies: research in reverse,
K. F. Schulz and D. A. Grimes, “Case-control studies: research in reverse,” The Lancet, vol. 359, no. 9304, pp. 431–434, 2002
2002
-
[27]
Study design iii: Cross-sectional studies,
K. A. Levin, “Study design iii: Cross-sectional studies,” Evidence-based dentistry, vol. 7, no. 1, p. 24, 2006
2006
-
[28]
Evolution of statistical analysis in empirical software engineering research: Current state and steps forward,
F. G. de Oliveira Neto, R. Torkar, R. Feldt, L. Gren, C. A. Furia, and Z. Huang, “Evolution of statistical analysis in empirical software engineering research: Current state and steps forward,” Journal of Systems and Software , vol. 156, pp. 246 – 267, 2019
2019
-
[30]
On the diffuseness of code technical debt in open source projects of the apache ecosystem,
N. Saarim ¨aki, V. Lenarduzzi, and D. Taibi, “On the diffuseness of code technical debt in open source projects of the apache ecosystem,” International Conference on Technical Debt (TechDebt 2019) , 2019
2019
-
[33]
On the fault proneness of sonarqube technical debt violations: A comparison of eight machine learning techniques,
V. Lenarduzzi, F. Lomio, D. Taibi, and H. Huttunen, “On the fault proneness of sonarqube technical debt violations: A comparison of eight machine learning techniques,” 2019
2019
-
[34]
A continuous software quality monitoring approach for small and medium enterprises,
A. Janes, V. Lenarduzzi, and A. C. Stan, “A continuous software quality monitoring approach for small and medium enterprises,” in International Conference on Performance Engineering Companion, ser. ICPE ’17, 2017
2017
-
[35]
When do changes induce fixes?
A. Z. J. Sliwerski, T. Zimmermann, “When do changes induce fixes?” ser. MSR ’05. New York, NY, USA: ACM, 2005, pp. 1–5
2005
-
[36]
On the diffuseness of code technical debt in java projects of the apache ecosystem,
N. Saarim ¨aki, V. Lenarduzzi, and D. Taibi, “On the diffuseness of code technical debt in java projects of the apache ecosystem,” in International Conference on Technical Debt (TechDebt 2019) , 2019
2019
-
[37]
The technical debt dataset,
V. Lenarduzzi, N. Saarim ¨aki, and D. Taibi, “The technical debt dataset,” in 15th conference on PREdictive Models and data analysis In Software Engineering , ser. PROMISE ’19, 2019
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.