{"id":"2e0a8673-730d-456f-99e6-0688a353131f","arxiv_id":"1908.04366","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"An early-stage thesis proposal outlining a methodology for applying cohort, case-control, and cross-sectional observational studies to empirical software engineering.","lead":"This paper is an early-stage PhD plan proposing to adapt observational studies from medicine to software engineering research, with guidelines for running and reporting them. A generalist might read it to see how researchers hope to make causal claims about software development without randomized controlled experiments.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Plan's causal premise is unexamined: observational studies support but do not prove causality; Section 2.1.3 contradicts the abstract, and no causal identification strategy or validation against known causal effects is described.","rationale":"The reader's verdict of UNVERDICTED is appropriate because this is an early-stage PhD plan with no completed methodology, empirical results, or guidelines. My concern does not make the current non-result more rejectable; rather, it identifies a condition that any future acceptance must satisfy. The reader's weakest assumption focused on the availability of suitable replication datasets (Section 3.2). I agree that is a valid feasibility risk, but the more load-bearing concern is conceptual: the paper treats observational studies as capable of proving causality without articulating the causal identification assumptions required. The abstract's claim that controlled experiments are currently the only method for proving cause-effect relationships contradicts Section 2.1.3's admission that cross-sectional studies cannot prove causality. This internal inconsistency, combined with the absence of any confounding-control or sensitivity-analysis strategy, means the proposed replication-based validation might merely reproduce correlational results and still be mistaken for causal validation. Therefore the plan, as written, does not yet demonstrate a sound basis for its central promise. This supports keeping UNVERDICTED but with a clear warning that the methodological foundation, not just the data availability, needs substantial work.","tokens_in":8135,"tokens_out":3681,"duration_ms":39315,"concrete_test":"Select the strongest causal claim involving the Technical Debt Dataset (e.g., tech debt violations cause faults). Reanalyze the same data using the proposed cohort or case-control design and compare the estimated effect and its confidence interval against a benchmark from a randomized or otherwise identified counterfactual (e.g., a natural experiment or an instrumental-variable analysis). If the observational estimate is materially biased or if the guidelines omit any robustness check for unmeasured confounding, the causal-validity premise fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central promised contribution—guidelines enabling software engineering researchers to 'prove causality' via observational studies—requires that the planned methodology actually support causal inference. The paper never specifies the conditions under which an observational estimate is causal. The abstract and Section 1 assert that observational studies 'prove causality,' yet Section 2.1.3 states cross-sectional studies 'cannot be used to prove causality as there is no information about when events took place.' This is not a minor wording issue: cohort, case-control, and cross-sectional designs all require explicit assumptions about confounding, selection, and measurement (e.g., exchangeability, positivity, no unmeasured confounding) to support causal conclusions. The only validation described is 'replicating existing studies' (Section 3.1.3) and comparing results; but if the original studies were correlational or flawed, a replicated match does not validate causal claims. A legitimate causal claim would need, at minimum, adjustment for confounders, a strategy for unmeasured confounding (sensitivity analysis, instrumental variables, negative controls), and careful handling of time-dependent data. None of this appears. Section 3.2 even concedes results may not change when applying observational designs, which would leave the causal promise unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript, labeled an early-stage PhD plan, proposes to adapt epidemiological observational study designs (cohort, case-control, and cross-sectional) to empirical software engineering so that causality can be investigated without controlled experiments. It poses three research questions: which observational designs are applicable, which analysis techniques should be used, and how such studies should be reported. The proposed approach has four steps: a literature analysis, identification of observational methodologies, validation by replicating existing studies, and the production of reporting guidelines. The paper describes current progress including a dataset of Apache Java projects and two related publications, but it does not yet present a methodology, analysis techniques, or guidelines. The central causal claim, the validation strategy, and the plan for selecting analysis techniques are the main points of concern.","tokens_in":8352,"tokens_out":3817,"duration_ms":40571,"significance":"If the promised methodology were delivered and validated, it could provide empirical software engineering with a principled way to draw causal conclusions from observational repository data, with potential applications to mining software repositories, effort estimation, and technical debt research. The manuscript is transparent about several threats to validity, and the plan to build on existing datasets such as the Technical Debt Dataset and on medical reporting standards such as STROBE is sensible. However, as submitted, the paper is a research proposal rather than a substantive contribution: no methodology, analysis technique, or guideline is actually presented, and the central causal promise is not backed by a causal identification strategy or a validation design capable of supporting causal claims. The paper also contains an internal inconsistency about whether observational studies can 'prove causality.' These issues need to be resolved before the contribution can be evaluated.","major_comments":[{"comment":"The manuscript uses 'prove causality' in incompatible ways. The abstract states that 'Other fields use observational studies for proving causality,' and Section 3.1.4 says the guidelines will enable researchers to conduct studies 'which can prove causality.' In contrast, Section 2.1.3 states that cross-sectional studies 'cannot be used to prove causality as there is no information about when events took place,' and Section 1 states that 'causality cannot be proven without controlled experiments.' Because the central promised contribution is a methodology for causal inference, the paper must state precisely which observational designs are claimed to support causal conclusions and under which assumptions (for example, exchangeability, positivity, no unmeasured confounding, and correct handling of time-dependent confounding). The text as written is internally inconsistent and does not provide this.","section":"Abstract, Section 1, Section 2.1.3, Section 3.1.4"},{"comment":"The validation plan is not sufficient to support the causal promise. Replicating existing studies and comparing results (Section 3.1.3) can only check whether the new methodology reproduces the original findings; if the original studies are correlational or flawed, agreement does not demonstrate that the estimates are causal. The paper does not describe any identification strategy for confounding, selection, measurement error, or time-varying exposures, nor does it mention validation against known causal effects, negative controls, or simulations. Section 3.2's concession that results may not change when applying observational designs underscores that the plan as stated cannot validate causal claims.","section":"Section 3.1.3, Section 3.2"},{"comment":"The plan for selecting analysis techniques is underspecified. Step 2 says the author will 'investigate appropriate data analysis techniques' and suggests Markov chains or time series as examples, but it does not specify criteria for choosing among techniques, how dependent-data methods will be evaluated, or how the analysis will address confounding rather than just autocorrelation. Without this, RQ2 ('Which analysis techniques should be applied?') remains unanswered even as a research plan, and the promised 'validated analysis techniques for handling dependent data' are not concretely defined.","section":"Section 3.1.2, Section 3.2"}],"minor_comments":[{"comment":"The text says 'odd-ratio' but should read 'odds ratio.'","section":"Section 2.1.2"},{"comment":"The sentence 'A major concern with case-studies is that is that if are not done properly, they can suffer from biases stemming from several sources' is grammatically incomplete and appears to refer to case-control studies rather than case studies; it should be rewritten.","section":"Section 2.1.2"},{"comment":"The sentence 'For example, case studies cannot achieve this and thus guidelines for such studies are not enough' conflates the 'case study' methodology, which is a separate research design, with the 'case-control' observational design discussed earlier; this should be clarified.","section":"Section 3.1.4"},{"comment":"Because the manuscript is explicitly an early-stage plan, it would help to state at the start, in both the abstract and the introduction, that the paper describes a research proposal and that no methodology or guidelines have yet been delivered. Some formulations, such as 'we will propose a set of methodologies' and 'the guidelines will be assessed,' are clear, but the abstract's conclusion in particular reads as if the methods already exist.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is a PhD proposal rather than a completed study, and the editor may wish to consider whether the venue accepts research plans as standalone papers. The central issues are fixable in a revision: the causal language needs to be made consistent, the validation design needs to be upgraded to include identification assumptions and benchmarks, and RQ2 needs a concrete plan for evaluating analysis techniques. If the venue expects completed empirical or methodological contributions, the paper may be out of scope regardless of revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a doctoral research plan, not a completed study. There is no new result here—no guidelines, no validated analysis techniques, no causal finding. What is genuinely useful is the problem identification: empirical software engineering lacks an accepted way to make causal claims from observational data, and the two cited cohort studies [11][12] show people are already improvising. A systematic methodology, analysis guidance for dependent data, and reporting guidelines would fill a real gap.\n\nThe plan is structured and honest. Step 1's literature review and Step 2's design identification are reasonable. Section 3.2 concedes explicitly that suitable replication studies may not exist because data may not be public or complete—that is a real threat and the author names it. The self-citations to the Technical Debt Dataset [25]-[29],[31],[32] are legitimate: they point to a concrete, shared artifact that could support replication.\n\nThe soft spots are equally clear. The abstract says observational studies \"prove causality,\" but Section 2.1.3 says cross-sectional studies cannot prove causality, and Section 1 says causality cannot be proven without controlled experiments. The stress-test note is right: this is not a wording nit. The whole promised contribution is causal methodology, yet the plan never states what would make an observational estimate causal—no discussion of confounding, exchangeability, sensitivity analysis, negative controls, or instrumental variables. And the validation step (3.1.3) is under-specified: replicating existing studies and seeing whether results match does not validate causal inference unless the original studies had known causal effects. If the original studies are correlational, a matched replication is still correlational.\n\nThese gaps are normal for an early-stage PhD plan, and they are fixable. The bigger issue is that as a preprint, there is nothing to evaluate yet. The background is standard epidemiology material, and the software-engineering-specific contribution is still in progress. Section 4 reports a dataset and some published analyses, but that work does not yet exercise the proposed methodology.\n\nWho is this for? A thesis advisor, a doctoral-symposium audience, or someone assessing the student's progress. For a regular SE journal or conference, I would desk-reject this today. For a workshop or early-stage research track, it deserves a serious look and constructive feedback. I would not cite it in my own work yet.\n\nNet: engage with the author if you can, but don't send this to peer review as a research contribution.","headline":"Early-stage PhD plan with a real gap and a sound structure, but no deliverable yet and the causal claims are over-stated.","tokens_in":8821,"tokens_out":3091,"would_cite":false,"duration_ms":29926,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An early-stage PhD plan would bring epidemiology's observational studies to software engineering to test causality without controlled experiments.","keywords":["observational studies","empirical software engineering","causality","cohort study","case-control study","cross-sectional study","dependent data","reporting guidelines"],"falsifier":"An exhaustive search of the empirical software engineering literature would falsify the feasibility premise if it found fewer than two studies with complete public datasets and fully reported analysis techniques; alternatively, a head-to-head comparison in which a cohort analysis of a development practice and a randomized controlled trial of the same practice yield conflicting causal estimates would refute the claim that observational studies can substitute for controlled experiments.","tokens_in":7939,"feed_emoji":"🩺","tokens_out":8026,"duration_ms":71289,"temperature":0.7,"pith_summary":"This early-stage thesis argues that empirical software engineering should adopt observational study designs from epidemiology—cohort, case-control, and cross-sectional studies—to investigate cause-effect relationships without running controlled experiments. The author's planned contribution is a methodology for applying these designs to software engineering data, validated analysis techniques for dependent data such as commit histories, and reporting guidelines comparable to those used in medical epidemiology. The motivation is that controlled experiments are currently seen as the only way to prove causality in software engineering, while medicine routinely uses observational studies for causal claims. This document is a research plan: the methodology and guidelines are promised, not yet delivered.","feed_headline":"Observational studies could prove causality without experiments","feed_subtitle":"A PhD plan adapts cohort and case-control designs to software engineering data, then validates them by replication.","key_machinery":"The carrier of the argument is the trio of observational study designs borrowed from medicine: cohort studies compare exposed and unexposed groups over time and analyze from exposure to outcome; case-control studies start with cases that have the outcome and controls that do not, then look backward at exposures; cross-sectional studies take a simultaneous snapshot of exposure and outcome. To these the author adds bias-control practices from epidemiology (selection, information, and recall bias) and the proposal that dependent software engineering data—commits in a repository, for example—should be analyzed with techniques such as time-series analysis or Markov chains rather than the methods used in typical case studies and experiments. The cohort design's clear temporal ordering is what would allow the methodology to claim causal insight.","core_discovery":"The central claim is that observational studies can be imported from epidemiology into empirical software engineering and used to test causal hypotheses when controlled experiments are impractical or not feasible. Cohort studies, which follow exposed and unexposed groups from exposure to outcome, are presented as the most immediately applicable—two recent software engineering studies have already used them—and case-control studies are suggested for rare outcomes, while cross-sectional studies are described as only able to establish prevalence, not causality. The proposed validation strategy is replication: re-run published software engineering studies that have public data and clearly reported analysis techniques, but conducted as epidemiological designs with analysis techniques suited to dependent data, and compare the results with the original findings. As released, the paper sets out this program for a doctoral thesis rather than reporting completed results.","pith_inferences":["Editorial inference: If this methodology matures, much retrospective software repository mining could be redescribed as observational epidemiology, with causal language justified by exposure-outcome timing and confounding control.","Editorial inference: The same approach could transfer beyond software engineering to any discipline with dense timestamped trace data, such as ML operations or hardware telemetry.","Editorial inference: A natural testable extension would be a case-control study linking static-analysis warnings to production failures, then comparing its causal estimate with a randomized experiment on the same intervention."],"forward_implications":["Researchers could investigate causal questions—such as whether code smells lead to faults or whether test-driven development improves retention—without assigning treatments or waiting for controlled experiments.","Repository mining and effort-estimation studies could be recast as cohort or case-control studies, making their causal assumptions explicit and their evidence levels comparable to medical observational studies.","A reporting standard for observational studies in software engineering would reduce the methodological divergence seen in the first two published cohort studies.","Validated analysis techniques for dependent data would improve how commit histories and other time-ordered data are analyzed across empirical software engineering."],"supporting_citations":[{"why":"Supplies the evidence that well-designed observational studies can give results similar to randomized controlled trials, the key premise for substituting them for experiments.","marker":"[16]"},{"why":"Corroborates the similarity of observational studies and randomized trials, reinforcing the paper's central claim.","marker":"[17]"},{"why":"Provides the medical evidence-level hierarchy that positions observational studies just below randomized controlled trials.","marker":"[14]"},{"why":"Defines cohort and case-control designs and the bias concerns that structure the proposed methodology.","marker":"[15]"},{"why":"One of the first software engineering cohort studies, used as evidence that the approach is applicable and as a target for methodology development.","marker":"[11]"},{"why":"Another early software engineering cohort study, cited to show the inconsistency in methods that motivates reporting guidelines.","marker":"[12]"},{"why":"The STROBE-ME reporting guideline from epidemiology that the proposed software engineering guidelines would build on.","marker":"[24]"},{"why":"Documents statistical analysis practice in empirical software engineering, motivating the need for better data analysis techniques.","marker":"[23]"},{"why":"The SZZ algorithm used to identify fault-inducing and fault-fixing commits in the author's baseline dataset.","marker":"[30]"}],"fun_headline_variants":["Borrowing epidemiology's toolkit to prove causality in software studies","Cohort and case-control designs for causal claims without experiments","Observational studies: a PhD plan to test causality in software engineering","When experiments are impractical, observational studies may prove causality","Observational designs from epidemiology to test causality in software engineering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The plan depends on finding already-published software engineering studies that have complete, publicly available datasets and clearly reported analysis techniques to replicate; the paper itself concedes that suitable studies may not exist because data is often not public.","fun_headline_variants_meta":{"raw":{"variants":["Borrowing epidemiology's toolkit to prove causality in software studies","Cohort and case-control designs for causal claims without experiments","Observational studies: a PhD plan to test causality in software engineering","When experiments are impractical, observational studies may prove causality","Observational designs from epidemiology to test causality in software engineering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000659,"raw_usage":{"total_tokens":3004,"prompt_tokens":928,"completion_tokens":2076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1992}},"tokens_in":544,"tokens_out":2076,"duration_ms":13559,"temperature":1.0,"reasoning_tokens":1992,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:43:27.097607+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An exhaustive search of the empirical software engineering literature would falsify the feasibility premise if it found fewer than two studies with complete public datasets and fully reported analysis techniques; alternatively, a head-to-head comparison in which a cohort analysis of a development practice and a randomized controlled trial of the same practice yield conflicting causal estimates would refute the claim that observational studies can substitute for controlled experiments.","supporting_citations":[{"cited_title":"A longitudinal cohort study on the retainment of test-driven development,","cited_arxiv_id":null,"evidence_quote":"Supplies the evidence that well-designed observational studies can give results similar to randomized controlled trials, the key premise for substituting them for experiments."},{"cited_title":"A reﬂection on diversity and inclusivity eﬀorts in a software engineering program,","cited_arxiv_id":null,"evidence_quote":"Corroborates the similarity of observational studies and randomized trials, reinforcing the paper's central claim."},{"cited_title":"Guidelines for performing systematic literature reviews in software engineering,","cited_arxiv_id":null,"evidence_quote":"Provides the medical evidence-level hierarchy that positions observational studies just below randomized controlled trials."},{"cited_title":"Guidelines for including grey literature and conducting multivocal literature reviews in software engineering,","cited_arxiv_id":null,"evidence_quote":"Defines cohort and case-control designs and the bias concerns that structure the proposed methodology."},{"cited_title":"Experimentation in software engineering,","cited_arxiv_id":null,"evidence_quote":"One of the first software engineering cohort studies, used as evidence that the approach is applicable and as a target for methodology development."},{"cited_title":"Reporting experiments in software engineering,","cited_arxiv_id":null,"evidence_quote":"Another early software engineering cohort study, cited to show the inconsistency in methods that motivates reporting guidelines."},{"cited_title":"Cohort studies: marching towards outcomes,","cited_arxiv_id":null,"evidence_quote":"The STROBE-ME reporting guideline from epidemiology that the proposed software engineering guidelines would build on."},{"cited_title":"Introduction to epidemiology: Ray m. merrill, thomas c. timmreck,","cited_arxiv_id":null,"evidence_quote":"Documents statistical analysis practice in empirical software engineering, motivating the need for better data analysis techniques."},{"cited_title":"On the diﬀuseness of code technical debt in open source projects of the apache ecosystem,","cited_arxiv_id":null,"evidence_quote":"The SZZ algorithm used to identify fault-inducing and fault-fixing commits in the author's baseline dataset."}],"review_version":1}