REVIEW 3 major objections 4 minor 64 references
Does Code Quality Affect Pull Request Acceptance? An empirical study
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Code quality measured by PMD rule violations in changed code does not affect whether a maintainer accepts or rejects a pull request.
desk verdict Good question, big dataset, but the paper's own contingency table contradicts its null claim; needs major revision before it can be believed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is PMD, an open-source static analysis tool whose default Java rule set flags more than 300 potential issues, including code smells, anti-patterns, style and documentation lapses, and possible bugs, each with a priority from P1 (critically broken) to P5 (cosmetic). The authors convert every violation appearing in a pull request's diff into a technical-debt item, producing 4.7 million observations across 253 rule types. The argument then runs through two complementary mechanisms that both register the same null result: a chi-square test on the contingency table of accepted/rejected versus violations/no violations, and eight classifiers whose accuracy metrics (AUC, precision, recall, MCC, F-measure) quantify the predictive signal in the violation counts. The machinery demonstrates the result by failing to exceed chance: near-50% AUC and near-zero MCC indicate that violation counts are statistically independent of the merge decision.
What would settle it
Look at pull requests whose human review comments explicitly criticize code quality; if those comments predict rejection while PMD counts do not, the paper's null result is a measurement artifact rather than evidence that quality is ignored.
Extended reading notes
Core claim
Using PMD's default Java rule set, the authors analyze each pull request's diff against the master branch and count every rule violation introduced by the contribution, treating each as a technical-debt item. They then test whether these counts separate accepted from rejected pull requests using a chi-square test on a contingency matrix and eight machine-learning classifiers (logistic regression, decision tree, bagging, random forest, extremely randomized trees, AdaBoost, gradient boosting, XGBoost). The central finding is that the counts do not separate the two groups: the chi-square statistic is 0.12, all classifiers hover at roughly 50% AUC, and grouping violations by PMD priority from P1 to P4 does not improve prediction. Of the 253 distinct PMD rules found in the data, 243 appear in both accepted and rejected pull requests, and the ten found only in rejected requests are too rare to move the models. The paper concludes that the presence of technical-debt items in pull request code does not influence acceptance or rejection at all.
Load-bearing premise
The assumption that carries the result is that automated rule violations in a pull request's changed lines capture the code quality maintainers actually judge when they accept or reject it.
Editorial extensions
If this is right
- In the 28 studied projects, maintainers do not appear to use PMD-detectable code quality as a screening criterion, so contributors cannot reliably raise acceptance odds by cleaning up style and design violations.
- Even the most severe PMD priorities, the ones closest to potential bugs, show no relationship with rejection, so the null result is not an artifact of flooding by low-severity style rules.
- The finding aligns with earlier evidence that acceptance is driven more by contributor reputation and the importance of the delivered feature than by measured code quality.
- Because every project independently shows the same pattern, the null result is not explained by one project's unusual review culture but appears general across the sampled ecosystem.
- The authors suggest that pre-submission quality checks can still improve code maintainability, even though on this evidence they will not raise the chance of acceptance.
Reading between the lines
- A natural extension is to test quality signals that focus on potential bugs and design flaws, such as compiler warnings, failing tests, or human reviewer comments, since the default PMD rule set is dominated by style-oriented priority-4 rules.
- The binary merged/not-merged outcome may hide quality effects that appear earlier in the process: maintainers might send low-quality pull requests back for revisions, so quality would shape latency and revision cycles rather than final acceptance.
- Contributor reputation may moderate the quality effect: trusted contributors could get the benefit of the doubt while newcomers with identical code are rejected, and aggregate models would average that interaction away.
- A replication on projects with enforced quality gates, where continuous integration blocks merging on rule violations, would clarify whether the null result reflects maintainer indifference or simply the absence of such gates in the sampled projects.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a case study of pull request acceptance in 28 Java open-source projects. Using PMD with its default rule set, the authors count technical-debt (TD) items in 36,344 pull requests and relate these counts to whether each pull request was merged. RQ1 describes the distribution of TD items across projects and priorities; RQ2 tests whether TD item presence is associated with acceptance using a chi-square test and eight classifiers; RQ3 repeats the modeling after grouping TD items by PMD priority. The paper's central conclusion is that "the presence of TD items of all types in the pull request code, does not influence the acceptance or rejection of pull requests at all." The authors also provide a replication package and use multiple machine learning methods to support the null result.
Significance. If valid, this negative result would be a noteworthy contribution to empirical software engineering, because it challenges the intuition that code quality as measured by static analysis affects maintainers' merge decisions. The study is large in scale, covers a diverse set of mature Java projects, and the authors provide a replication package, which are strengths. However, the paper's single reported chi-square test of independence is contradicted by its own printed contingency table, and the machine-learning results, while suggestive, cannot by themselves overcome that arithmetic inconsistency. The significance of the finding is therefore conditional on a resolution of the Table 11 discrepancy; as printed, the central claim is unsupported.
major comments (3)
- [Section 5 (RQ2), Table 11] The printed contingency counts are incompatible with the reported chi-square value of 0.12. Interpreting the entries as counts (10,563; 8,558; 11,228; 5,528), Pearson's chi-square is approximately 518 on 1 degree of freedom, with p << 0.001, not 0.12. The acceptance rate for pull requests with TD items is 48.5% (10,563/21,791), while for pull requests without TD items it is 60.8% (8,558/14,086), a 12.2 percentage-point difference. This directly contradicts the sentence "the presence of TD items does not affect pull request acceptance" and the summary statement "χ2 0.12 and AUC 50%." Since Table 11 is the only direct statistical test of independence reported, the central claim is unsupported until the counts or the test statistic are reconciled.
- [Section 5 (RQ2), Figure 1 and Section 8] The conclusion "The same results are verified in all the 28 projects independently" is not backed by any per-project chi-square test or per-project machine-learning table in the manuscript. The text states that "some projects showed some moderate success" but dismisses them as outliers without reporting how many projects, what threshold was used, or what the per-project statistics were. Given that the pooled contingency test is invalid as printed, this per-project evidence is needed to support the universal formulation of the conclusion.
- [Sections 4.3, 4.4, and 7] The independent variable is a count of PMD default-rule violations, yet the research questions, abstract, and conclusions are phrased in terms of "code quality." The most frequent violations listed in Table 8 are predominantly priority-4 style rules, and Section 7 concedes both that PMD's detection accuracy has not been empirically assessed and that no comparison was made with maintainers' actual quality judgments. The null result should therefore be scoped to "PMD-detected TD item counts," not to code quality generally; as written, the "at all" wording in Section 8 overstates what the measurement supports.
minor comments (4)
- [Table 11] The table uses period separators (10.563) while other tables use commas (1,270), and the four counts sum to 35,877 rather than the 36,344 pull requests reported in Section 5; the formatting and the total should be clarified.
- [Abstract and Section 4.4] The abstract says "seven machine learning techniques" but lists six, while Section 4.4 states that eight classifiers were used; the count should be made consistent.
- [Table 2] Several entries appear truncated or erroneous, such as "1,27", "2,19", "5,52", "4,12", and "3,07", and the time frame for apache/cassandra is listed as "2018/10-2011/09", which appears reversed.
- [Throughout] There are typographical errors that should be corrected, including "Gausios" for Gousios, "Change Pronenes" in Table 1, and "3,6022" in Table 8.
Circularity Check
No circularity: the empirical design derives the null claim from measured PMD counts and held-out cross-validated classifiers, with no fitted input renamed as a prediction and no load-bearing self-citation.
full rationale
The paper is an empirical measurement study rather than a derivation from assumed axioms. The independent variables are PMD-detected Technical Debt item counts for each pull request, collected as described in Section 4.3, and the dependent variable is whether the pull request was merged, also determined in Section 4.3. The central claim that TD items do not affect pull request acceptance is an inference from the observed data, supported by the chi-square test, logistic regression, and seven machine learning classifiers (Sections 4.4 and 5). No model parameter is fitted to a subset of acceptance outcomes and then reported as a prediction of a closely related quantity: the classifiers are evaluated with 5-fold cross-validation, so the reported AUC values are out-of-sample estimates. The only self-citation, reference [18], supports the background statement that PMD is one of the four tools used most frequently for software analysis; it is not load-bearing for the null result. The internal arithmetic inconsistency in Table 11, where the printed contingency counts imply a chi-square near 518 rather than 0.12, is a correctness and reporting concern, not circularity, because the conclusion is not forced by construction from the inputs. No step in the paper reduces by definition or by self-citation to the target claim.
Assumptions & free parameters
assumptions (4)
- domain assumption PMD default Java rule set violations are a valid measure of the code quality relevant to pull request review.
- domain assumption GitHub merged status and closed-by-commit events correctly identify pull request acceptance.
- domain assumption PMD issues can be filtered to those introduced by the pull request by comparing against the master branch diff.
- domain assumption The 28 selected projects, 22 from Apache, are representative enough to support a general statement about open-source Java pull requests.
Cite this review
Pith. "Pith review of Does Code Quality Affect Pull Request Acceptance? An empirical study." pith.science (2026). https://pith.science/paper/NHGWIKWH
@misc{pith2026190809321,
author = {Pith},
title = {Pith review of: Does Code Quality Affect Pull Request Acceptance? An empirical study},
year = {2026},
howpublished = {\url{https://pith.science/paper/NHGWIKWH}},
note = {Machine review of arXiv:1908.09321}
}
read the original abstract
Background. Pull requests are a common practice for contributing and reviewing contributions, and are employed both in open-source and industrial contexts. One of the main goals of code reviews is to find defects in the code, allowing project maintainers to easily integrate external contributions into a project and discuss the code contributions. Objective. The goal of this paper is to understand whether code quality is actually considered when pull requests are accepted. Specifically, we aim at understanding whether code quality issues such as code smells, antipatterns, and coding style violations in the pull request code affect the chance of its acceptance when reviewed by a maintainer of the project. Method. We conducted a case study among 28 Java open-source projects, analyzing the presence of 4.7 M code quality issues in 36 K pull requests. We analyzed further correlations by applying Logistic Regression and seven machine learning techniques (Decision Tree, Random Forest, Extremely Randomized Trees, AdaBoost, Gradient Boosting, XGBoost). Results. Unexpectedly, code quality turned out not to affect the acceptance of a pull request at all. As suggested by other works, other factors such as the reputation of the maintainer and the importance of the feature delivered might be more important than code quality in terms of pull request acceptance. Conclusions. Researchers already investigated the influence of the developers' reputation and the pull request acceptance. This is the first work investigating if quality of the code in pull requests affects the acceptance of the pull request or not. We recommend that researchers further investigate this topic to understand if different measures or different tools could provide some useful measures.
Figures
Reference graph
Works this paper leans on
-
[1]
A. F. Ackerman, P. J. Fowler, R. G. Ebenau, Software inspections and the industrial production of software, in: Proc. Of a Symposium on Software Validation: Inspection-testing-verification-alternatives, pp. 13–40
-
[2]
A. F. Ackerman, L. S. Buchwald, F. H. Lewski, Software inspections: an effective verification process, IEEE Software 6 (1989) 31–36
work page 1989
-
[3]
M. E. Fagan, Design and code inspections to reduce errors in program development, IBM Systems Journal 15 (1976) 182–211
work page 1976
- [4]
-
[5]
D. G. Feitelson, E. Frachtenberg, K. L. Beck, Development and deploy- ment at facebook, IEEE Internet Computing 17 (2013) 8–17
work page 2013
- [6]
-
[7]
A. Bacchelli, C. Bird, Expectations, outcomes, and challenges of mod- ern code review, in: Proceedings of the 2013 International Conference on Software Engineering, ICSE ’13, pp. 712–721
work page 2013
- [8]
Show all 64 references
-
[9]
Gousios, M
G. Gousios, M. Pinzger, A. van Deursen, An exploratory study of the pull-based software development model, in: 36th International Confer- ence on Software Engineering, ICSE 2014, pp. 345–355. 27
2014
-
[10]
Gousios, A
G. Gousios, A. Zaidman, M. Storey, A. van Deursen, Work practices and challenges in pull-based development: The integrator’s perspec- tive, in: 37th IEEE International Conference on Software Engineering, volume 1, pp. 358–368
-
[11]
E. v. d. Veen, G. Gousios, A. Zaidman, Automatically prioritizing pull requests, in: 12th Working Conference on Mining Software Reposito- ries, pp. 357–361
-
[12]
Zampetti, L
F. Zampetti, L. Ponzanelli, G. Bavota, A. Mocci, M. D. Penta, M. Lanza, How developers document pull requests with external refer- ences, in: 25th International Conference on Program Comprehension (ICPC), volume 00, pp. 23–33
-
[13]
Y. Yu, H. Wang, G. Yin, C. X. Ling, Reviewer recommender of pull- requests in github, in: IEEE International Conference on Software Maintenance and Evolution, pp. 609–612
-
[14]
M. M. Rahman, C. K. Roy, An insight into the pull requests of github, in: 11th Working Conference on Mining Software Repositories, MSR 2014, pp. 364–367
2014
-
[15]
D. M. Soares, M. L. d. L. Jnior, L. Murta, A. Plastino, Rejection factors of pull requests filed by core team developers in software projects with high acceptance rates, in: 14th International Conference on Machine Learning and Applications (ICMLA), pp. 960–965
-
[16]
Kononenko, T
O. Kononenko, T. Rose, O. Baysal, M. Godfrey, D. Theisen, B. de Wa- ter, Studying pull request merges: A case study of shopify’s active merchant, in: 40th International Conference on Software Engineering: Software Engineering in Practice, ICSE-SEIP ’18, pp. 124–133
-
[17]
Calefato, F
F. Calefato, F. Lanubile, N. Novielli, A preliminary analysis on the effects of propensity to trust in distributed software development, in: 2017 IEEE 12th International Conference on Global Software Engineer- ing (ICGSE), pp. 56–60
2017
-
[18]
Lenarduzzi, A
V. Lenarduzzi, A. Sillitti, D. Taibi, A survey on code analysis tools for software maintenance prediction, in: 6th International Conference in Software Engineering for Defence Applications, Springer International Publishing, 2020, pp. 165–175. 28
2020
-
[19]
Beller, R
M. Beller, R. Bholanath, S. McIntosh, A. Zaidman, Analyzing the state of static analysis: A large-scale evaluation in open source software, in: 23rd International Conference on Software Analysis, Evolution, and Reengineering (SANER), volume 1, pp. 470–481
-
[20]
Fowler, K
M. Fowler, K. Beck, Refactoring: Improving the design of existing code, Addison-Wesley Longman Publishing Co., Inc. (1999)
1999
-
[21]
W. J. Brown, R. C. Malveau, H. W. S. McCormick, T. J. Mowbray, An- tiPatterns: Refactoring Software, Architectures, and Projects in Crisis: Refactoring Software, Architecture and Projects in Crisis, John Wiley and Sons, 1998
1998
-
[22]
Lanza, R
M. Lanza, R. Marinescu, S. Ducasse, Object-Oriented Metrics in Prac- tice, Springer-Verlag, Berlin, Heidelberg, 2005
2005
-
[23]
Cunningham, The wycash portfolio management system, OOPSLA ’92
W. Cunningham, The wycash portfolio management system, OOPSLA ’92
-
[25]
Olbrich, D
S. Olbrich, D. S. Cruzes, V. Basili, N. Zazworka, The evolution and impact of code smells: A case study of two open source systems, in: 2009 3rd International Symposium on Empirical Software Engineering and Measurement, pp. 390–400
2009
-
[26]
D’Ambros, A
M. D’Ambros, A. Bacchelli, M. Lanza, On the impact of design flaws on software defects, in: 2010 10th International Conference on Quality Software, pp. 23–31
2010
-
[27]
Fontana Arcelli, S
F. Fontana Arcelli, S. Spinelli, Impact of refactoring on quality code evaluation, in: Proceedings of the 4th Workshop on Refactoring Tools, WRT ’11, pp. 37–40
-
[28]
W. H. Brown, R. C. Malveau, H. W. S. McCormick, T. J. Mowbray, An- tiPatterns: Refactoring Software, Architectures, and Projects in Crisis, New York, NY, USA, 1st edition, 1998
1998
-
[29]
S. R. Chidamber, C. F. Kemerer, A metrics suite for object oriented design, IEEE Trans. Softw. Eng. 20 (1994) 476–493. 29
1994
-
[30]
Al Dallal, A
J. Al Dallal, A. Abdin, Empirical evaluation of the impact of object- oriented code refactoring on quality attributes: A systematic literature review, IEEE Transactions on Software Engineering 44 (2018) 44–69
2018
-
[31]
T. J. McCabe, A complexity measure, IEEE Trans. Softw. Eng. 2 (1976) 308–320
1976
-
[32]
W. Li, R. Shatnawi, An empirical study of the bad smells and class error probability in the post-release object-oriented system evolution, J. Syst. Softw. 80 (2007) 1120–1128
2007
-
[33]
D. I. K. Sjberg, A. Yamashita, B. C. D. Anda, A. Mockus, T. Dyb, Quantifying the effect of code smells on maintenance effort, IEEE Transactions on Software Engineering 39 (2013) 1144–1156
2013
-
[34]
Yamashita, Assessing the capability of code smells to explain mainte- nance problems: An empirical study combining quantitative and qual- itative data, Empirical Softw
A. Yamashita, Assessing the capability of code smells to explain mainte- nance problems: An empirical study combining quantitative and qual- itative data, Empirical Softw. Engg. 19 (2014) 1111–1143
2014
-
[35]
Palomba, G
F. Palomba, G. Bavota, M. D. Penta, F. Fasano, R. Oliveto, A. D. Lucia, On the diffuseness and the impact on maintainability of code smells: A large scale empirical investigation, Empirical Softw. Engg. 23 (2018) 1188–1221
2018
-
[36]
Khomh, M
F. Khomh, M. Di Penta, Y. Gueheneuc, An exploratory study of the impact of code smells on software change-proneness, in: 2009 16th Working Conference on Reverse Engineering, pp. 75–84
2009
-
[37]
Jaafar, Y.-G
F. Jaafar, Y.-G. Gu´ eh´ eneuc, S. Hamel, F. Khomh, M. Zulkernine, Eval- uating the impact of design pattern and anti-pattern dependencies on changes and faults, Empirical Softw. Engg. 21 (2016) 896–931
2016
-
[38]
S. M. Olbrich, D. S. Cruzes, D. I. K. Sjberg, Are all code smells harmful? a study of god classes and brain classes in the evolution of three open source systems, in: 2010 IEEE International Conference on Software Maintenance, pp. 1–10
2010
-
[39]
Schumacher, N
J. Schumacher, N. Zazworka, F. Shull, C. Seaman, M. Shaw, Building empirical support for automated code smell detection, in: Proceedings of the 2010 ACM-IEEE International Symposium on Empirical Soft- ware Engineering and Measurement, ESEM ’10, pp. 8:1–8:10. 30
2010
-
[40]
Zazworka, M
N. Zazworka, M. A. Shaw, F. Shull, C. Seaman, Investigating the impact of design debt on software quality, in: Proceedings of the 2Nd Workshop on Managing Technical Debt, MTD ’11, pp. 17–23
-
[41]
Du Bois, S
B. Du Bois, S. Demeyer, J. Verelst, T. Mens, M. Temmerman, Does god class decomposition affect comprehensibility?, pp. 346–355
-
[42]
H. Aman, S. Amasaki, T. Sasaki, M. Kawahara, Empirical analysis of fault-proneness in methods by focusing on their comment lines, in: 2014 21st Asia-Pacific Software Engineering Conference, volume 2, pp. 51–56
2014
-
[43]
Aman, An empirical analysis on fault-proneness of well-commented modules, in: 2012 Fourth International Workshop on Empirical Soft- ware Engineering in Practice, pp
H. Aman, An empirical analysis on fault-proneness of well-commented modules, in: 2012 Fourth International Workshop on Empirical Soft- ware Engineering in Practice, pp. 3–9
2012
-
[44]
D. R. Cox, The regression analysis of binary sequences, Journal of the Royal Statistical Society. Series B (Methodological) 20 (1958) 215–242
1958
-
[45]
Breiman, J
L. Breiman, J. Friedman, C. Stone, R. Olshen, Classification and Re- gression Trees, The Wadsworth and Brooks-Cole statistics-probability series, Taylor and Francis, 1984
1984
-
[46]
Breiman, Random forests, Machine Learning 45 (2001) 5–32
L. Breiman, Random forests, Machine Learning 45 (2001) 5–32
2001
-
[47]
Geurts, D
P. Geurts, D. Ernst, L. Wehenkel, Extremely randomized trees, Ma- chine Learning 63 (2006) 3–42
2006
-
[48]
Breiman, Bagging predictors, Machine Learning 24 (1996) 123–140
L. Breiman, Bagging predictors, Machine Learning 24 (1996) 123–140
1996
-
[49]
Freund, R
Y. Freund, R. E. Schapire, A decision-theoretic generalization of on- line learning and an application to boosting, Journal of Computer and System Sciences 55 (1997) 119 – 139
1997
-
[50]
J. H. Friedman, Greedy function approximation: A gradient boosting machine., Ann. Statist. 29 (2001) 1189–1232
2001
-
[51]
T. Chen, C. Guestrin, Xgboost: A scalable tree boosting system, in: Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794
-
[52]
Y. Yu, H. Wang, V. Filkov, P. Devanbu, B. Vasilescu, Wait for it: Determinants of pull request evaluation latency on github, in: 12th Working Conference on Mining Software Repositories, pp. 367–371. 31
-
[53]
V. J. Hellendoorn, P. T. Devanbu, A. Bacchelli, Will they like this? evaluating code contributions with language models, in: 12th Working Conference on Mining Software Repositories, pp. 157–167
-
[54]
P. C. Rigby, M. Storey, Understanding broadcast based peer review on open source software projects, in: 33rd International Conference on Software Engineering (ICSE), pp. 541–550
-
[55]
Zampetti, G
F. Zampetti, G. Bavota, G. Canfora, M. Di Penta, A study on the interplay between pull request review and continuous integration builds, pp. 38–48
-
[56]
M. M. Rahman, C. K. Roy, J. A. Collins, Correct: Code reviewer recommendation in github based on cross-project and technology ex- perience, in: 38th International Conference on Software Engineering Companion (ICSE-C), pp. 222–231
-
[57]
J. Tsay, L. Dabbish, J. Herbsleb, Influence of social and technical factors for evaluating contribution in github, in: 36th International Conference on Software Engineering, ICSE 2014, pp. 356–366
2014
-
[58]
Runeson, M
P. Runeson, M. H¨ ost, Guidelines for conducting and reporting case study research in software engineering, Empirical Softw. Engg. 14 (2009) 131–164
2009
-
[59]
V. R. Basili, G. Caldiera, H. D. Rombach, The goal question metric approach, Encyclopedia of Software Engineering (1994)
1994
-
[60]
Patton, Qualitative Evaluation and Research Methods, Sage, New- bury Park, 2002
M. Patton, Qualitative Evaluation and Research Methods, Sage, New- bury Park, 2002
2002
-
[61]
Nagappan, T
M. Nagappan, T. Zimmermann, C. Bird, Diversity in software engi- neering research, ESEC/FSE 2013, pp. 466–476
2013
-
[62]
Kalliamvakou, G
E. Kalliamvakou, G. Gousios, K. Blincoe, L. Singer, D. M. German, D. Damian, An in-depth study of the promises and perils of mining github, Empirical Software Engineering 21 (2016) 2035–2071
2016
-
[63]
Powers, Evaluation: From precision, recall and f-factor to roc, in- formedness, markedness & correlation, Mach
D. Powers, Evaluation: From precision, recall and f-factor to roc, in- formedness, markedness & correlation, Mach. Learn. Technol. 2 (2008)
2008
-
[64]
A. P. Bradley, The use of the area under the roc curve in the evaluation of machine learning algorithms, Pattern Recognition 30 (1997) 1145 – 1159. 32
1997
-
[65]
Calefato, F
F. Calefato, F. Lanubile, B. Vasilescu, A large-scale, in-depth analysis of developers personalities in the apache ecosystem, Information and Software Technology 114 (2019) 1 – 20. 33
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.