REVIEW 3 major objections 4 minor 43 references
Enhancing Software Maintenance: A Learning to Rank Approach for Co-changed Method Identification
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A learning-to-rank model trained on pull-request histories ranks co-changed methods well enough to beat five baselines, reaching a mean NDCG@5 of 0.84 across 150 Java projects.
desk verdict A solid LtR application with a large dataset, undermined by an evaluation filter that drops all zero-label queries, so the headline NDCG is optimistic. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is a learning-to-rank pipeline that turns each query-method and candidate-method pair into a feature vector and trains a ranker over lists of candidates. The Random Forest model is the mechanism that carries the result; with features such as historical co-change count, author similarity, path similarity, code dependency, hierarchy similarity, clone similarity, and argument type and name similarity, and with labels equal to co-change counts in the next six months, it orders candidates by predicted relevance. The pull-request level is the unit of analysis, and NDCG@k is the evaluation device used to measure how well the top of the ranked list matches the true co-change labels.
What would settle it
Re-run evaluation without excluding zero-label queries: take every query method from a held-out period, rank all candidates, and compute NDCG@5; if the ranking on those queries is near chance, the reported performance does not extend to the arbitrary queries a developer would actually issue.
Extended reading notes
Core claim
The paper claims that co-change relationships between methods can be predicted and ranked by a learning-to-rank model. Each query method is paired with every other non-test method in the same repository; ten features describe the pair, from historical co-change count and shared authors to file-path similarity, code dependency, inheritance, clone similarity, and argument and signature similarity. The relevance label is the number of future pull-request co-changes in a six-month window. Trained on 150 open-source Java projects, a Random Forest model ranks candidates with a mean NDCG@5 of 0.84 and a median of 0.91, outperforming five baselines, with the historical co-change count by far the strongest feature.
Load-bearing premise
The claim relies on treating future co-change frequency in a fixed six-month window as the ground truth for relevance, and it removes any query whose methods never co-change in that window before measuring NDCG.
Editorial extensions
If this is right
- A developer editing one method can inspect a short ranked list of the methods most likely to need the same edit, reducing missed dependencies during maintenance.
- Pull-request-level analysis captures changes spread over multiple commits that commit-level co-change detection would miss.
- Ninety days of history is enough to train a usable model, and longer histories do not improve its performance significantly.
- Prediction quality degrades after about 60 days of unretrained use, so the model should be retrained roughly every two months.
- The best results appear on medium-size and younger Java projects; long-lived projects see significantly lower ranking quality.
Reading between the lines
- Editorial inference: since the co-change count feature dominates with permutation importance 0.38, much of the top-5 signal is the same signal a developer would get from counting past co-edits, and the added value of the machine-learning model lies mainly in path and author similarity.
- Editorial inference: the low importance of semantic and clone similarity suggests the ranking is not finding new conceptual couplings, so a testable next step is to fine-tune a code model on pull-request co-change data instead of using off-the-shelf cosine similarity.
- Editorial inference: a bi-monthly retraining rule follows directly from the reported 60-day performance decline, and it can be tested prospectively by comparing a model retrained every 60 days against one retrained every 180 days on live pull-request streams.
- Editorial inference: because labels are counted at pull-request level, the method may favour changes that developers deliberately bundle and could miss co-changes split across separate pull requests; measuring against commit-level labels would reveal how much of the gain comes from the pull-request unit itself.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a learning-to-rank (LtR) approach for identifying co-changed methods at the pull-request level. It extracts ten features from source-code structure and version-control history, labels method pairs by the number of future co-changes in a six-month window, and trains seven LtR models on 150 open-source Java projects. The Random Forest model is reported to achieve the best NDCG@5 (mean 0.84) and to outperform five baselines (support ranking, file proximity, code clone detection, FCP2Vec, and StarCoder 2). The paper also analyzes feature importance and studies how the training and testing label periods affect performance, concluding that models should be retrained roughly every two months.
Significance. If the reported evaluation is valid, the paper would be a useful empirical contribution: it is large-scale (150 projects, 634,216 pull requests), uses a non-overlapping temporal split, provides a replication package, and compares multiple ranking models and baselines. The permutation-importance analysis and the study of labeling periods are also constructive. However, two evaluation choices compromise the central quantitative claims: the exclusion of queries with no future co-changes in Algorithm 1, and the mismatched project sets for the StarCoder 2 baseline. These issues are fixable by re-analysis, so the underlying idea remains defensible, but the current numbers overstate the practical value of the tool.
major comments (3)
- [Section 3.5, Algorithm 1 (line 21)] The dataset creation excludes all ranking lists whose labels are all zero, with the comment that NDCG would be 0/0 and thus not meaningful. This is a selection on the outcome: only queries that are known to have at least one future co-change are scored. In real use, a developer can query any method, and many methods will have no co-change in the next six months; NDCG@5 for those queries should be defined (e.g., as 0) and included. Without that, the reported mean NDCG@5 of 0.84 (Section 4.1.3) and the margins in Table 5 are conditional on the query having at least one relevant item. Please re-run the evaluation on the full query set with an explicit convention for the undefined denominator, and report the fraction of queries excluded by the current filter.
- [Section 4.2.2, Table 5] StarCoder 2 was applied to only 45 randomly selected projects while the RF model was evaluated on all 150 projects. Table 5 compares aggregate means over different project sets, so the reported margin over StarCoder 2 confounds model quality with project selection. Please report a paired comparison on the same 45 projects, or run StarCoder 2 on all 150 projects, and state the number of projects underlying each row of Table 5.
- [Section 3.2, Section 4.3, Table 6] The most important feature, 'Number of co-changes,' is exactly the quantity used by the support-ranking baseline and is closely related to the label (future co-change count). The feature and label periods are sequential, so this is not direct label leakage, but the RQ2 comparison to support ranking (4.7% at NDCG@5) measures only the marginal value of the remaining features. Because the paper claims a learning-to-rank advantage, please add an ablation that removes the number-of-co-changes feature or uses only static features, and discuss how the 0.38 permutation importance bears on the interpretation of the RQ2 results. The threat stated in Section 6 about historical relevance biasing toward frequently modified methods applies directly here.
minor comments (4)
- [Section 3.2, Semantic similarity] The text says CodeBERT produces a vector with 764 scalar values, but the standard CodeBERT base model produces 768-dimensional embeddings; please verify and correct this number.
- [Abstract, Section 4.2.3, Conclusion] The margin over code-clone detection is reported as 537.5% in the abstract but as 573.5% in Section 4.2.3 and the conclusion; Table 5's NDCG@5 values imply 573.5%, so the abstract appears inconsistent.
- [Section 4.1.3] The sentence 'The RF model achieves the highest performance of 0.91 NDCG@5' should specify whether 0.91 is the median, a per-project maximum, or the mean; Table 4 lists the mean NDCG@5 as 0.8394 and the median as 0.9106.
- [Figure 6] The figure caption should state the axis labels explicitly, since the current text refers to 'days of testing data' but the figure itself is not self-contained.
Circularity Check
No significant circularity; the historical co-change feature is a legitimate time-split predictor and the evaluation compares against future labels.
full rationale
The derivation is self-contained. The only apparent near-circular element is that the 'Number of co-changes' feature (Table 3) is the same historical quantity used by the support-ranking baseline, and the label is a future co-change count. However, the paper constructs a temporal split: features are computed from repository creation to t_d, and labels are co-change occurrences between t_d and t_e (Section 3.4, Algorithm 1). The historical count is therefore a legitimate predictive feature for the future count, not the same quantity as the label. The support-ranking baseline uses the same historical count, so the 4.7% NDCG@5 advantage over that baseline is an empirical result of the model's combination of features, not an identity by construction. Algorithm 1's exclusion of ranking lists with all-zero labels is an evaluation-coverage threat and may inflate absolute NDCG, but it is not a circularity: it does not make the predicted ranking equivalent to an input. There are no load-bearing self-citations, imported uniqueness theorems, or renamed known results.
Assumptions & free parameters
free parameters (3)
- Historical labeling period =
180 days
- Spearman correlation threshold =
0.7
- Random Forest hyperparameters =
not reported
assumptions (4)
- domain assumption FinerGit accurately tracks method identities across commits.
- domain assumption Future co-change frequency is the appropriate ground truth for co-change recommendations.
- ad hoc to paper Queries with no co-changes in the test period can be excluded from evaluation.
- domain assumption CodeBERT embeddings and cosine similarity capture semantic similarity of methods.
Cite this review
Pith. "Pith review of Enhancing Software Maintenance: A Learning to Rank Approach for Co-changed Method Identification." pith.science (2026). https://pith.science/paper/ZVOF5PIN
@misc{pith2026241119099,
author = {Pith},
title = {Pith review of: Enhancing Software Maintenance: A Learning to Rank Approach for Co-changed Method Identification},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZVOF5PIN}},
note = {Machine review of arXiv:2411.19099}
}
read the original abstract
With the increasing complexity of large-scale software systems, identifying all necessary modifications for a specific change is challenging. Co-changed methods, which are methods frequently modified together, are crucial for understanding software dependencies. However, existing methods often produce large results with high false positives. Focusing on pull requests instead of individual commits provides a more comprehensive view of related changes, capturing essential co-change relationships. To address these challenges, we propose a learning-to-rank approach that combines source code features and change history to predict and rank co-changed methods at the pull-request level. Experiments on 150 open-source Java projects, totaling 41.5 million lines of code and 634,216 pull requests, show that the Random Forest model outperforms other models by 2.5 to 12.8 percent in NDCG@5. It also surpasses baselines such as file proximity, code clones, FCP2Vec, and StarCoder 2 by 4.7 to 537.5 percent. Models trained on longer historical data (90 to 180 days) perform consistently, while accuracy declines after 60 days, highlighting the need for bi-monthly retraining. This approach provides an effective tool for managing co-changed methods, enabling development teams to handle dependencies and maintain software quality.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Hamdi Abdurhman Ahmed and Jihwan Lee. 2023. FCP2Vec: Deep Learning-Based Approach to Software Change Prediction by Learning Co-Changing Patterns from Changelogs. Applied Sciences 13, 11 (2023). https://doi.org/10.3390/app13116453
-
[3]
Abdulkareem Alali, Brian Bartman, Christian D. Newman, and Jonathan I. Maletic. 2013. A preliminary investigation of using age and distance measures in the detection of evolutionary couplings. 2013 10th Working Conference on Mining Software Repositories (MSR) (2013). https://doi.org/10. 1109/msr.2013.6624024
arXiv 2013
-
[4]
Oscar Alejo, Juan M Fernández-Luna, Juan F Huete, and Ramiro Pérez-Vázquez. 2010. Direct optimization of evaluation measures in learning to rank using particle swarm. In 2010 Workshops on Database and Expert Systems Applications . IEEE, 42–46
work page 2010
-
[5]
L Breiman. 2001. Random Forests. Machine Learning 45 (10 2001), 5–32. https://doi.org/10.1023/A:1010950718922
-
[6]
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005. Learning to rank using gradient descent. In Proceedings of the 22nd international conference on Machine learning . 89–96
work page 2005
-
[7]
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007. Learning to Rank: From Pairwise Approach to Listwise Approach. Proceedings of the 24th International Conference on Machine Learning 227, 129–136. https://doi.org/10.1145/1273496.1273513
arXiv 2007
-
[8]
Vitor Cerqueira, Luis Torgo, and Igor Mozetič. 2020. Evaluating time series forecasting models: An empirical study on performance estimation methods. Machine Learning 109, 11 (2020), 1997–2028
work page 2020
-
[9]
James R Cordy and Chanchal K Roy. 2011. The NiCad clone detector. In 2011 IEEE 19th International Conference on Program Comprehension . IEEE, 219–220
work page 2011
Show all 43 references
- [10]
-
[11]
Jerome Friedman. 2000. Greedy Function Approximation: A Gradient Boosting Machine. The Annals of Statistics 29 (11 2000). https://doi.org/10. 1214/aos/1013203451
2000
-
[12]
H. Gall, K. Hajek, and M. Jazayeri. 1998. Detection of logical coupling based on product release history. In Proceedings. International Conference on Software Maintenance (Cat. No. 98CB36272) . 190–198. https://doi.org/10.1109/ICSM.1998.738508
1998
-
[13]
Francis Galton. 1886. Regression towards mediocrity in hereditary stature. The Journal of the Anthropological Institute of Great Britain and Ireland 15 (1886), 246–263
-
[14]
Yoshiki Higo, Shinpei Hayashi, and Shinji Kusumoto. 2020. On tracking Java methods with Git mechanisms. Journal of Systems and Software 165 (2020), 110571. https://doi.org/10.1016/j.jss.2020.110571
2020
-
[15]
Paul Jaccard. 1901. Etude de la distribution florale dans une portion des Alpes et du Jura. Bulletin de la Societe Vaudoise des Sciences Naturelles 37 (01 1901), 547–579. https://doi.org/10.5169/seals-266450
1901 doi
-
[16]
Kalervo Järvelin and Jaana Kekäläinen. 2000. IR evaluation methods for retrieving highly relevant documents. In Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Athens, Greece) (SIGIR ’00). Association for ...
2000
-
[17]
Huzefa Kagdi, Malcom Gethers, Denys Poshyvanyk, and Michael L. Collard. 2010. Blending Conceptual and Evolutionary Couplings to Support Change Impact Analysis in Source Code. In 2010 17th Working Conference on Reverse Engineering . 119–128. https://doi.org/10.1109/WCRE.2010.21
2010 doi
-
[18]
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Z...
2024 arXiv
-
[19]
Metzler and W.B
D. Metzler and W.B. Croft. 2007. Linear feature-based models for information retrieval. Inf. Retr. 16 (01 2007), 1–23
2007
-
[20]
Roy, and Kevin A
Manishankar Mondal, Banani Roy, Chanchal K. Roy, and Kevin A. Schneider. 2019. Ranking Co-Change Candidates of Micro-Clones. In Proceedings of the 29th Annual International Conference on Computer Science and Software Engineering (Toronto, Ontario, Canada) (CASCON ’19). IBM Cor...
2019
-
[21]
Roy, and Kevin A
Manishankar Mondal, Banani Roy, Chanchal K. Roy, and Kevin A. Schneider. 2020. HistoRank: History-Based Ranking of Co-change Candidates. In 2020 IEEE 27th International Conference on Software Analysis, Evolution and Reengineering (SANER) . 240–250. https://doi.org/10.1109/SANE...
2020
-
[22]
Roy, and Kevin A
Manishankar Mondal, Chanchal K. Roy, and Kevin A. Schneider. 2013. Improving the detection accuracy of evolutionary coupling. In 2013 21st International Conference on Program Comprehension (ICPC) . 223–226. https://doi.org/10.1109/ICPC.2013.6613853 Manuscript submitted to ACM ...
2013
-
[23]
Roy, and Kevin A
Manishankar Mondal, Chanchal K. Roy, and Kevin A. Schneider. 2014. Prediction and Ranking of Co-Change Candidates for Clones. In Proceedings of the 11th Working Conference on Mining Software Repositories (Hyderabad, India) (MSR 2014). Association for Computing Machinery, New Y...
2014
-
[24]
Md Nadim, Manishankar Mondal, Chanchal K Roy, and Kevin A Schneider. 2022. Evaluating the performance of clone detection tools in detecting cloned co-change candidates. Journal of Systems and Software 187 (2022), 111229
2022
-
[25]
Sleiman Rabah, Jiang Li, Mingzhi Liu, and Yuanwei Lai. 2010. Comparative Studies of 10 Programming Languages within 10 Diverse Criteria – a Team 7 COMP6411-S10 Term Report. arXiv:1009.0305 [cs.PL]
2010 arXiv
-
[26]
Thomas Rolfsnes, Stefano Di Alesio, Razieh Behjati, Leon Moonen, and Dave W. Binkley. 2016. Generalizing the Analysis of Evolutionary Coupling for Software Change Impact Analysis. In 2016 IEEE 23rd International Conference on Software Analysis, Evolution, and Reengineering (SA...
2016 doi
-
[27]
Abdullah Saydemir, Muhammed Esad Simitcioglu, and Hasan Sozer. 2021. On the Use of Evolutionary Coupling for Software Architecture Recovery. In 2021 15th Turkish National Software Engineering Symposium (UYMS) . 1–6. https://doi.org/10.1109/UYMS54260.2021.9659761
2021
-
[28]
Jeongju Sohn and Mike Papadakis. 2022. CEMENT: On the Use of Evolutionary Coupling Between Tests and Code Units. A Case Study on Fault Localization. In 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE) . 133–144. https://doi.org/10.1109/ISSRE55...
2022
-
[29]
Jeongju Sohn and Mike Papadakis. 2022. Using Evolutionary Coupling to Establish Relevance Links Between Tests and Code Units. A case study on fault localization. arXiv:2203.11343 [cs.SE]
2022 arXiv
-
[30]
Charles Spearman. 1961. The proof and measurement of association between two things. (1961)
1961
-
[31]
Jeffrey Svajlenko and Chanchal K. Roy. 2014. Evaluating Modern Clone Detection Tools. In2014 IEEE International Conference on Software Maintenance and Evolution. 321–330. https://doi.org/10.1109/ICSME.2014.54
2014 doi
-
[32]
Chakkrit Tantithamthavorn, Akinori Ihara, and Ken-Ichi Matsumoto. 2013. Using Co-change Histories to Improve Bug Localization Performance. In 2013 14th ACIS International Conference on Software Engineering, Artificial Intelligence, Networking and Parallel/Distributed Computing...
2013 doi
-
[33]
Hamed Valizadegan, Rong Jin, Ruofei Zhang, and Jianchang Mao. 2009. Learning to Rank by Optimizing NDCG Measure. In Advances in Neural Information Processing Systems , Y. Bengio, D. Schuurmans, J. Lafferty, C. Williams, and A. Culotta (Eds.), Vol. 22. Curran Associates, Inc. h...
2009
-
[34]
Feng Wang, Jinxiao Huang, and Yutao Ma. 2018. A Top-k Learning to Rank Approach to Cross-Project Software Defect Prediction. In 2018 25th Asia-Pacific Software Engineering Conference (APSEC) . 335–344. https://doi.org/10.1109/APSEC.2018.00048
2018
-
[35]
Yining Wang, Liwei Wang, Yuanzhi Li, Di He, and Tie-Yan Liu. 2013. A theoretical analysis of NDCG type ranking measures. In Conference on learning theory. PMLR, 25–54
2013
-
[36]
Yining Wang, Liwei Wang, Yuanzhi Li, Di He, Tie-Yan Liu, and Wei Chen. 2013. A Theoretical Analysis of NDCG Type Ranking Measures. arXiv:1304.6480 [cs.LG]
2013 arXiv
-
[37]
Frank Wilcoxon. 1945. Individual Comparisons by Ranking Methods. Biometrics Bulletin 1, 6 (1945), 80–83. http://www.jstor.org/stable/3001968
1945
-
[38]
Qiang Wu, Christopher Burges, Krysta Svore, and Jianfeng Gao. 2010. Adapting boosting for information retrieval measures. Inf. Retr. 13 (06 2010), 254–270. https://doi.org/10.1007/s10791-009-9112-1
2010 doi
-
[39]
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li. 2008. Listwise approach to learning to rank: theory and algorithm. InProceedings of the 25th international conference on Machine learning . 1192–1199
2008
-
[40]
Jun Xu and Hang Li. 2007. AdaRank: A Boosting Algorithm for Information Retrieval. In Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Amsterdam, The Netherlands) (SIGIR ’07). Association for Computing Mach...
2007
-
[41]
Yan Zheng, Zan Wang, Xiangyu Fan, Xiang Chen, and Zijiang Yang. 2018. Localizing multiple software faults based on evolution algorithm. Journal of Systems and Software 139 (2018), 107–123. https://doi.org/10.1016/j.jss.2018.02.001
2018 doi
-
[42]
Jiangang Zhu, Beijun Shen, and Fanghuai Hu. 2015. A Learning to Rank Framework for Developer Recommendation in Software Crowdsourcing. In 2015 Asia-Pacific Software Engineering Conference (APSEC) . 285–292. https://doi.org/10.1109/APSEC.2015.50
2015 doi
-
[43]
Zimmermann, P
T. Zimmermann, P. Weibgerber, S. Diehl, and A. Zeller. 2004. Mining version histories to guide software changes. In Proceedings. 26th International Conference on Software Engineering . 563–572. https://doi.org/10.1109/ICSE.2004.1317478
2004 arXiv
-
[44]
Thomas Zimmermann, Peter Weisgerber, Stephan Diehl, and Andreas Zeller. 2004. Mining Version Histories to Guide Software Changes. In Proceedings of the 26th International Conference on Software Engineering (ICSE ’04) . IEEE Computer Society, USA, 563–572. Received 20 February ...
2004
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.