REVIEW 2 major objections 5 minor 198 references
When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read This survey of 61 deep-learning-based IR bug localization studies argues that DL turns bug localization from lexical matching into semantic and syntactic understanding, easing lexical gap, code-structure blindness, and cold-start problems.
desk verdict Useful qualitative map of DL-based IRBL, but Table 9's impossible Top-k values sink the performance meta-analysis until fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery of the survey is its systematic study-selection procedure coupled with a classification scheme. Following established systematic-review guidelines, the authors built a 61-study pool from 440 initial hits using a PICO-derived search string, inclusion and exclusion criteria, forward and backward snowballing, and a quality checklist covering data, model, and evaluation criteria. Onto this pool they map a taxonomy of approaches - heterogeneous networks (separate encoders for bug reports and source code), homogeneous networks (one shared encoder), relevance-matching models (vector similarity), and other structures such as encoder-decoder and adversarial models - and a three-way categorization of code representation spanning token-based, syntactic-based, and semantic-based forms. This scheme is what turns a scattered literature into a meta-analysis: the same labels applied across all 61 studies produce the distribution counts, the representation statistics, and the compiled performance table from which the high-performance highlights are drawn.
What would settle it
Check the compiled performance table for internal consistency: Top-$k$ accuracy cannot decrease as $k$ grows, yet the FBL-BERT (changeset-file) row reports Top-1 = 29.3% and Top-5 = 13.8%, and the HMCBL (commit) row reports Top-1 = 28.4% and Top-5 = 16.0%; going back to the two original papers to see which value comes from which experimental setting would determine whether the survey's quantitative map of the field is reliable.
Extended reading notes
Core claim
The paper's central claim is that integrating deep learning into IR-based bug localization has moved the field from surface lexical matching to a third generation of approaches that extract semantic and syntactic information from both bug reports and source code. On this account, DL-based IRBL mitigates three long-standing problems: the lexical gap between how users describe bugs and how developers name code, the neglect of structural information in code (addressed through AST, CFG, and dependency-graph representations), and the cold-start problem in projects without enough bug-fixing history (addressed through cross-project and cross-language transfer). Analyzing 61 primary studies, the survey reports that heterogeneous networks, which use separate models for bug reports and source code, dominate the literature (33 of 61 studies), that file-level localization remains the dominant granularity, and that 18 approaches reach an average MAP of 50% or higher on the studied benchmarks, with five approaches highlighted as meeting all five reported performance thresholds at once. It also finds that evaluation remains concentrated on a few Java projects, that random cross-validation is giving way to chronological validation as the more accurate scheme, and that large-language-model-based localization is the field's clearest open frontier.
Load-bearing premise
The survey's quantitative overview assumes the performance numbers it extracted from the 61 papers faithfully match what those papers reported - and the paper itself admits the comparison is only rough - so a single misread or misaligned value would make the ranking of high-performance approaches unreliable.
Editorial extensions
If this is right
- Researchers entering the field can use the survey's taxonomy of model structures and code representations to position new work immediately, rather than re-deriving known designs from scattered papers.
- The compiled performance overview gives practitioners a shortlist: the five approaches that clear all reported thresholds become reasonable defaults for adoption or for baseline comparison in new evaluations.
- The survey's methodological comparison implies that future studies reporting only random k-fold validation will face scrutiny, because chronological validation avoids leaking future bug-fixing information into training.
- The gap analysis redirects the field's agenda toward non-Java languages, finer-grained localization at method, hunk, and changeset levels, and real-world industrial evaluation, where current evidence is thinnest.
- The observation that only a few studies use large language models, while several already use BERT, CodeBERT, and CodeSage families, sets LLM-based bug localization up as the next major test of the survey's core claim.
Reading between the lines
- The compiled table contains rows that violate the mathematical property that Top-$k$ accuracy is non-decreasing in $k$ (FBL-BERT changeset-file shows Top-1 = 29.3% but Top-5 = 13.8%; HMCBL commit shows Top-1 = 28.4% but Top-5 = 16.0%), which suggests those numbers were drawn from different experimental settings; a corrected extraction could change which approaches look high-performing.
- The survey's own preference for chronological validation predicts that approaches evaluated with random splits will look stronger than they would under temporal splits, since random splits can leak future bug-fixing information into training.
- Because most of the 61 studies share a small set of Java benchmarks, the performance map is best read as a ranking on those benchmarks; cross-language and LLM-based methods are still evaluated on too little common ground to rank fairly against each other.
- The recurring finding that mixing 20% of target-project data with source-project training outperforms pure transfer suggests a concrete next benchmark: a zero-label cold-start test in which no target-project data is allowed, which would sharpen the distinction between cold-start mitigated and cold-start solved.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a systematic literature survey of 61 primary studies (published up to November 2024) that apply deep learning (DL) to information retrieval-based bug localization (IRBL). It organizes the field along three research questions: RQ1 identifies the techniques, model structures, and text/code representations; RQ2 analyzes evaluation datasets, metrics, granularity, and validation approaches; RQ3 synthesizes challenges and open problems. The survey follows a Kitchenham-style protocol with search string, inclusion/exclusion criteria, snowballing, quality assessment, dual data extraction, and a meta-analysis of reported performance metrics in Table 9. The authors claim to be the first to provide a comprehensive survey dedicated to DL-based IRBL and derive several future directions, with particular attention to large language models.
Significance. If the survey were accurate, it would fill a genuine gap: previous IRBL surveys predate the DL wave, and this paper brings together 61 DL-based approaches with a structured map of model taxonomies, representations, evaluation practices, and open challenges. The systematic-review methodology (explicit protocol, independent extraction, quality assessment) is a strength, and the paper usefully documents the shift from relevance-based to semantic, structural, and graph-based code representations. The claimed contribution, however, rests heavily on the quantitative meta-analysis in Table 9 and Section 6.1. Because that table contains internally inconsistent values that violate the definition of Top-k ranking, the performance overview and the counts/highlights derived from it are not currently trustworthy. The paper's qualitative taxonomy and future-directions discussion remain useful, but the central quantitative claim needs to be corrected and made auditable before the survey can be relied upon.
major comments (2)
- [Table 9 / Section 6.1] Table 9 contains multiple rows in which the reported Top-1 value exceeds the Top-5 value for the same approach and granularity, e.g., FBL-BERT (changeset-file) Top-1=29.3 and Top-5=13.8; HMCBL (commit) Top-1=28.4 and Top-5=16.0; HMCBL (method) Top-1=37.2 and Top-5=20.2; Ciborowska et al. [29] Top-1=26.9 and Top-5=19.2; IBL(commit) Top-1=25.0 and Top-5=11.0. Top-k accuracy is, by definition, monotonically non-decreasing in k, so these values cannot originate from a single ranked list. This internal inconsistency indicates that the extracted metrics were averaged or concatenated across incompatible setups (e.g., different projects, granularities, or ranking definitions). Since Section 6.1 uses Table 9 to count approaches with MAP>=50%, Top-5>=70%, etc., and to bold 'high-performance' approaches, the quantitative performance overview and all derived conclusions are invalid as presented. The authors should re-extract and re-verify the per-study values, provide the underlying per-project data or a clear audit trail, and state how averages were computed across heterogeneous evaluation settings.
- [Section 2.4.2 / Section 6] The challenges taxonomy (RQ3) is derived explicitly from self-reported limitations and future-work sections of the primary studies ('Any discussion in a paper that explicitly mentioning a challenge or future work was extracted to the data extraction form'). This approach risks summarizing the authors' own claims rather than independently verifying them, and it may miss challenges that are observable from the survey's own cross-study analysis (e.g., the inconsistent evaluation protocols documented in Section 5.4 are themselves a challenge that few primary studies acknowledge). The paper should explicitly distinguish between challenges reported by the primary studies and challenges inferred by the survey authors, and should justify why the self-report-based taxonomy is considered complete.
minor comments (5)
- [Figure 2] The y-axis label '论文发表数量年份' is in Chinese; it should be translated to English (e.g., 'Number of publications by year') for a venue with an international readership.
- [Section 4.4] Equations (1) and (2) are both labeled as (1). The BugCache formula and the Recency formula need distinct equation numbers, and the in-text reference to 'Equotion (1)' should be corrected.
- [Introduction] The list of prior surveys cites '[2, 148, 152, 152]', with reference [152] duplicated. This suggests a missing or miscited reference and should be fixed.
- [Table 2] The caption states that 'we only present the mode of dataset sizes across all the studies', but the table mixes bug-report counts, file counts, method counts, and changeset counts without clear column semantics; a short explanation of what 'mode' means and which cells refer to which granularity would improve readability.
- [Section 5.1] The text says the search was conducted on November 30, 2024 and includes studies 'published up to that date', while Figure 5's timeline ends at '2024' without a month and the text elsewhere says 'published until November 2024' and 'from 2015 to October 2024'; these temporal statements should be made consistent.
Circularity Check
No circularity: the survey's taxonomy and performance overview are explicit aggregations of the 61 primary studies, not predictions derived from a fitted input.
full rationale
This is a systematic literature survey, not a derivation. Its output—a taxonomy of DL-based IRBL approaches, an inventory of representations/models/features/datasets, and a list of challenges—is explicitly synthesized from the 61 selected primary studies via data extraction (Section 2.4) and is not presented as an independently derived prediction. The closest thing to a quantitative claim, Table 9, is a direct aggregation of metrics reported by the primary studies; even if individual rows are internally inconsistent (e.g., Top-1 greater than Top-5), that is an extraction/accuracy threat, not a circular reduction of the survey's conclusions into its inputs. The challenge taxonomy is built from the primary studies' self-reported limitations and future work, but the survey does not claim this taxonomy was derived from first principles; it explicitly says 'Any discussion in a paper that explicitly mentioning a challenge or future work was extracted to the data extraction form' (Section 2.4.2), so no hidden equivalence is involved. The survey does include one self-citation: TRANP-CNN [56] is co-authored by one of the survey authors and is highlighted as the first deep transfer solution to cold-start. However, that citation points to a peer-reviewed TSE paper whose method and results are externally falsifiable, and the survey's overall map does not depend on any unverified assumption from that paper; hence it is not load-bearing circularity. No fitted parameter is renamed as a prediction, no ansatz is smuggled via citation, and no uniqueness claim is imported from the authors' prior work. Consequently the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The 61 selected primary studies are representative of DL-based IRBL research and their reported results are extracted accurately.
- domain assumption The search and snowballing process identified all relevant studies up to November 2024.
Cite this review
Pith. "Pith review of When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey." pith.science (2026). https://pith.science/paper/L35FRIUA
@misc{pith2026250500144,
author = {Pith},
title = {Pith review of: When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey},
year = {2026},
howpublished = {\url{https://pith.science/paper/L35FRIUA}},
note = {Machine review of arXiv:2505.00144}
}
read the original abstract
Bug localization is a crucial aspect of software maintenance, running through the entire software lifecycle. Information retrieval-based bug localization (IRBL) identifies buggy code based on bug reports, expediting the bug resolution process for developers. Recent years have witnessed significant achievements in IRBL, propelled by the widespread adoption of deep learning (DL). To provide a comprehensive overview of the current state of the art and delve into key issues, we conduct a survey encompassing 61 IRBL studies leveraging DL. We summarize best practices in each phase of the IRBL workflow, undertake a meta-analysis of prior studies, and suggest future research directions. This exploration aims to guide further advancements in the field, fostering a deeper understanding and refining practices for effective bug localization. Our study suggests that the integration of DL in IRBL enhances the model's capacity to extract semantic and syntactic information from both bug reports and source code, addressing issues such as lexical gaps, neglect of code structure information, and cold-start problems. Future research avenues for IRBL encompass exploring diversity in programming languages, adopting fine-grained granularity, and focusing on real-world applications. Most importantly, although some studies have started using large language models for IRBL, there is still a need for more in-depth exploration and thorough investigation in this area.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[152]
Xin Xia and David Lo. 2023. Information Retrieval-Based Techniques for Software Fault Localization. Handbook of Software Fault Localization: Foundations and Advances (2023), 365–391
work page 2023
-
[29]
Agnieszka Ciborowska and Kostadin Damevski. 2023. Too Few Bug Reports? Exploring Data Augmentation for Improved Changeset-based Bug Localization. arXiv:2305.16430 (2023)
arXiv 2023
-
[1]
Rui Abreu, Peter Zoeteweij, and Arjan JC Van Gemund. 2007. On the accuracy of spectrum-based fault localization. In TAICPART-MUTATION 2007. IEEE, 89–98
2007
-
[2]
Pragya Agarwal and Arun Prakash Agrawal. 2014. Fault-localization techniques for software systems: A literature review. ACM SIGSOFT Software Engineering Notes 39, 5 (2014), 1–8
2014
-
[3]
Hiralal Agrawal, Joseph R Horgan, Saul London, and W Eric Wong. 1995. Fault localization using execution slices and dataflow tests. In Proceedings of Sixth ISSRE . IEEE, 143–151
1995
-
[4]
Aminu A Ahmad, Lasheng Yu, Mohamed Kholief, and Abba Garba. 2023. AttentiveBugLocator: A Bug Localization Model using Attention-based SemanticFeatures and Information Retrieval. Research Square (2023)
2023
-
[5]
Shayan A Akbar and Avinash C Kak. 2020. A large-scale comparative evaluation of IR-based tools for bug localization. In Proceedings of the 17th MSR . 21–31
2020
-
[6]
Ahmed Sheikh Al-Aidaroos and Sara Mohammed Bamzahem. 2023. The Impact of GloVe and Word2Vec Word- Embedding Technologies on Bug Localization with Convolutional Neural Network. IJSEA (2023), 108–111
2023
Show all 198 references
-
[7]
Waqas Ali, Lili Bo, Xiaobing Sun, Xiaoxue Wu, Aakash Ali, and Ying Wei. 2024. Software bug localization based on optimized and ensembled deep learning models. Journal of Software: Evolution and Process 36, 8 (2024), e2654
2024
-
[8]
Waqas Ali, Lili Bo, Xiaobing Sun, Xiaoxue Wu, Saifullah Memon, Saima Siraj, and Ann Suwaree Ashton. 2023. Automated Software Bug Localization enabled by Meta-heuristic-based Convolutional Neural Network and Improved Deep Neural Network. Expert Systems with Applications (2023), 120562
2023
-
[9]
Shatha Alsaedi, Ahmed AA Gad-Elrab, Amin Noaman, and Fathy Eassa. 2024. Two-Level Information-Retrieval-Based Model for Bug Localization Based on Bug Reports. Electronics 13, 2 (2024), 321
2024
-
[10]
Bui Thi Mai Anh and Nguyen Viet Luyen. 2021. An Imbalanced Deep Learning Model for Bug Localization. In 2021 28th APSEC Workshops. IEEE, 32–40
2021
-
[11]
Le, David Lo, Claire Le Goues, and Lars Grunske
Tien-Duy B. Le, David Lo, Claire Le Goues, and Lars Grunske. 2016. A learning-to-rank based fault localization approach using likely invariants. In 25th ISSTA. 177–188
2016
-
[12]
Adrian Bachmann and Abraham Bernstein. 2009. Software process data quality and characteristics: a historical view on open and closed source projects. In Proceedings of the joint IWPSE and Evol workshops . 119–128
2009
-
[13]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. arXiv:1409.0473 (2014)
2014 arXiv
-
[14]
Anupam Baliyan, Akshit Batra, and Sunil Pratap Singh. 2021. Multilingual sentiment analysis using RNN-LSTM and neural machine translation. In 2021 8th INDIACom. IEEE, 710–713
2021
-
[15]
Samuel Benton, Ali Ghanbari, and Lingming Zhang. 2019. Defexts: A curated dataset of reproducible real-world bugs for modern jvm languages. In 2019 IEEE/ACM 41st ICSE-Companion. IEEE, 47–50
2019
-
[16]
Nicolas Bettenburg, Sascha Just, Adrian Schröter, Cathrin Weiß, Rahul Premraj, and Thomas Zimmermann. 2007. Quality of bug reports in eclipse. In Proceedings of the 2007 OOPSLA workshop on eclipse technology eXchange . 21–25
2007
-
[17]
Nicolas Bettenburg, Sascha Just, Adrian Schröter, Cathrin Weiss, Rahul Premraj, and Thomas Zimmermann. 2008. What makes a good bug report?. In Proceedings of the 16th ACM SIGSOFT FSE . 308–318
2008
-
[18]
Nicolas Bettenburg, Rahul Premraj, Thomas Zimmermann, and Sunghun Kim. 2008. Extracting structural information from bug reports. In Proceedings of the 2008 MSR . 27–30
2008
-
[19]
Kirti Bhandari, Kuldeep Kumar, and Amrit Lal Sangal. 2023. Data quality issues in software fault prediction: a systematic literature review. Artificial Intelligence Review 56, 8 (2023), 7839–7908
2023
-
[20]
Irina Ioana Brudaru and Andreas Zeller. 2008. What is the Long-Term Impact of Changes?. In RSSE’08 (Atlanta, Georgia). Association for Computing Machinery, New York, NY, USA, 30–32. https://doi.org/10.1145/1454247.1454257
2008
-
[21]
Junming Cao, Shouliang Yang, Wenhui Jiang, Hushuang Zeng, Beijun Shen, and Hao Zhong. 2020. Bugpecker: Locating faulty methods with deep learning on revision graphs. In Proceedings of the 35th IEEE/ACM ASE . 1214–1218
2020
-
[22]
Cagatay Catal, Görkem Giray, and Bedir Tekinerdogan. 2022. Applications of deep learning for mobile malware detection: A systematic literature review. Neural Computing and Applications (2022), 1–26
2022
-
[23]
Partha Chakraborty, Mahmoud Alfadel, and Meiyappan Nagappan. 2024. BLAZE: Cross-Language and Cross-Project Bug Localization via Dynamic Chunking and Hard Example Learning. arXiv:2407.17631 (2024)
2024 arXiv
-
[24]
Partha Chakraborty, Mahmoud Alfadel, and Meiyappan Nagappan. 2024. RLocator: Reinforcement learning for bug localization. IEEE TSE (2024)
2024
-
[25]
Mahinthan Chandramohan, Dai Quoc Nguyen, Padmanabhan Krishnan, and Jovan Jancic. 2024. Supporting Cross- language Cross-project Bug Localization Using Pre-trained Language Models. arXiv:2407.02732 (2024)
2024 arXiv
-
[26]
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer. 2002. SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16 (2002), 321–357
2002
-
[27]
Hao Chen, Haiyang Yang, Zilun Yan, Li Kuang, and Lingyan Zhang. 2022. CGMBL: Combining GAN and Method Name for Bug Localization. In 2022 IEEE 22nd QRS . IEEE, 231–241. , Vol. 1, No. 1, Article . Publication date: May 2025. 30 Niu et al
2022
-
[28]
Agnieszka Ciborowska and Kostadin Damevski. 2022. Fast changeset-based bug localization with BERT. InProceedings of the 44th ICSE . 946–957
2022
-
[30]
Roland Croft, Yongzheng Xie, and Muhammad Ali Babar. 2022. Data preparation for software vulnerability prediction: A systematic literature review. IEEE TSE 49, 3 (2022), 1044–1063
2022
-
[31]
Valentin Dallmeier and Thomas Zimmermann. 2007. Extraction of bug localization benchmarks from history. In Proceedings of the 22nd IEEE/ACM ASE . 433–436
2007
-
[32]
Steven Davies and Marc Roper. 2013. Bug localisation through diverse sources of information. In 2013 IEEE ISSREW. IEEE, 126–131
2013
-
[33]
Steven Davies, Marc Roper, and Murray Wood. 2012. Using bug report similarity to enhance bug localisation. In 2012 19th Working Conference on Reverse Engineering . IEEE, 125–134
2012
-
[34]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805 (2018)
2018 arXiv
-
[35]
Georgios Douzas, Fernando Bacao, and Felix Last. 2018. Improving imbalanced learning through a heuristic oversam- pling method based on k-means and SMOTE. Information Sciences 465 (2018), 1–20
2018
-
[36]
Yali Du and Zhongxing Yu. 2023. Pre-training Code Representation with Semantic Flow Graph for Effective Bug Localization. In Proceedings of the 31st ACM FSE . 579–591
2023
-
[37]
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al. 2020. Codebert: A pre-trained model for programming and natural languages. arXiv:2002.08155 (2020)
2020 arXiv
-
[38]
Joel Escudé Font and Marta R Costa-Jussa. 2019. Equalizing gender biases in neural machine translation with word embeddings techniques. arXiv:1901.03116 (2019)
2019 arXiv
-
[39]
Görkem Giray, Kwabena Ebo Bennin, Ömer Köksal, Önder Babur, and Bedir Tekinerdogan. 2023. On the use of deep learning in software defect prediction. Journal of Systems and Software 195 (2023), 111537
2023
-
[40]
Daya Guo, Shuai Lu, Nan Duan, Yanlin Wang, Ming Zhou, and Jian Yin. 2022. Unixcoder: Unified cross-modal pre-training for code representation. arXiv:2203.03850 (2022)
2022 arXiv
-
[41]
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie Liu, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, et al. 2020. Graphcodebert: Pre-training code representations with data flow. arXiv:2009.08366 (2020)
2020 arXiv
-
[42]
Tracy Hall, Sarah Beecham, David Bowes, David Gray, and Steve Counsell. 2011. A systematic literature review on fault prediction performance in software engineering. IEEE TSE 38, 6 (2011), 1276–1304
2011
-
[43]
Jiaxuan Han, Cheng Huang, and Jiayong Liu. 2024. bjEnet: a fast and accurate software bug localization method in natural language semantic space. Software Quality Journal (2024), 1–24
2024
-
[44]
Jiaxuan Han, Cheng Huang, Siqi Sun, Zhonglin Liu, and Jiayong Liu. 2023. bjXnet: an improved bug localization model based on code property graph and attention mechanism. Automated Software Engineering 30, 1 (2023), 12
2023
-
[45]
Mary Jean Harrold, Gregg Rothermel, Kent Sayre, Rui Wu, and Liu Yi. 2000. An empirical investigation of the relationship between spectra differences and regression faults. Software Testing, Verification and Reliability 10, 3 (2000), 171–194
2000
-
[46]
Haibo He, Yang Bai, Edwardo A Garcia, and Shutao Li. 2008. ADASYN: Adaptive synthetic sampling approach for imbalanced learning. In 2008 IEEE international joint conference on neural networks . Ieee, 1322–1328
2008
-
[47]
Kim Herzig, Sascha Just, and Andreas Zeller. 2013. It’s not a bug, it’s a feature: how misclassification impacts bug prediction. In 2013 35th ICSE. IEEE, 392–401
2013
-
[48]
Kim Herzig and Andreas Zeller. 2013. The impact of tangled code changes. In 2013 10th MSR. IEEE, 121–130
2013
-
[49]
S Hochreiter. 1997. Long Short-term Memory. Neural Computation MIT-Press (1997)
1997
-
[50]
Seyedrebvar Hosseini, Burak Turhan, and Dimuthu Gunarathna. 2017. A systematic literature review and meta- analysis on cross project defect prediction. IEEE TSE 45, 2 (2017), 111–147
2017
-
[51]
Xuxiang Huang, Chen Xiang, Hua Li, and Peng He. 2022. SBugLocater: Bug Localization Based on Deep Matching and Information Retrieval. Mathematical Problems in Engineering 2022 (2022)
2022
-
[52]
Bob Hunt, Bryn Turner, and Karen McRitchie. 2008. Software Maintenance Implications on Cost and Schedule. In 2008 IEEE Aerospace Conference. 1–6. https://doi.org/10.1109/AERO.2008.4526688
2008
-
[53]
Xuan Huo and Ming Li. 2017. Enhancing the Unified Features to Locate Buggy Files by Exploiting the Sequential Nature of Source Code.. In IJCAI. 1909–1915
2017
-
[54]
Xuan Huo, Ming Li, and Zhi-Hua Zhou. 2020. Control flow graph embedding based on multi-instance decomposition for bug localization. In AAAI, Vol. 34. 4223–4230
2020
-
[55]
Xuan Huo, Ming Li, Zhi-Hua Zhou, et al. 2016. Learning unified features from natural and programming languages for locating buggy source code.. In IJCAI, Vol. 16. 1606–1612. , Vol. 1, No. 1, Article . Publication date: May 2025. When Deep Learning Meets Information Retrieval-b...
2016
-
[56]
Xuan Huo, Ferdian Thung, Ming Li, David Lo, and Shu-Ting Shi. 2019. Deep transfer bug localization. IEEE TSE 47, 7 (2019), 1368–1380
2019
-
[57]
Shahid Iqbal, Rashid Naseem, Salman Jan, Sami Alshmrany, Muhammad Yasar, and Arshad Ali. 2020. Determining bug prioritization using feature reduction and clustering with classification. IEEE Access 8 (2020), 215661–215678
2020
-
[58]
Darryl Jarman, Jeffrey Berry, Riley Smith, Ferdian Thung, and David Lo. 2021. Legion: Massively composing rankers for improved bug localization at adobe. IEEE TSE 48, 8 (2021), 3010–3024
2021
-
[59]
Bo Jiang, Pengfei Liu, and Jie Xu. 2020. A deep learning approach to locate buggy files. In 11th DESSERT. IEEE
2020
-
[60]
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023. Swe-bench: Can language models resolve real-world github issues? arXiv:2310.06770 (2023)
2023 arXiv
-
[61]
René Just, Darioush Jalali, and Michael D Ernst. 2014. Defects4J: A database of existing faults to enable controlled testing studies for Java programs. In Proceedings of the 2014 ISSTA . 437–440
2014
-
[62]
Nal Kalchbrenner, Edward Grefenstette, and Phil Blunsom. 2014. A convolutional neural network for modelling sentences. arXiv:1404.2188 (2014)
2014 arXiv
-
[63]
Staffs Keele et al. 2007. Guidelines for performing systematic literature reviews in software engineering
2007
-
[64]
Sunghun Kim, Thomas Zimmermann, E James Whitehead Jr, and Andreas Zeller. 2007. Predicting faults from cached history. In 29th ICSE. IEEE, 489–498
2007
-
[65]
Thomas N Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv:1609.02907 (2016)
2016 arXiv
-
[66]
Barbara Kitchenham. 2004. Procedures for performing systematic reviews. Keele, UK, Keele University 33 (2004), 1–26
2004
-
[67]
Pavneet Singh Kochhar, Tien-Duy B Le, and David Lo. 2014. It’s not a bug, it’s a feature: Does misclassification affect bug localization?. In Proceedings of the 11th MSR . 296–299
2014
-
[68]
Pavneet Singh Kochhar, Yuan Tian, and David Lo. 2014. Potential biases in bug localization: Do they matter?. In Proceedings of the 29th ACM/IEEE ASE . 803–814
2014
-
[69]
Pavneet Singh Kochhar, Xin Xia, David Lo, and Shanping Li. 2016. Practitioners’ expectations on automated fault localization. In Proceedings of the 25th ISSTA . 165–176
2016
-
[70]
Sotiris Kotsiantis, Dimitris Kanellopoulos, Panayiotis Pintelas, et al. 2006. Handling imbalanced datasets: A review. GESTS international transactions on computer science and engineering 30, 1 (2006), 25–36
2006
-
[71]
Adrian Kuhn, Stéphane Ducasse, and Tudor Gîrba. 2007. Semantic clustering: Identifying topics in source code. IST 49, 3 (2007), 230–243
2007
-
[72]
An Ngoc Lam, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N Nguyen. 2015. Combining deep learning with information retrieval to localize buggy files for bug reports (n). In 2015 30th IEEE/ACM ASE. IEEE, 476–481
2015
-
[73]
An Ngoc Lam, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N Nguyen. 2017. Bug localization with combination of deep learning and information retrieval. In 2017 IEEE/ACM 25th ICPC. IEEE, 218–229
2017
-
[74]
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019. Albert: A lite bert for self-supervised learning of language representations. arXiv:1909.11942 (2019)
2019 arXiv
-
[75]
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. nature 521, 7553 (2015), 436–444
2015
-
[76]
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998. Gradient-based learning applied to document recognition. Proc. IEEE 86, 11 (1998), 2278–2324
1998
-
[77]
Jaekwon Lee, Dongsun Kim, Tegawendé F Bissyandé, Woosung Jung, and Yves Le Traon. 2018. Bench4bl: repro- ducibility study on the performance of ir-based bug localization. In 27th ISSTA. 61–72
2018
-
[78]
Chris Lewis, Zhongpeng Lin, Caitlin Sadowski, Xiaoyan Zhu, Rong Ou, and E James Whitehead. 2013. Does bug prediction support human developers? findings from a google case study. In 2013 35th ICSE. IEEE, 372–381
2013
-
[79]
Wei Li, Qingan Li, Yunlong Ming, Weijiao Dai, Shi Ying, and Mengting Yuan. 2022. An empirical study of the effectiveness of IR-based bug localization for large-scale industrial projects. EMSE 27, 2 (2022), 47
2022
-
[80]
Xia Li, Wei Li, Yuqun Zhang, and Lingming Zhang. 2019. Deepfl: Integrating multiple fault diagnosis dimensions for deep fault localization. In 28th ISSTA. 169–180
2019
-
[81]
Yujia Li, Daniel Tarlow, Marc Brockschmidt, and Richard Zemel. 2015. Gated graph sequence neural networks. arXiv:1511.05493 (2015)
2015 arXiv
-
[82]
Hongliang Liang, Dengji Hang, and Xiangyu Li. 2022. Modeling function-level interactions for file-level bug localization. EMSE 27, 7 (2022), 186
2022
-
[83]
Hongliang Liang, Lu Sun, Meilin Wang, and Yuxing Yang. 2019. Deep learning with customized abstract syntax tree for bug localization. IEEE Access 7 (2019), 116309–116320
2019
-
[84]
Guangliang Liu, Yang Lu, Ke Shi, Jingfei Chang, and Xing Wei. 2019. Convolutional neural networks-based locating relevant buggy code files for bug reports affected by data imbalance. IEEE Access 7 (2019), 131304–131316
2019
-
[85]
Pablo Loyola, Kugamoorthy Gajananan, and Fumiko Satoh. 2018. Bug localization by learning to rank and represent bug inducing changes. In Proceedings of the 27th ACM CIKM . 657–665. , Vol. 1, No. 1, Article . Publication date: May 2025. 32 Niu et al
2018
-
[86]
Stacy K Lukins, Nicholas A Kraft, and Letha H Etzkorn. 2008. Source code retrieval for bug localization using latent dirichlet allocation. In 2008 15Th working conference on reverse engineering . IEEE, 155–164
2008
-
[87]
Zhengmao Luo, Wenyao Wang, and Caichun Cen. 2022. Improving Bug Localization With Effective Contrastive Learning Representation. IEEE Access 11 (2022), 32523–32533
2022
-
[88]
Yi-Fan Ma, Yali Du, and Ming Li. 2023. Capturing the long-distance dependency in the control flow graph via structural-guided attention for bug localization. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI
2023
-
[89]
Yi-Fan Ma and Ming Li. 2022. The flowing nature matters: feature learning from the control flow graph of source code for bug localization. Machine Learning 111, 3 (2022), 853–870
2022
-
[90]
Yi-Fan Ma and Ming Li. 2022. Learning from the Multi-Level Abstraction of the Control Flow Graph via Alternating Propagation for Bug Localization. In 2022 IEEE ICDM. IEEE, 299–308
2022
-
[91]
Christopher D Manning, Prabhakar Raghavan, and Hinrich Schütze. 2008. Introduction to information retrieval . Cambridge university press
2008
-
[92]
Andrian Marcus, Andrey Sergeyev, Vaclav Rajlich, and Jonathan I Maletic. 2004. An information retrieval approach to concept location in source code. In 11th working conference on reverse engineering . IEEE, 214–223
2004
-
[93]
Matias Martinez and Martin Monperrus. 2016. Astor: A program repair library for java. In 25th ISSTA. 441–444
2016
-
[94]
E Peters Matthew, N Mark, I Mohit, G Matt, C Christopher, and L Kenton. 1802. Deep contextualized word represen- tations (2018). arXiv:1802.05365 (1802)
2018 arXiv
-
[95]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv:1301.3781 (2013)
2013 arXiv
-
[96]
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. Advances in neural information processing systems 26 (2013)
2013
-
[97]
Amr Mansour Mohsen, Hesham Hassan, Khaled Wassif, Ramadan Moawad, and Soha Makady. 2023. Enhancing Bug Localization using Phase-Based Approach. IEEE Access (2023)
2023
-
[98]
Seokhyeon Moon, Yunho Kim, Moonzoo Kim, and Shin Yoo. 2014. Ask the mutants: Mutating faulty programs for fault localization. In 2014 IEEE Seventh ICST . IEEE, 153–162
2014
-
[99]
Laura Moreno, John Joseph Treadway, Andrian Marcus, and Wuwei Shen. 2014. On the use of stack traces to improve text retrieval-based bug localization. In 2014 IEEE ICSME. IEEE, 151–160
2014
-
[100]
Alejandro Moreo, Andrea Esuli, and Fabrizio Sebastiani. 2021. Word-class embeddings for multiclass text classification. Data Mining and Knowledge Discovery 35 (2021), 911–963
2021
-
[101]
Vijayaraghavan Murali, Lee Gross, Rebecca Qian, and Satish Chandra. 2021. Industry-scale ir-based bug localization: A perspective from facebook. In 2021 IEEE/ACM 43rd ICSE-SEIP. IEEE, 188–197
2021
-
[102]
Andrew Ng. 2024. UNBIGGEN AI. https://spectrum.ieee.org/andrew-ng-data-centric-ai
2024
-
[103]
Chao Ni, Wei Wang, Kaiwen Yang, Xin Xia, Kui Liu, and David Lo. 2022. The best of both worlds: integrating semantic features with expert features for defect prediction and localization. In Proceedings of the 30th ACM FSE . 672–683
2022
-
[104]
Brent D Nichols. 2010. Augmented bug localization using past bug information. In Proceedings of the 48th Annual Southeast Regional Conference. 1–6
2010
-
[105]
Feifei Niu, Wesley KG Assunçao, LiGuo Huang, Christoph Mayr-Dorn, Jidong Ge, Bin Luo, and Alexander Egyed
-
[106]
Feifei Niu, Christoph Mayr-Dorn, Wesley KG Assunção, LiGuo Huang, Jidong Ge, Bin Luo, and Alexander Egyed
-
[107]
Feifei Niu, Enshuo Zhang, Christoph Mayr-Dorn, Wesley Klewerton Guez Assunção, Liguo Huang, Jidong Ge, Bin Luo, and Alexander Egyed. 2024. An extensive replication study of the ABLoTS approach for bug localization. EMSE 29, 6 (2024), 1–37
2024
-
[108]
In 20th MSR
The ABLoTS Approach for Bug Localization: is it replicable and generalizable?. In 20th MSR. IEEE, 576–587
-
[109]
Mike Papadakis and Yves Le Traon. 2015. Metallaxis-FL: mutation-based fault localization.Software Testing, Verification and Reliability 25, 5-7 (2015), 605–628
2015
-
[110]
Matteo Pagliardini, Prakhar Gupta, and Martin Jaggi. 2017. Unsupervised learning of sentence embeddings using compositional n-gram features. arXiv:1703.02507 (2017)
2017 arXiv
-
[111]
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In Proceedings of the 2014 EMNLP . 1532–1543
2014
-
[112]
2019.{TESSERACT}: Eliminating experimental bias in malware classification across space and time
Feargus Pendlebury, Fabio Pierazzi, Roberto Jordaney, Johannes Kinder, and Lorenzo Cavallaro. 2019.{TESSERACT}: Eliminating experimental bias in malware classification across space and time. In 28th USENIX Security. 729–746
2019
-
[113]
Binhang Qi, Hailong Sun, Wei Yuan, Hongyu Zhang, and Xiangxin Meng. 2021. Dreamloc: A deep relevance matching-based framework for bug localization. IEEE Transactions on Reliability 71, 1 (2021), 235–249. , Vol. 1, No. 1, Article . Publication date: May 2025. When Deep Learning...
2021
-
[114]
Sravya Polisetty, Andriy Miranskyy, and Ayşe Başar. 2019. On usefulness of the deep-learning-based bug localization models to practitioners. In Proceedings of the Fifteenth International Conference on Predictive Models and Data Analytics in Software Engineering. 16–25
2019
-
[115]
Shivani Rao and Avinash Kak. 2011. Retrieval from software libraries for bug localization: a comparative study of generic and composite text models. In Proceedings of the 8th MSR . 43–52
2011
-
[116]
Foyzur Rahman, Daryl Posnett, Abram Hindle, Earl Barr, and Premkumar Devanbu. 2011. BugCache for inspections: hit or miss?. In FSE. 322–331
2011
-
[117]
Nils Reimers. 2023. SentenceTransformer. https://huggingface.co/sentence-transformers
2023
-
[118]
Michael Rath, David Lo, and Patrick Mäder. 2018. Analyzing requirements and traceability information to improve bug localization. In Proceedings of the 15th MSR . 442–453
2018
-
[119]
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986. Learning representations by back-propagating errors. nature 323, 6088 (1986), 533–536
1986
-
[120]
Manos Renieres and Steven P Reiss. 2003. Fault localization with nearest neighbor queries. In 18th IEEE ASE. IEEE, 30–39
2003
-
[121]
Ripon K Saha, Matthew Lease, Sarfraz Khurshid, and Dewayne E Perry. 2013. Improving bug localization using structured information retrieval. In 28th IEEE/ACM ASE. IEEE, 345–355
2013
-
[122]
Ripon K Saha, Julia Lawall, Sarfraz Khurshid, and Dewayne E Perry. 2014. On the effectiveness of information retrieval based bug localization for c programs. In 2014 IEEE ICSME. IEEE, 161–170
2014
-
[123]
Shubham Sangle, Sandeep Muvva, Sridhar Chimalakonda, Karthikeyan Ponnalagu, and Vijendran Gopalan Venkoparao
-
[124]
Gerard Salton. 1989. Automatic text processing: The transformation, analysis, and retrieval of.Reading: Addison-Wesley 169 (1989)
1989
-
[125]
Hinrich Schütze, Christopher D Manning, and Prabhakar Raghavan. 2008. Introduction to information retrieval. Vol. 39. Cambridge University Press Cambridge
2008
-
[126]
S Selva Birunda and R Kanniga Devi. 2021. A review on word embedding techniques for text classification. Innovative Data Communication Technologies and Application: Proceedings of ICIDCA 2020 (2021), 267–281
2021
-
[127]
Adrian Schroter, Adrian Schröter, Nicolas Bettenburg, and Rahul Premraj. 2010. Do stack traces help developers fix bugs?. In 2010 7th IEEE MSR . IEEE, 118–121
2010
-
[128]
Bunyamin Sisman and Avinash C Kak. 2012. Incorporating version histories in information retrieval based bug localization. In 2012 9th IEEE MSR . IEEE, 50–59
2012
-
[129]
Jacek Śliwerski, Thomas Zimmermann, and Andreas Zeller. 2005. When do changes induce fixes? ACM sigsoft software engineering notes 30, 4 (2005), 1–5
2005
-
[130]
Xiangyu Shi, Xiaolin Ju, Xiang Chen, Guilong Lu, and Mengqi Xu. 2022. SemirFL: Boosting Fault Localization via Combining Semantic Information and Information Retrieval. In 2022 IEEE 22nd QRS-C . IEEE, 324–332
2022
-
[131]
Roger Alan Stein, Patricia A Jaques, and Joao Francisco Valiati. 2019. An analysis of hierarchical text classification using word embeddings. Information Sciences 471 (2019), 216–232
2019
-
[132]
Jeniya Tabassum, Mounica Maddela, Wei Xu, and Alan Ritter. 2020. Code and named entity recognition in stackover- flow. arXiv:2005.01634 (2020)
2020 arXiv
-
[133]
Riad Sonbol, Ghaida Rebdawi, and Nada Ghneim. 2022. The use of nlp-based text representation techniques to support requirement engineering tasks: A systematic mapping review. IEEE Access (2022)
2022
-
[134]
Shizuka Tsumita, Shinpei Hayashi, and Sousuke Amasaki. 2023. Large-Scale Evaluation of Method-Level Bug Localization with FinerBench4BL. In 2023 IEEE SANER. IEEE, 815–824
2023
-
[135]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[136]
Wei Tao, Yucheng Zhou, Yanlin Wang, Wenqiang Zhang, Hongyu Zhang, and Yu Cheng. 2024. Magis: Llm-based multi-agent framework for github issue resolution. arXiv:2403.17927 (2024)
2024 arXiv
-
[137]
Bei Wang, Ling Xu, Meng Yan, Chao Liu, and Ling Liu. 2020. Multi-dimension convolutional neural network for bug localization. IEEE TSC 15, 3 (2020), 1649–1663
2020
-
[138]
Qianqian Wang, Chris Parnin, and Alessandro Orso. 2015. Evaluating the usefulness of ir-based fault localization techniques. In Proceedings of the 2015 ISSTA . 1–11
2015
-
[139]
Ellen M Voorhees et al. 1999. The trec-8 question answering track report.. In Trec, Vol. 99. 77–82
1999
-
[140]
Shaowei Wang and David Lo. 2016. Amalgam+: Composing rich information sources for accurate bug localization. Journal of Software: Evolution and Process 28, 10 (2016), 921–942
2016
-
[141]
Xiaoyin Wang, Lu Zhang, Tao Xie, John Anvik, and Jiasu Sun. 2008. An approach to detecting duplicate bug reports using natural language and execution information. In Proceedings of the 30th ICSE . 461–470
2008
-
[142]
Shaowei Wang and David Lo. 2014. Version history, similar report, and structure: Putting them together for improved bug localization. In Proceedings of the 22nd ICPC . 53–63
2014
-
[143]
Ming Wen, Rongxin Wu, and Shing-Chi Cheung. 2016. Locus: Locating bugs from software changes. In Proceedings of the 31st IEEE/ACM ASE. 262–273
2016
-
[144]
Ratnadira Widyasari, Stefanus Agus Haryono, Ferdian Thung, Jieke Shi, Constance Tan, Fiona Wee, Jack Phan, and David Lo. 2022. On the influence of biases in bug localization: Evaluation and benchmark. In IEEE SANER. 128–139
2022
-
[145]
Yue Wang, Hung Le, Akhilesh Deepak Gotmare, Nghi DQ Bui, Junnan Li, and Steven CH Hoi. 2023. Codet5+: Open code large language models for code understanding and generation. arXiv:2305.07922 (2023). , Vol. 1, No. 1, Article . Publication date: May 2025. 34 Niu et al
2023 arXiv
-
[146]
Claes Wohlin. 2014. Guidelines for snowballing in systematic literature studies and a replication in software engineering. In Proceedings of the 18th EASE . 1–10
2014
-
[147]
Chu-Pan Wong, Yingfei Xiong, Hongyu Zhang, Dan Hao, Lu Zhang, and Hong Mei. 2014. Boosting bug-report-oriented fault localization with segmentation and stack-trace analysis. In 2014 IEEE ICSME. IEEE, 181–190
2014
-
[148]
Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok, Haodi Qi, Jack Phan, Qijin Tay, Constance Tan, Fiona Wee, Jodie Ethelda Tan, Yuheng Yieh, et al. 2020. Bugsinpy: a database of existing bugs in python programs to enable controlled testing and debugging studies. In Proceedings ...
2020
-
[149]
W Eric Wong, Ruizhi Gao, Yihao Li, Rui Abreu, Franz Wotawa, and Dongcheng Li. 2023. Software fault localization: An overview of research, techniques, and tools. Handbook of Software Fault Localization: Foundations and Advances (2023), 1–117
2023
-
[150]
Rongxin Wu, Hongyu Zhang, Shing-Chi Cheung, and Sunghun Kim. 2014. Crashlocator: Locating crashing faults based on crash stacks. In Proceedings of the 2014 ISSTA . 204–214
2014
-
[151]
W Eric Wong, Ruizhi Gao, Yihao Li, Rui Abreu, and Franz Wotawa. 2016. A survey on software fault localization. IEEE TSE 42, 8 (2016), 707–740
2016
-
[153]
Xin Xia, David Lo, Xingen Wang, Chenyi Zhang, and Xinyu Wang. 2014. Cross-language bug localization. In Proceedings of the 22nd ICPC . 275–278
2014
-
[154]
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin. 2018. Unsupervised feature learning via non-parametric instance discrimination. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3733–3742
2018
-
[155]
Yan Xiao and Jacky Keung. 2018. Improving bug localization with character-level convolutional neural network and recurrent neural network. In 2018 25th APSEC. IEEE, 703–704
2018
-
[156]
Yan Xiao, Jacky Keung, Kwabena E Bennin, and Qing Mi. 2018. Machine translation-based bug localization technique for bridging lexical gap. IST 99 (2018), 58–61
2018
-
[157]
Xi Xiao, Renjie Xiao, Qing Li, Jianhui Lv, Shunyan Cui, and Qixu Liu. 2023. BugRadar: Bug localization by knowledge graph link prediction. IST (2023), 107274
2023
-
[158]
Yan Xiao, Jacky Keung, Qing Mi, and Kwabena E Bennin. 2017. Improving bug localization with an enhanced convolutional neural network. In 2017 24th APSEC. IEEE, 338–347
2017
-
[159]
Yan Xiao, Jacky Keung, Qing Mi, and Kwabena E Bennin. 2018. Bug localization with semantic and structural features using convolutional neural network and cascade forest. In Proceedings of the 22nd EASE . 101–111
2018
-
[160]
Yan Xiao, Jacky Keung, Kwabena E Bennin, and Qing Mi. 2019. Improving bug localization with word embedding and enhanced convolutional neural networks. IST 105 (2019), 17–29
2019
-
[161]
Guoqing Xu, Xingqi Wang, Dan Wei, Yanli Shao, and Bin Chen. 2023. Bug Localization with Features Crossing and Structured Semantic Information Matching. SEKE (2023)
2023
-
[162]
Fabian Yamaguchi, Nico Golde, Daniel Arp, and Konrad Rieck. 2014. Modeling and discovering vulnerabilities with code property graphs. In 2014 IEEE Symposium on Security and Privacy . IEEE, 590–604
2014
-
[163]
Xiaoyuan Xie, Tsong Yueh Chen, Fei-Ching Kuo, and Baowen Xu. 2013. A theoretical analysis of the risk evaluation formulas for spectrum-based fault localization. ACM TOSEM 22, 4 (2013), 1–40
2013
-
[164]
Geunseok Yang and Byungjeong Lee. 2021. Utilizing topic-based similar commit information and CNN-LSTM algorithm for bug localization. Symmetry 13, 3 (2021), 406
2021
-
[165]
Geunseok Yang, Kyeongsic Min, and Byungjeong Lee. 2020. Applying deep learning algorithm to automatic bug localization and repair. In Proceedings of the 35th Annual ACM symposium on applied computing . 1634–1641
2020
-
[166]
Xuefeng Yan, Shasha Cheng, and Liqin Guo. 2023. Bug localization based on syntactical and semantic information of source code. Journal of Systems Engineering and Electronics 34, 1 (2023), 236–246
2023
-
[167]
Xin Ye, Razvan Bunescu, and Chang Liu. 2014. Learning to rank relevant files for bug reports using domain knowledge. In Proceedings of the 22nd ACM SIGSOFT FSE . 689–699
2014
-
[168]
Xin Ye, Razvan Bunescu, and Chang Liu. 2015. Mapping bug reports to relevant files: A ranking model, a fine-grained benchmark, and feature evaluation. IEEE TSE 42, 4 (2015), 379–402
2015
-
[169]
Shouliang Yang, Junming Cao, Hushuang Zeng, Beijun Shen, and Hao Zhong. 2021. Locating faulty methods with a mixed RNN and attention model. In 2021 IEEE/ACM 29th ICPC. IEEE, 207–218
2021
-
[170]
Klaus Changsun Youm, June Ahn, Jeongho Kim, and Eunseok Lee. 2015. Bug localization based on code change histories and bug reports. In 2015 APSEC. IEEE, 190–197
2015
-
[171]
Liang-Chih Yu, Jin Wang, K Robert Lai, and Xuejie Zhang. 2017. Refining word embeddings for sentiment analysis. In Proceedings of the 2017 EMNLP . 534–539
2017
-
[172]
Jian Yong, Ziye Zhu, and Yun Li. 2023. Decomposing Source Codes by Program Slicing for Bug Localization. In 2023 IJCNN. IEEE, 1–8. , Vol. 1, No. 1, Article . Publication date: May 2025. When Deep Learning Meets Information Retrieval-based Bug Localization: A Survey 35
2023
-
[173]
Wei Yuan, Binhang Qi, Hailong Sun, and Xudong Liu. 2020. Dependloc: A dependency-based framework for bug localization. In 2020 27th APSEC. IEEE, 61–70
2020
-
[174]
Abubakar Zakari, Sai Peck Lee, Rui Abreu, Babiker Hussien Ahmed, and Rasheed Abubakar Rasheed. 2020. Multiple fault localization of software programs: A systematic literature review. IST 124 (2020), 106312
2020
-
[175]
Xiao Yu, Zexian Zhang, Feifei Niu, Xing Hu, Xin Xia, and John Grundy. 2024. What Makes a High-Quality Training Dataset for Large Language Models: A Practitioners’ Perspective. In Proceedings of the 39th IEEE/ACM ASE . 656–668
2024
-
[176]
Filip Zamfirov. 2022. A literature review on different types of empirically evaluated bug localization approaches. arXiv:2212.11774 (2022)
2022 arXiv
-
[177]
Zhengran Zeng, Yuqun Zhang, Haotian Zhang, and Lingming Zhang. 2021. Deep just-in-time defect prediction: how far are we?. In 30th ISSTA. 427–438
2021
-
[178]
Abubakar Zakari, Sai Peck Lee, Khubaib Amjad Alam, and Rodina Ahmad. 2019. Software fault localisation: a systematic mapping study. IET Software 13, 1 (2019), 60–74
2019
-
[179]
He Zhang, Muhammad Ali Babar, and Paolo Tell. 2011. Identifying relevant studies in software engineering. IST 53, 6 (2011), 625–637
2011
-
[180]
Jinglei Zhang, Rui Xie, Wei Ye, Yuhan Zhang, and Shikun Zhang. 2020. Exploiting code knowledge graph for bug localization via bi-directional attention. In Proceedings of the 28th ICPC . 219–229
2020
-
[181]
Dejiao Zhang, Wasi Uddin Ahmad, Ming Tan, Hantian Ding, Ramesh Nallapati, Dan Roth, Xiaofei Ma, and Bing Xiang. 2024. CODE REPRESENTATION LEARNING AT SCALE. In The Twelfth ICLR
2024
-
[182]
Xiangyu Zhang, Neelam Gupta, and Rajiv Gupta. 2006. Locating faults through automated predicate switching. In Proceedings of the 28th ICSE . 272–281
2006
-
[183]
Xia Zhang, Ziye Zhu, and Yun Li. 2023. Enhancing Bug Localization through Bug Report Summarization. In 2023 IEEE ICDM. IEEE, 1541–1546
2023
-
[184]
Lisa Zhang, Zhe Kang, Xiaoxin Sun, Hong Sun, Bangzuo Zhang, and Dongbing Pu. 2021. KCRec: Knowledge-aware representation graph convolutional network for recommendation. Knowledge-Based Systems 230 (2021), 107399
2021
-
[185]
Chunying Zhou, Xiaoyuan Xie, Gong Chen, Peng He, and Bing Li. 2024. Multi-View Adaptive Contrastive Learning for Information Retrieval Based Fault Localization. arXiv:2409.12519 (2024)
2024
-
[186]
Jian Zhou, Hongyu Zhang, and David Lo. 2012. Where should the bugs be fixed? more accurate information retrieval-based bug localization based on bug reports. In 2012 34th ICSE. IEEE, 14–24
2012
-
[187]
Yaqiang Zhao, Xiaozhuo Li, Wei Deng, Ying Li, Xiaobo Guo, Qing Tian, and Ying Fan. 2024. Fine-Grained Bug Localization Based on Rich Context using Attention Tree-GRU. In 2024 5th ICCEA. IEEE, 640–646
2024
-
[188]
Ziye Zhu, Yun Li, Yu Wang, Yaojing Wang, and Hanghang Tong. 2021. A deep multimodal model for bug localization. Data Mining and Knowledge Discovery 35, 4 (2021), 1369–1392
2021
-
[189]
Ziye Zhu, Hanghang Tong, Yu Wang, and Yun Li. 2022. BL-GAN: Semi-Supervised Bug Localization Via Generative Adversarial Network. IEEE TKDE (2022)
2022
-
[190]
Ziye Zhu, Yun Li, Hanghang Tong, and Yu Wang. 2020. Cooba: Cross-project bug localization via adversarial transfer learning. In IJCAI
2020
-
[191]
Ziye Zhu, Yu Wang, and Yun Li. 2021. TroBo: A Novel Deep Transfer Model for Enhancing Cross-Project Bug Localization. In International Conference on Knowledge Science, Engineering and Management . Springer, 529–541
2021
-
[192]
Thomas Zimmermann, Nachiappan Nagappan, Harald Gall, Emanuel Giger, and Brendan Murphy. 2009. Cross-project defect prediction: a large scale experiment on data vs. domain vs. process. In FSE. 91–100
2009
-
[193]
Ziye Zhu, Hanghang Tong, Yu Wang, and Yun Li. 2022. Enhancing bug localization with bug report decomposition and code hierarchical network. Knowledge-Based Systems 248 (2022), 108741
2022
-
[194]
Weiqin Zou, Enming Li, and Chunrong Fang. 2021. BLESER: Bug localization based on enhanced semantic retrieval. arXiv:2109.03555 (2021)
2021 arXiv
-
[195]
Weiqin Zou, David Lo, Zhenyu Chen, Xin Xia, Yang Feng, and Baowen Xu. 2018. How practitioners perceive automated bug report management techniques. IEEE TSE 46, 8 (2018), 836–862. Appendix A ADDITIONAL MATERIAL , Vol. 1, No. 1, Article . Publication date: May 2025. 36 Niu et al...
2018
-
[196]
Daming Zou, Jingjing Liang, Yingfei Xiong, Michael D Ernst, and Lu Zhang. 2019. An empirical study of fault localization families and their combinations. IEEE TSE 47, 2 (2019), 332–347
2019
-
[2020]
arXiv:2011.03449 (2020)
DRAST–A Deep Learning and AST Based Approach for Bug Localization. arXiv:2011.03449 (2020)
2020 arXiv
-
[2023]
In 2023 IEEE/ACM 45th ICSE
Rat: A refactoring-aware traceability model for bug localization. In 2023 IEEE/ACM 45th ICSE. IEEE, 196–207
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.