REVIEW 2 major objections 5 minor 300 references
Towards a Science of Causal Interpretability in Deep Learning for Software Engineering
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This dissertation claims that neural code models can be understood causally by treating code properties as interventions in structural causal models, and that docode makes this practical.
desk verdict The dissertation's central confounding-bias illustration compares a bounded distributional divergence to a mean-difference ATE, and that mismatch, not just unmeasured confounders, is the real obstacle to accepting its flagship claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Structural Causal Model (SCM), a directed acyclic graph with structural equations that encodes assumptions about which software properties influence which model outcomes. docode's machinery is the four-step pipeline built on it: model the causal problem as an SCM, identify a causal estimand such as p(Y|do(T)) via do-calculus and back-door adjustment, estimate the effect with metrics like the Average Treatment Effect (ATE), and refute the estimate through placebo and robustness tests. Interventions are software-engineering-based perturbations, such as buggy versus fixed code, commented versus uncommented code, clone types, and AST node masking, applied to parallel code corpora.
What would settle it
A concrete check: take a model and dataset from the case study, randomly assign code snippets to buggy versus fixed treatments, and measure the difference in cross-entropy under this randomized assignment; if the randomized effect is non-negligible while docode's back-door ATE is near zero, then the SCM's confounder set was incomplete and the causal claim fails.
Extended reading notes
Core claim
The central claim is the Causal Interpretability Hypothesis: docode is a causal interpretability method that aims to make NCMs and their decision-making understandable by interpreting model prediction performance through Pearl's Ladder of Causation. Concretely, the dissertation reports that under docode's interventions, bugginess of code did not appear to causally influence prediction performance even when association was high, while Type II clones in training data did causally affect cross-entropy. It also reports that masking random tokens hurt BERT-like models more than masking grammar-based categories, that most studied models learned code-block tokens with less confounding bias, and that prompt semantics causally influenced ChatGPT performance in the Galeras benchmark. The author presents these as evidence that causal estimands can reveal confounding bias that associational interpretability misses.
Load-bearing premise
The load-bearing premise is that each SCM includes all relevant confounders, so that back-door adjustment recovers the true causal effect; the dissertation itself flags in Sec 3.6.3 that unidentified confounders could bias the relationship.
Editorial extensions
If this is right
- If the case study is right, benchmark comparisons of NCMs that rely on correlations or accuracy alone can rank models on spurious signal; causal estimates should be reported alongside.
- If docode estimates are valid, practitioners can distinguish code properties worth fixing from harmless correlates, guiding data curation and model debugging.
- The finding that grammar-based masking matters less than random masking for BERT-like models would imply these models may not encode full syntactic structure.
- The result that buggy code shows high association but near-zero causal effect would imply that some commonly reported 'bugginess hurts models' findings may be confounded by features like token count.
- The success of refutation testing in the pipeline would imply that causal interpretability can be made falsifiable, unlike purely descriptive explanation methods.
Reading between the lines
- Editorial inference: if the same back-door adjustment were applied to other code properties, such as identifier naming style, comment density, or code churn, many reported correlation-based findings in DL4SE might shrink or reverse.
- Editorial inference: docode's framework suggests a testable standardization where model cards report ATE, association, and the confounder set, so causal claims accumulate across studies.
- Editorial inference: the four-confounder set used in the case study is only a starting point; adding token frequency or training-data leakage as confounders could change the estimates, which is a direct empirical extension using the same pipeline.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This dissertation introduces docode, a post hoc causal interpretability method for Neural Code Models (NCMs), and evaluates it across multiple studies. The method is presented as a four-step pipeline: structural causal model construction, causal estimand identification, effect estimation via Average Treatment Effect (ATE), and refutation testing. The empirical core is a deep-code-generation case study spanning several NCM architectures and data interventions (buggy/fixed code, comments, clone types, AST node masking), supplemented by chapters on ASTrust, COMET, TraceXplainer, CodeQ, and the Galeras benchmark. The headline claim is that observed associations and causal effects diverge—for example, buggy code is correlated with NCM performance but allegedly does not causally influence it—and that docode can identify such confounding bias.
Significance. If the causal estimates were externally validated, docode would be a genuinely useful addition to DL4SE: it is explicitly grounded in Pearl's do-calculus, the case study covers a diverse set of architectures and datasets, several replication packages are released, and the proposed four-step framework gives practitioners a concrete starting point. The dissertation also states important limitations, including the possibility of unidentified confounders in Sec. 3.6.3. However, the central demonstration of confounding bias rests on a scale mismatch between a bounded distributional divergence (Jensen-Shannon distance) and a mean difference in cross-entropy units, and the hand-specified SCMs are not validated against any external causal ground truth. These issues materially weaken the empirical conclusions as currently stated, although they appear addressable with re-analysis and additional sensitivity testing.
major comments (2)
- [§2.4.1, Table 8.4, Definitions 7.4 and 8.1] The flagship example of confounding bias compares a Jensen-Shannon distance (bounded on [0,1]) with an ATE in cross-entropy units. These quantities are not on a common scale: a non-zero JSD can coexist with a zero mean difference when the treatment changes the variance, shape, or a small subset of the outcome distribution. Consequently, the reported contrast between JSD ≈ 0.67 and ATE ≈ −2E−4 does not by itself establish spurious correlation or confounding bias. The same issue affects Table 8.5, where Pearson correlations are compared with ATEs. To support the claim that 'buggy code did not appear to causally influence prediction performance', the authors should re-express association and causal effect on the same scale—for example, comparing the associational mean difference E[Y|T=1] − E[Y|T=0] with the back-door adjusted ATE, or comparing interventional and observational distributions with the same distributional metric.
- [§3.4.4, §3.6.3, §7.4] The causal conclusions depend on the correctness of hand-specified SCMs whose adjustment sets are not validated against any external benchmark. In the ASTrust validity study, only four confounders are used (cyclomatic complexity, AST levels, AST nodes, sequence size), and the text itself acknowledges in §3.6.3 that unidentified confounders may bias the causal relationship. The placebo refutations described in §7.4 test internal consistency of an assumed SCM but do not test whether the adjustment set is sufficient. The authors should add sensitivity analyses for unmeasured confounding (e.g., E-values or bounds under an omitted confounder), report uncertainty intervals for the ATEs in Tables 8.2–8.6, and, where possible, compare the back-door estimates against a randomized or quasi-experimental intervention.
minor comments (5)
- [§3.4.4, Table 3.4] The text states that 14 sub-categories were analyzed, but Table 3.4 lists 13 rows; the count and the table should be reconciled.
- [Table of Contents, Chapter 5] The heading '5.7 A Case Study in Industry' appears twice; the duplicated heading should be corrected.
- [Table 3.3, Fig. 3.9] The metric name is inconsistently spelled as 'BLUE-4' and 'CodeBLUE'; it should be BLEU-4 and CodeBLEU.
- [Fig. 2.4 caption] The caption uses p(Y|Z) ≈ 0.87 for what is described as a correlation; this notation is confusing and should be replaced with a Pearson correlation coefficient or equivalent.
- [Chapter 10] The autopoietic Ψ-Arch chapter is a speculative position with no empirical evaluation; the introduction should clearly frame it as a vision statement rather than a contribution on par with the evaluated tools.
Circularity Check
One validation claim (ASTrust RQ3) reduces partly by construction because treatment and outcome are both derived from the same token-level prediction probabilities, making the negative causal effect an algebraic consequence of the definitions; the core docode methodology itself is externally grounded in Pearl's framework and doWhy, while the flagship confounding-bias contrast (JSD vs ATE) is a…
-
self definitional
[Sec. 3.4.4 (Fig. 3.6 SCM for RQ3) and Sec. 3.5.3 (RQ3 conclusion)]
"We consider that the learning error (i.e., cross-entropy loss) of an LLM is causally impacted by the predicted probabilities of syntax elements. ... Negative effects indicate that the better a syntax category is predicted, the lower the learning error associated."
The outcome Y (cross-entropy loss) is computed as the negative log of each token's predicted probability summed over a snippet, while the treatment T (ASTrust performance for a syntax category) is the median of those same token-level prediction probabilities restricted to tokens in that category. Raising T mechanically lowers the -log(p) contributions of the category's tokens, hence lowers Y. The claimed causal relation T -> Y with a negative sign is therefore baked into the shared definitions: 'the better a syntax category is predicted, the lower the learning error' is an identity of how both quantities are constructed from the same NTP values, not an independent empirical discovery. Only the ATE magnitude is data-driven; its sign is forced.
-
self citation load bearing
[Sec. 3.4.4 (ASTrust validation) and Sec. 8.4 (docode case-study syntax clustering)]
"We conducted a causal inference analysis using the docode technique [PCR+23] to estimate SCs influence."
ASTrust's causal validity (RQ3) is established by applying docode, which is this dissertation's own prior work [PCR+23], and docode's case study (Sec. 8.4, 'Syntax Clustering of Code Predictions') relies on the syntax categories developed in Ch. 3. The two self-citations form a mutual-support loop: ASTrust validates docode's machinery and docode's machinery validates ASTrust, with no external causal ground truth for either. The loop is only partial, not total, because the estimation engine is externally anchored (Pearl's do-calculus and the doWhy implementation), the ATE magnitudes are data-driven, and the artifacts are publicly available and externally falsifiable in principle; nonetheless, the load-bearing validity argument for syntax-grounded explanations is internally generated.
full rationale
The core contribution, docode, is a methodological application of Pearl's Structural Causal Model / do-calculus framework, estimated through the external doWhy library on public NCMs and datasets; that machinery is independent of the paper's own conclusions, so the pipeline itself is not circular by construction. The empirical case-study findings (near-zero ATE for buggy code, positive ATE for Type II clones, token-masking effects) are data-driven estimates, not fitted parameters renamed as predictions. However, two issues warrant the moderate score. First, the RQ3 validation of ASTrust is partially circular: the treatment (per-category prediction confidence) and outcome (cross-entropy loss) are both defined from the same token-level predicted probabilities, so the negative causal effect is largely an algebraic consequence of the definitions; the paper even states the conclusion tautologically ('the better a syntax category is predicted, the lower the learning error'). Second, the dissertation's flagship demonstration of confounding bias (Sec. 2.4.1, Tables 8.2/8.4) contrasts a Jensen-Shannon Distance bounded by [0,1] with an ATE in cross-entropy units; a JSD near 0.67 and an ATE near -2E-4 are incommensurable, so the inference of 'spurious correlation' is a correctness/validity flaw rather than circularity, and per the review rules it is noted here rather than scored as a circular step. The self-citation loop between ASTrust and docode is present but only partially load-bearing given the external Pearl/doWhy anchor and public, externally falsifiable artifacts. Overall, the central claim retains independent empirical content, never reducing to a pure self-citation or definitional identity, hence a score of 4 rather than 6 or higher.
Assumptions & free parameters
free parameters (3)
- Syntax Categories set =
10 hand-defined categories (Decisions, Data Structures, Exceptions, Iterations, Functional Programming, Operators…
- ASTrust confidence threshold =
0.6
- Confounder set for causal analyses =
Number of Subwords (ProgramRepair); Cyclomatic Complexity, AST Levels, #AST Nodes, Sequence Size (ASTrust causal study)
assumptions (3)
- domain assumption The back-door criterion is satisfied by the chosen confounders, so the ATE is identifiable from observational data.
- domain assumption Token probabilities are well-calibrated proxies for the likelihood that a token prediction is correct.
- domain assumption Parallel code corpora (buggy/fixed, commented/uncommented, clone pairs, masked AST nodes) realize the SE-based interventions described.
invented entities (1)
-
Psi-Arch autopoietic architecture
Cite this review
Pith. "Pith review of Towards a Science of Causal Interpretability in Deep Learning for Software Engineering." pith.science (2026). https://pith.science/paper/N45TVOQY
@misc{pith2026250515023,
author = {Pith},
title = {Pith review of: Towards a Science of Causal Interpretability in Deep Learning for Software Engineering},
year = {2026},
howpublished = {\url{https://pith.science/paper/N45TVOQY}},
note = {Machine review of arXiv:2505.15023}
}
read the original abstract
This dissertation addresses achieving causal interpretability in Deep Learning for Software Engineering (DL4SE). While Neural Code Models (NCMs) show strong performance in automating software tasks, their lack of transparency in causal relationships between inputs and outputs limits full understanding of their capabilities. To build trust in NCMs, researchers and practitioners must explain code predictions. Associational interpretability, which identifies correlations, is often insufficient for tasks requiring intervention and change analysis. To address this, the dissertation introduces DoCode, a novel post hoc interpretability method for NCMs. DoCode uses causal inference to provide programming language-oriented explanations of model predictions. It follows a four-step pipeline: modeling causal problems using Structural Causal Models (SCMs), identifying the causal estimand, estimating effects with metrics like Average Treatment Effect (ATE), and refuting effect estimates. Its framework is extensible, with an example that reduces spurious correlations by grounding explanations in programming language properties. A case study on deep code generation across interpretability scenarios and various deep learning architectures demonstrates DoCode's benefits. Results show NCMs' sensitivity to code syntax changes and their ability to learn certain programming concepts while minimizing confounding bias. The dissertation also examines associational interpretability as a foundation, analyzing software information's causal nature using tools like COMET and TraceXplainer for traceability. It highlights the need to identify code confounders and offers practical guidelines for applying causal interpretability to NCMs, contributing to more trustworthy AI in software engineering.
Reference graph
Works this paper leans on
-
[1]
Asuncion, Arthur U
Hazeline U. Asuncion, Arthur U. Asuncion, and Richard N. Taylor. Software traceability with topic modeling. In Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering - Volume 1 , ICSE '10, pages 95--104, 2010
2010
-
[2]
Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Mart\' n Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. Tensorflow: a system for lar...
2016
-
[3]
Learning to represent programs with graphs
Miltiadis Allamanis, Marc Brockschmidt, and Mahmoud Khademi. Learning to represent programs with graphs. In International Conference on Learning Representations , 2018
2018
-
[4]
Information retrieval models for recovering traceability links between code and documentation
Antoniol, Canfora, Casazza, and De Lucia. Information retrieval models for recovering traceability links between code and documentation. In Proceedings of the International Conference on Software Maintenance , ICSE'00, pages 40--49, Oct 2000
2000
-
[5]
Unified pre-training for program understanding and generation
Wasi Uddin Ahmad, Saikat Chakraborty, Baishakhi Ray, and Kai-Wei Chang. Unified pre-training for program understanding and generation
-
[6]
Albrecht, Filippos Christianos, and Lukas Sch\"afer
Stefano V. Albrecht, Filippos Christianos, and Lukas Sch\"afer. Multi-Agent Reinforcement Learning: Foundations and Modern Approaches . MIT Press, 2024
2024
-
[7]
A literature review of automatic traceability links recovery for software change impact analysis
Thazin Win Win Aung, Huan Huo, and Yulei Sui. A literature review of automatic traceability links recovery for software change impact analysis. 2020 IEEE/ACM 28th International Conference on Program Comprehension (ICPC) , pages 14--24, 2020
2020
-
[8]
The adverse effects of code duplication in machine learning models of code
Miltiadis Allamanis. The adverse effects of code duplication in machine learning models of code. In Onward! OOPLSA 2019 , pages 143--153, 2019
2019
Show all 300 references
-
[9]
Abu-Mostafa, Malik Magdon-Ismail, and Hsuan-Tien Lin
Yaser S. Abu-Mostafa, Malik Magdon-Ismail, and Hsuan-Tien Lin. Learning from Data: A Short Course . AMLBook, United States, 2012
2012
-
[10]
The claude 3 model family: Opus, sonnet, haiku, 2024
Anthropic. The claude 3 model family: Opus, sonnet, haiku, 2024. Preprint
2024
-
[11]
Program synthesis with large language models, 2021
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models, 2021
2021
-
[12]
Bishop and Hugh Bishop
Christopher M. Bishop and Hugh Bishop. Deep Learning: Foundations and Concepts . Springer International Publishing, 2024
2024
-
[13]
Gpt-neox-20b: An open-source autoregressive language model, 2022
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. Gpt-neox-20b: An ope...
2022
-
[14]
Maximum a posteriori estimators as a limit of Bayes estimators , 2018
Robert Bassett and Julio Deride. Maximum a posteriori estimators as a limit of Bayes estimators , 2018
2018
-
[15]
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. A neural probabilistic language model . Advances in Neural Information Processing Systems , 3:1137--1155, 2003
2003
-
[16]
Artificial Life
M A Bedau. Artificial Life . Philosophy of Biology , pages 585--603, 2007
2007
-
[17]
Jasmijn Bastings and Katja Filippova. The elephant in the interpretability room: Why use attention as explanation when we have saliency methods? In Afra Alishahi, Yonatan Belinkov, Grzegorz Chrupa a, Dieuwke Hupkes, Yuval Pinter, and Hassan Sajjad, editors, Proceedings of the ...
2020
-
[18]
Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Margaret Mitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , FAccT '21, page 610–623, New York, NY,...
2021
-
[19]
Christopher M. Bishop. Pattern Recognition and Machine Learning (Information Science and Statistics) . Springer-Verlag New York, Inc., 2006
2006
-
[20]
Natural Language Processing with Python
Steven Bird, Ewan Klein, and Edward Loper. Natural Language Processing with Python . O'Reilly Media Inc., 2009
2009
-
[21]
NLTK : The natural language toolkit
Steven Bird and Edward Loper. NLTK : The natural language toolkit. In Proceedings of the ACL Interactive Poster and Demonstration Sessions , pages 214--217, Barcelona, Spain, July 2004. Association for Computational Linguistics
2004
-
[22]
API design for machine learning software: experiences from the scikit-learn project
Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexandre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Ga \" e l Varoquaux. API design for machine lear...
2013
-
[23]
A model-driven architecture approach to accelerate software code generation
Mayuri Bhadra, Daniela Sanchez Lopera, Robert Kunzelmann, and Wolfgang Ecker. A model-driven architecture approach to accelerate software code generation. In 2024 7th International Conference on Software and System Engineering (ICoSSE) , pages 23--30, 2024
2024
-
[24]
On identifiability in transformers, 2020
Gino Brunner, Yang Liu, Damián Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. On identifiability in transformers, 2020
2020
-
[25]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[26]
Biggerstaff, Bharat G
Ted J. Biggerstaff, Bharat G. Mitbander, and Dallas E. Webster. Program understanding and the concept assignment problem. Commun. ACM , 37(5):72--82, May 1994
1994
-
[27]
tree-sitter/tree-sitter: v0.25.3, March 2025
Max Brunsfeld, Amaan Qureshi, Andrew Hlynskyi, Patrick Thomson, ObserverOfTime, Will Lillis, Josh Vera, dundargoc, Phil Turnbull, Timothy Clem, Douglas Creager, Andrew Helwer, Rob Rix, Daumantas Kavolis, Hendrik van Antwerpen, Michael Davis, Christian Clason, Ika, Amin Ya, Ril...
2025
-
[28]
Sampling in Software Engineering Research : A Critical Review and Guidelines , October 2021
Sebastian Baltes and Paul Ralph. Sampling in Software Engineering Research : A Critical Review and Guidelines , October 2021. arXiv:2002.07764 [cs]
2021 arXiv
-
[29]
J. Brooke. SUS : A quick and dirty usability scale. In P. W. Jordan, B. Weerdmeester, A. Thomas, and I. L. Mclelland, editors, Usability Evaluation in Industry . Taylor and Francis , London, 1996
1996
-
[30]
Ullman, Fernando Martinez-Plumed, Joshua B
Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Rutar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernandez...
2023
-
[31]
Causality: The Place of the Causal Principle in Modern Science
Mario Bunge. Causality: The Place of the Causal Principle in Modern Science . Harvard University Press, 1959
1959
-
[32]
Emergence and Convergence: Qualitative Novelty and the Unity of Knowledge
Mario Bunge. Emergence and Convergence: Qualitative Novelty and the Unity of Knowledge . University of Toronto Press, 2003
2003
-
[33]
Philosophy of Science: Volume 1 and 2
Mario Bunge. Philosophy of Science: Volume 1 and 2 . Routledge, 2011
2011
-
[34]
The care and feeding of wild-caught mutants
David Bingham Brown, Michael Vaughn, Ben Liblit, and Thomas Reps. The care and feeding of wild-caught mutants. In Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering , ESEC/FSE 2017, page 511–522, New York, NY, USA, 2017. Association for Computing...
2017
-
[35]
An empirical exploration of trust dynamics in llm supply chains, 2024
Agathe Balayn, Mireia Yurrita, Fanny Rancourt, Fabio Casati, and Ujwal Gadiraju. An empirical exploration of trust dynamics in llm supply chains, 2024
2024
-
[36]
Nature’s Capacities and Their Measurement
Nancy Cartwright. Nature’s Capacities and Their Measurement . Oxford University Press, 1989
1989
-
[37]
The Dappled World: A Study of the Boundaries of Science
Nancy Cartwright. The Dappled World: A Study of the Boundaries of Science . Cambridge University Press, 1999
1999
-
[38]
Hunting Causes and Using Them: Approaches in Philosophy and Economics
Nancy Cartwright. Hunting Causes and Using Them: Approaches in Philosophy and Economics . Cambridge University Press, 2007
2007
-
[39]
An empirical study on the usage of bert models for code completion
Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. An empirical study on the usage of bert models for code completion. In 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR) , pages 108-...
2021
-
[40]
An empirical study on the usage of BERT models for code completion
Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. An empirical study on the usage of BERT models for code completion. CoRR , abs/2103.07115, 2021
2021 arXiv
-
[41]
An empirical study on the usage of transformer models for code completion
Matteo Ciniselli, Nathan Cooper, Luca Pascarella, Antonio Mastropaolo, Emad Aghajani, Denys Poshyvanyk, Massimiliano Di Penta, and Gabriele Bavota. An empirical study on the usage of transformer models for code completion. IEEE Transactions on Software Engineering , 48(12):481...
2022
-
[42]
C. S. Corley, K. Damevski, and N. A. Kraft. Exploring the use of deep learning for feature location. In 2015 IEEE International Conference on Software Maintenance and Evolution ( ICSME ) , ICSME'15, pages 556--560, September 2015. ISSN:
2015
-
[43]
Counterfactual explanations for models of code
Jürgen Cito, Isil Dillig, Vijayaraghavan Murali, and Satish Chandra. Counterfactual explanations for models of code. In Proceedings of the 44th International Conference on Software Engineering : Software Engineering in Practice , pages 125--134, Pittsburgh Pennsylvania, May 2022. ACM
2022
-
[44]
Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda
Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q. Feldman, Arjun Guha, Michael Greenberg, and Abhinav Jangda. MultiPL - E : A Scalable and Extensible Approach to Benchmark...
2022 arXiv
-
[45]
Constructing Grounded Theory: A Practical Guide through Qualitative Analysis
Kathy Charmaz. Constructing Grounded Theory: A Practical Guide through Qualitative Analysis . SAGE Publications Inc., 2006
2006
-
[46]
Chang, and Mark Christensen
Jane Cleland-Huang, Carl K. Chang, and Mark Christensen. Event-based traceability for managing evolutionary change. IEEE Trans. Softw. Eng. , 29(9), September 2003
2003
-
[47]
A machine learning approach for tracing regulatory codes to product specific requirements
Jane Cleland-Huang, Adam Czauderna, Marek Gibiec, and John Emenecker. A machine learning approach for tracing regulatory codes to product specific requirements. In Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering , ICSE'10, pages 155--164. ACM, 2010
2010
-
[48]
Connor, A
A. Connor, A. Harris, N. Cooper, and D. Poshyvanyk. Can we automatically fix bugs by learning edit operations? In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER) , pages 782--792, Los Alamitos, CA, USA, mar 2022. IEEE Computer Society
2022
-
[49]
Jane Cleland-Huang, Orlena C. Z. Gotel, Jane Huffman Hayes, Patrick M\" a der, and Andrea Zisman. Software traceability: Trends and future directions. In Proceedings of the on Future of Software Engineering , FOSE'14, pages 55--69. ACM, 2014
2014
-
[50]
Software and Systems Traceability
Jane Cleland-Huang, Orlena Gotel, and Andrea Zisman. Software and Systems Traceability . Springer Publishing Company, Incorporated, 2012
2012
-
[51]
Achieving lightweight trustworthy traceability
Jane Cleland-Huang, Mona Rahimi, and Patrick M\" a der. Achieving lightweight trustworthy traceability. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering , FSE'14, pages 849--852, 2014
2014
-
[52]
Sequencer: Sequence-to-sequence learning for end-to-end program repair
Zimin Chen, Steve Kommrusch, Michele Tufano, Louis-Noël Pouchet, Denys Poshyvanyk, and Martin Monperrus. Sequencer: Sequence-to-sequence learning for end-to-end program repair. IEEE Transactions on Software Engineering , 47(9):1943--1959, 2021
1943
-
[53]
Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, and Alexey Svyatkovskiy
Colin B. Clement, Shuai Lu, Xiaoyu Liu, Michele Tufano, Dawn Drain, Nan Duan, Neel Sundaresan, and Alexey Svyatkovskiy. Long-range modeling of source code files with eWASH : Extended window access by syntax hierarchy
-
[54]
Bigquery: Serverless, highly scalable, and cost-effective multi-cloud data warehouse, 2025
Google Cloud. Bigquery: Serverless, highly scalable, and cost-effective multi-cloud data warehouse, 2025. Accessed: 2025-04-07
2025
-
[55]
Debugging tool for code generation neural language models, March 2024
Colin Bruce Clement, David Alberto Nader Palacio, Neelakantan Sundaresan, Alexey Svyatkovskiy, and Michele Tufano. Debugging tool for code generation neural language models, March 2024. Application No. 18/082366, filed on December 15, 2022
2024
-
[56]
Cisco Systems
Inc. Cisco Systems. Libest: Enrollment over secure transport (est) library, 2013. Accessed: 2025-04-07
2013
-
[57]
Improving smart contract security with contrastive learning-based vulnerability detection
Yizhou Chen, Zeyu Sun, Zhihao Gong, and Dan Hao. Improving smart contract security with contrastive learning-based vulnerability detection. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , ICSE '24, New York, NY, USA, 2024. Association for...
2024
-
[58]
What Makes a Good Explanation ?: A Harmonized View of Properties of Explanations , December 2022
Zixi Chen, Varshini Subhash, Marton Havasi, Weiwei Pan, and Finale Doshi-Velez. What Makes a Good Explanation ?: A Harmonized View of Properties of Explanations , December 2022. arXiv:2211.05667 [cs]
2022 arXiv
-
[59]
Snopy: Bridging sample denoising with causal graph learning for effective vulnerability detection
Sicong Cao, Xiaobing Sun, Xiaoxue Wu, David Lo, Lili Bo, Bin Li, Xiaolei Liu, Xingwei Lin, and Wei Liu. Snopy: Bridging sample denoising with causal graph learning for effective vulnerability detection. In 2024 39th IEEE/ACM International Conference on Automated Software Engin...
2024
-
[60]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott G...
2021
-
[61]
On the Properties of Neural Machine Translation : Encoder - Decoder Approaches , October 2014
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio. On the Properties of Neural Machine Translation : Encoder - Decoder Approaches , October 2014. arXiv:1409.1259 [cs, stat]
2014 arXiv
-
[62]
A survey on open source software trustworthiness
Vieri del Bianco, Luigi Lavazza, Sandro Morasca, and Davide Taibi. A survey on open source software trustworthiness. IEEE Software , 28(5):67--75, 2011
2011
-
[63]
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei. Knowledge neurons in pretrained transformers. CoRR , abs/2104.08696, 2021
2021 arXiv
-
[64]
Enhancing software traceability by automatically expanding corpora with relevant documentation
Tathagata Dasgupta, Mark Grechanik, Evan Moritz, Bogdan Dit, and Denys Poshyvanyk. Enhancing software traceability by automatically expanding corpora with relevant documentation . IEEE International Conference on Software Maintenance, ICSM , pages 320--329, 2013
2013
-
[65]
Technique integration for requirements assessment
Alex Dekhtyar, Jane Huffman Hayes, Senthil Karthikeyan Sundaram, Elizabeth Ashlee Holbrook, and Olga Dekhtyar. Technique integration for requirements assessment. Proceedings of the 15th IEEE International Requirements Engineering Conference , pages 141--150, 2007
2007
-
[66]
Automatic traceability link recovery via active learning
Tianlong Du, Guo hua Shen, Zhi qiu Huang, Yaoliang Yu, and De xiang Wu. Automatic traceability link recovery via active learning. Frontiers of Information Technology & Electronic Engineering , 21:1217 -- 1225, 2020
2020
-
[67]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...
2024
-
[68]
Incremental approach and user feedbacks: a silver bullet for traceability recovery
Andrea De Lucia, Rocco Oliveto, and Paola Sgueglia. Incremental approach and user feedbacks: a silver bullet for traceability recovery. In Proceedings of the International Conference on Software Maintenance , ICSM'06, pages 299--309, 2006
2006
-
[69]
Ir-based traceability recovery processes: An empirical comparison of one-shot and incremental processes
Andrea De Lucia, Rocco Oliveto, and Genoveffa Tortora. Ir-based traceability recovery processes: An empirical comparison of one-shot and incremental processes. In Proceedings of the 2008 23rd IEEE/ACM International Conference on Automated Software Engineering , pages 39--48. I...
2008
-
[70]
Assessing ir-based traceability recovery tools through controlled experiments
Andrea De Lucia, Rocco Oliveto, and Genoveffa Tortora. Assessing ir-based traceability recovery tools through controlled experiments. Empirical Software Engineering , 14(1):57--92, 2009
2009
-
[71]
Supporting and accelerating reproducible research in software maintenance using tracelab component library
Bogdan Dit, Evan Moritz, Mario Linares-V \' a squez, and Denys Poshyvanyk. Supporting and accelerating reproducible research in software maintenance using tracelab component library . IEEE International Conference on Software Maintenance, ICSM , pages 330--339, 2013
2013
-
[72]
Configuring topic models for software engineering tasks in tracelab
Bogdan Dit, Annibale Panichella, Evan Moritz, Rocco Oliveto, Massimilano Di Penta, Denys Poshyvanyk, and Andrea De Lucia. Configuring topic models for software engineering tasks in tracelab. In 2013 7th International Workshop on Traceability in Emerging Forms of Software Engin...
2013
-
[73]
Integrating information retrieval, execution and link analysis algorithms to improve feature location in software
Bogdan Dit, Meghan Revelle, and Denys Poshyvanyk. Integrating information retrieval, execution and link analysis algorithms to improve feature location in software . Empirical Software Engineering , 18(2):277--309, 2013
2013
-
[74]
What is artificial life today, and where should it go? Artificial Life , 30(1):1--15, 02 2024
Alan Dorin and Susan Stepney. What is artificial life today, and where should it go? Artificial Life , 30(1):1--15, 02 2024
2024
-
[75]
Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals
Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The benchmark lottery. CoRR , abs/2107.07002, 2021
2021 arXiv
-
[76]
Towards a rigorous science of interpretable machine learning, 2017
Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning, 2017
2017
-
[77]
Considerations for Evaluation and Generalization in Interpretable Machine Learning , pages 3--17
Finale Doshi-Velez and Been Kim. Considerations for Evaluation and Generalization in Interpretable Machine Learning , pages 3--17. Springer International Publishing, Cham, 2018
2018
-
[78]
Eick, T.L
S.G. Eick, T.L. Graves, A.F. Karr, J.S. Marron, and A. Mockus. Does code decay? assessing the evidence from change management data. IEEE Transactions on Software Engineering , 27(1):1--12, 2001
2001
-
[79]
Hadeel Eladawy, Claire Le Goues, and Yuriy Brun. Automated program repair, what is it good for? not absolutely nothing! In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , ICSE '24, New York, NY, USA, 2024. Association for Computing Machinery
2024
-
[80]
Estimating the number of remaining links in traceability recovery
Davide Falessi, Massimiliano Di Penta, Gerardo Canfora, and Giovanni Cantone. Estimating the number of remaining links in traceability recovery. Empirical Software Engineering , 22(3):996--1027, 2017
2017
-
[81]
Furia, Robert Feldt, and Richard Torkar
Carlo A. Furia, Robert Feldt, and Richard Torkar. Bayesian data analysis in empirical software engineering research. IEEE Transactions on Software Engineering , abs/1811.05422, 2019
2019 arXiv
-
[82]
Automated repair of programs from large language models
Zhiyu Fan, Xiang Gao, Martin Mirchev, Abhik Roychoudhury, and Shin Hwei Tan. Automated repair of programs from large language models. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , pages 1469--1481, 2023
2023
-
[83]
C ode BERT : A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. C ode BERT : A pre-trained model for programming and natural languages. In Trevor Cohn, Yulan He, and Yang Liu, editors, Findings of the Associ...
2020
-
[84]
Vulexplainer: A transformer-based hierarchical distillation for explaining vulnerability types
Michael Fu, Van Nguyen, Chakkrit Kla Tantithamthavorn, Trung Le, and Dinh Phung. Vulexplainer: A transformer-based hierarchical distillation for explaining vulnerability types. IEEE Transactions on Software Engineering , 49(10):4550--4565, 2023
2023
-
[85]
Comparing explanation methods for traditional machine learning models part 1: An overview of current methods and quantifying their disagreement
Montgomery Flora, Corey Potvin, Amy McGovern , and Shawn Handler. Comparing explanation methods for traditional machine learning models part 1: An overview of current methods and quantifying their disagreement
-
[86]
Can cooperative multi-agent reinforcement learning boost automatic web testing? an exploratory study
Yujia Fan, Sinan Wang, Zebang Fei, Yao Qin, Huaxuan Li, and Yepang Liu. Can cooperative multi-agent reinforcement learning boost automatic web testing? an exploratory study. In 2024 39th IEEE/ACM International Conference on Automated Software Engineering (ASE) , pages 14--26, 2024
2024
-
[87]
Trace++: A traceability approach to support transitioning to agile software engineering
Felipe Furtado and Andrea Zisman. Trace++: A traceability approach to support transitioning to agile software engineering. In Requirements Engineering Conference (RE), 2016 IEEE 24th International , pages 66--75. IEEE, 2016
2016
-
[88]
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2020
2020
-
[89]
Codeparrot, 2021
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. Codeparrot, 2021
2021
-
[90]
Semantically enhanced software traceability using deep learning techniques
Jin Guo, Jinghui Cheng, and Jane Cleland-Huang. Semantically enhanced software traceability using deep learning techniques. In Proceedings of the 39th International Conference on Software Engineering , ICSE'17, pages 3--14. IEEE Press, 2017
2017
-
[91]
Foundations for an expert system in domain-specific traceability
Jin Guo, Jane Cleland-Huang, and Brian Berenbach. Foundations for an expert system in domain-specific traceability. In 2013 21st IEEE International Requirements Engineering Conference (RE) , pages 42--51, 2013
2013
-
[92]
Reconciling Manual and Automatic Refactoring
Xi Ge, Quinton L Dubose, and Emerson Murphy-Hill. Reconciling Manual and Automatic Refactoring . In 2012 34th International Conference on Software Engineering (ICSE) , pages 211 -- 221, Zurich, 2012. IEEE
2012
-
[93]
Github, 2020
github. Github, 2020
2020
-
[94]
Github copilot, 2025
GitHub . Github copilot, 2025. Accessed: 2025-04-07
2025
-
[95]
Gethers, R
M. Gethers, R. Oliveto, D. Poshyvanyk, and A. D. Lucia. On integrating orthogonal information retrieval methods to improve traceability recovery. In Proceedings of the International Conference on Software Maintenance , ICSM'11, pages 133--142, 2011
2011
-
[96]
On integrating orthogonal information retrieval methods to improve traceability recovery
Malcom Gethers, Rocco Oliveto, Denys Poshyvanyk, and Andrea De Lucia. On integrating orthogonal information retrieval methods to improve traceability recovery. In 2011 27th IEEE International Conference on Software Maintenance (ICSM) , pages 133--142, 2011
2011
-
[97]
Cold-start Software Analytics
Jin Guo, Mona Rahimi, Jane Cleland-Huang, Alexander Rasin, Jane Huffman Hayes, and Michael Vierhauser. Cold-start Software Analytics . In Proceedings of the 13th International Conference on Mining Software Repositories , MSR '16, pages 142--153, Austin, Texas, 2016. ACM
2016
-
[98]
Graphcode \ bert \ : Pre-training code representations with data flow
Daya Guo, Shuo Ren, Shuai Lu, Zhangyin Feng, Duyu Tang, Shujie LIU, Long Zhou, Nan Duan, Alexey Svyatkovskiy, Shengyu Fu, Michele Tufano, Shao Kun Deng, Colin Clement, Dawn Drain, Neel Sundaresan, Jian Yin, Daxin Jiang, and Ming Zhou. Graphcode \ bert \ : Pre-training code rep...
2021
-
[99]
Traceability recovery between bug reports and test cases-a mozilla firefox case study
Guilherme Gadelha, Franklin Ramalho, and Tiago Lima Massoni. Traceability recovery between bug reports and test cases-a mozilla firefox case study. Automated Software Engineering , 28, 2021
2021
-
[100]
Do automatic test generation tools generate flaky tests?, 2023
Martin Gruber, Muhammad Firhard Roslan, Owain Parry, Fabian Scharnböck, Phil McMinn, and Gordon Fraser. Do automatic test generation tools generate flaky tests?, 2023
2023
-
[101]
Towards automatic concept-based explanations
Amirata Ghorbani, James Wexler, James Zou, and Been Kim. Towards automatic concept-based explanations . Curran Associates Inc., Red Hook, NY, USA, 2019
2019
-
[102]
Deep code search
Xiaodong Gu, Hongyu Zhang, and Sunghun Kim. Deep code search. In Proceedings of the 40th International Conference on Software Engineering , ICSE '18, pages 933--944, New York, NY, USA, 2018. ACM
2018
-
[103]
Measuring coding challenge competence with apps
Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, and Jacob Steinhardt. Measuring coding challenge competence with apps. In J. Vanschoren and S. Yeung, editors, Proceedings of the Neural Inf...
2021
-
[104]
Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu
Abram Hindle, Earl T. Barr, Zhendong Su, Mark Gabel, and Premkumar Devanbu. On the naturalness of software. In Proceedings of the 34th International Conference on Software Engineering , ICSE '12, page 837–847. IEEE Press, 2012
2012
-
[105]
Hellendoorn and Premkumar Devanbu
Vincent J. Hellendoorn and Premkumar Devanbu. Are deep neural networks the best choice for modeling source code? ESEC/FSE 2017: Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering , pages 763--773, 2017
2017
-
[106]
Advancing candidate link generation for requirements tracing: The study of methods
Jane Huffman Hayes, Alex Dekhtyar, and Senthil Karthikeyan Sundaram. Advancing candidate link generation for requirements tracing: The study of methods. IEEE Transactions on Software Engineering , 32(1):4, 2006
2006
-
[107]
Homan and Andrew Gelman
Matthew D. Homan and Andrew Gelman. The no-u-turn sampler: adaptively setting path lengths in hamiltonian monte carlo. J. Mach. Learn. Res. , 15(1):1593–1623, January 2014
2014
-
[108]
Deep transfer learning for source code modeling
Yasir Hussain, Zhiqiu Huang, Yu Zhou, and Senzhang Wang. Deep transfer learning for source code modeling. Int. J. Softw. Eng. Knowl. Eng. , 30:649--668, 2020
2020
-
[109]
Deep code comment generation
Xing Hu, Ge Li, Xin Xia, David Lo, and Zhi Jin. Deep code comment generation. In Proceedings of the 26th Conference on Program Comprehension , ICPC '18, page 200–210, New York, NY, USA, 2018. Association for Computing Machinery
2018
-
[110]
Hopcroft, Rajeev Motwani, and Jeffrey D
John E. Hopcroft, Rajeev Motwani, and Jeffrey D. Ullman. Introduction to Automata Theory, Languages, and Computation (3rd Edition) . Addison-Wesley Longman Publishing Co., Inc., USA, 2006
2006
-
[111]
Afshin Mansouri, and Yuanyuan Zhang
Mark Harman, S. Afshin Mansouri, and Yuanyuan Zhang. Search-based software engineering: Trends, techniques and applications. ACM Comput. Surv. , 45(1), December 2012
2012
-
[112]
Which explanation should i choose? a function approximation perspective to characterizing post hoc explanations
Tessa Han, Suraj Srinivas, and Himabindu Lakkaraju. Which explanation should i choose? a function approximation perspective to characterizing post hoc explanations
-
[113]
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Hanchi Sun, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric P....
2024
-
[114]
Neuron dependency graphs: A causal abstraction of neural networks
Yaojie Hu and Jin Tian. Neuron dependency graphs: A causal abstraction of neural networks. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning , volume 162...
2022
-
[115]
Codeparrot(codeparrot)
Hugginface. Codeparrot(codeparrot). Accessed: 23 July 2024
2024
-
[116]
CodeSearchNet challenge: Evaluating the state of semantic code search
Hamel Husain, Ho-Hsiang Wu, Tiferet Gazit, Miltiadis Allamanis, and Marc Brockschmidt. CodeSearchNet challenge: Evaluating the state of semantic code search. arXiv preprint arXiv:1909.09436 , 2019
1909 arXiv
-
[117]
R. G. James, C. J. Ellison, and J. P. Crutchfield. dit : a P ython package for discrete information theory. The Journal of Open Source Software , 3(25):738, 2018
2018
-
[118]
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?, 2020
Alon Jacovi and Yoav Goldberg. Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?, 2020
2020
-
[119]
Incremental latent semantic indexing for automatic traceability link evolution management
Hsin-Yi Jiang, Tien N Nguyen, Xiang Chen, Hojun Jaygarl, and Carl K Chang. Incremental latent semantic indexing for automatic traceability link evolution management. In Proceedings of the 2008 23rd IEEE/ACM International Conference on Automated Software Engineering , ASE'08, p...
2008
-
[120]
Ai alignment: A comprehensive survey, 2025
Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Lukas Vierling, Donghai Hong, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Juntao Dai, Xuehai Pan, Kwan Yee Ng, Aidan O'Gara, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang...
2025
-
[121]
Sarthak Jain and Byron C. Wallace. Attention is not Explanation , May 2019. arXiv:1902.10186 [cs]
2019 arXiv
-
[122]
Tabnine: Ai code assistant, 2018
Jacob Jackson, Dror Weiss, and Eran Yahav. Tabnine: Ai code assistant, 2018. AI code assistant that accelerates and simplifies software development while keeping code private, secure, and compliant
2018
-
[123]
Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues?, 2024
2024
-
[124]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR , abs/1412.6980, 2015
2015 arXiv
-
[125]
Big code != big vocabulary: Open-vocabulary models for source code
Rafael Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, and Andrea Janes. Big code != big vocabulary: Open-vocabulary models for source code . Proceedings - International Conference on Software Engineering , pages 1073--1085, 2020
2020
-
[126]
Open-Vocabulary Models for Source Code (Extended Abstract)
Rafael Michael Karampatsis, Hlib Babii, Romain Robbes, Charles Sutton, and Andrea Janes. Open-Vocabulary Models for Source Code (Extended Abstract) . Proceedings - 2020 ACM/IEEE 42nd International Conference on Software Engineering: Companion, ICSE-Companion 2020 , pages 294--...
2020
-
[127]
Jenkins: The leading open source automation server, 2025
Kohsuke Kawaguchi and Jenkins Community. Jenkins: The leading open source automation server, 2025. Accessed: 2025-04-07
2025
-
[128]
Decomposing information into copying versus transformation
Artemy Kolchinsky and Bernat Corominas-Murtra. Decomposing information into copying versus transformation . Journal of the Royal Society Interface , 17(162):1--17, 2020
2020
-
[129]
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys - Dowmunt, Hitokazu Matsushita, and Arul Menezes. To ship or not to ship: An extensive evaluation of automatic metrics for machine translation. CoRR , abs/2107.10821, 2021
2021 arXiv
-
[130]
Traceclipse: an eclipse plug-in for traceability link recovery and management
Samuel Klock, Malcom Gethers, Bogdan Dit, and Denys Poshyvanyk. Traceclipse: an eclipse plug-in for traceability link recovery and management. In Proceedings of the international workshop on traceability in emerging forms of software engineering , TEFSE'11, pages 24--30, 2011
2011
-
[131]
Sharp nearby, fuzzy far away: How neural language models use context
Urvashi Khandelwal, He He, Peng Qi, and Dan Jurafsky. Sharp nearby, fuzzy far away: How neural language models use context. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages 284--294, Melbourne, Australia...
2018
-
[132]
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Fei - Fei Li. Visualizing and understanding recurrent networks. CoRR , abs/1506.02078, 2015
2015 arXiv
-
[133]
Palacio, Yixuan Zhang, and Denys Poshyvanyk
Dipin Khati, Yijin Liu, David N. Palacio, Yixuan Zhang, and Denys Poshyvanyk. Mapping the trust terrain: Llms in software engineering -- insights and perspectives, 2025
2025
-
[134]
Robust Statistical Methods for Empirical Software Engineering
Barbara Kitchenham, Lech Madeyski, David Budgen, Jacky Keung, Pearl Brereton, Stuart Charters, Shirley Gibbs, and Amnart Pohthong. Robust Statistical Methods for Empirical Software Engineering . Empirical Software Engineering , 22(2):579--630, 2017
2017
-
[135]
On the relationship between explanation and prediction: A causal view
Amir-Hossein Karimi, Krikamol Muandet, Simon Kornblith, Bernhard Schölkopf, and Been Kim. On the relationship between explanation and prediction: A causal view
-
[136]
u , Alexander Egyed, and Patrick M \
Hongyu Kuang, Jia Nie, Hao Hu, Patrick Rempel, Jian L \"u , Alexander Egyed, and Patrick M \"a der. Analyzing closeness of code dependencies for improving ir-based traceability recovery. In Proceedings of the 24th International Conference on Software Analysis, Evolution and Re...
2017
-
[137]
Sentencepiece: A simple and language-independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. Sentencepiece: A simple and language-independent subword tokenizer and detokenizer for neural text processing. https://github.com/google/sentencepiece, 2024. Accessed: 2024-11-21
2024
-
[138]
On weighted shapley values
Ehud Kalai and Dov Samet. On weighted shapley values. International Journal of Game Theory , 16:205--222, 1983
1983
-
[139]
Maybe deep neural networks are the best choice for modeling source code, 2019
Rafael-Michael Karampatsis and Charles Sutton. Maybe deep neural networks are the best choice for modeling source code, 2019
2019
-
[140]
Interpretability Beyond Feature Attribution : Quantitative Testing with Concept Activation Vectors ( TCAV ), June 2018
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, and Rory Sayres. Interpretability Beyond Feature Attribution : Quantitative Testing with Concept Activation Vectors ( TCAV ), June 2018. arXiv:1711.11279 [stat]
2018 arXiv
-
[141]
Galeras benchmark: Benchmarking causal analysis for interpreting llms for source code
SEMERU Lab. Galeras benchmark: Benchmarking causal analysis for interpreting llms for source code. https://github.com/WM-SEMERU/galeras-benchmark, 2025. Accessed: 2025-04-07
2025
-
[142]
Christopher G. Langton. Artificial life: The proceedings of an interdisciplinary workshop on the synthesis and simulation of living systems. In Artificial Life , volume VI of SFI Studies in the Sciences of Complexity , Redwood City, CA, 1989. Addison-Wesley
1989
-
[143]
Starcoder: may the source be with you!, 2023
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, ...
2023
-
[144]
Improving trace accuracy through data-driven configuration and composition of tracing features
Sugandha Lohar, Sorawit Amornborvornwong, Andrea Zisman, and Jane Cleland-Huang. Improving trace accuracy through data-driven configuration and composition of tracing features. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering , ESEC/FSE 2013,...
2013
-
[145]
Manning, Christopher Ré, Diana Acosta-Navas, Drew A
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...
-
[146]
Ensemble Models for Neural Source Code Summarization of Subroutines , July 2021
Alexander LeClair, Aakash Bansal, and Collin McMillan. Ensemble Models for Neural Source Code Summarization of Subroutines , July 2021. arXiv:2107.11423 [cs]
2021 arXiv
-
[147]
Improving chatgpt prompt for code generation, 2023
Chao Liu, Xuanlin Bao, Hongyu Zhang, Neng Zhang, Haibo Hu, Xiaohong Zhang, and Meng Yan. Improving chatgpt prompt for code generation, 2023
2023
-
[148]
Gershman, and Finale Doshi-Velez
Isaac Lage, Emily Chen, Jeffrey He, Menaka Narayanan, Been Kim, Samuel J. Gershman, and Finale Doshi-Velez. Human Evaluation of Models Built for Interpretability . Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , 7:59--67, October 2019
2019
-
[149]
Programs, Life Cycles, and Laws of Software Evolution
Meir Lehman. Programs, Life Cycles, and Laws of Software Evolution . Proceedings of the IEEE , 1980
1980
-
[150]
De Lucia, F
A. De Lucia, F. Fasano, R. Oliveto, and G. Tortora. Enhancing an artefact management system with traceability recovery features. In Proceedings of the 20th IEEE International Conference on Software Maintenance, 2004. Proceedings. , ICSM'04, pages 306--315, Sept 2004
2004
-
[151]
Vera Liao, Daniel Gruen, and Sarah Miller
Q. Vera Liao, Daniel Gruen, and Sarah Miller. Questioning the ai: Informing design practices for explainable ai user experiences. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems , CHI '20, page 1–15, New York, NY, USA, 2020. Association for Comp...
2020
-
[152]
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambrosio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Lidong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and S...
2021 arXiv
-
[153]
Zachary C. Lipton. The mythos of model interpretability
-
[154]
Lundberg and Su-In Lee
Scott M. Lundberg and Su-In Lee. A unified approach to interpreting model predictions. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NIPS'17, page 4768–4777, Red Hook, NY, USA, 2017. Curran Associates Inc
2017
-
[155]
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, 2019
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension, 2019
2019
-
[156]
Traceability transformed: Generating more accurate links with pre-trained bert models
Jinfeng Lin, Yalin Liu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang. Traceability transformed: Generating more accurate links with pre-trained bert models. In Proceedings of the 43rd International Conference on Software Engineering , ICSE '21, page 324–335. IEEE Press, 2021
2021
-
[157]
Trustworthy and synergistic artificial intelligence for software engineering: Vision and roadmaps
David Lo. Trustworthy and synergistic artificial intelligence for software engineering: Vision and roadmaps
-
[158]
Luca Longo, editor. Explainable Artificial Intelligence: First World Conference, xAI 2023, Lisbon, Portugal, July 26–28, 2023, Proceedings, Part I , volume 1901 of Communications in Computer and Information Science . Springer, 2023
2023
-
[159]
Explainable ai: A review of machine learning interpretability methods
Pantelis Linardatos, Vasilis Papastefanopoulos, and Sotiris Kotsiantis. Explainable ai: A review of machine learning interpretability methods. Entropy , 23(1), 2021
2021
-
[160]
Poudel, Wenhao Yu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang
Jinfeng Lin, A. Poudel, Wenhao Yu, Qingkai Zeng, Meng Jiang, and Jane Cleland-Huang. Enhancing automated software traceability by transfer learning from open-world data. ArXiv , abs/2207.01084, 2022
2022 arXiv
-
[161]
Lee and Katrina A
John D. Lee and Katrina A. See. Trust in automation: Designing for appropriate reliance. Human Factors , 46(1):50--80, 2004. PMID: 15151155
2004
-
[162]
Clacer: A deep learning-based compilation error classification method for novice students’ programs
Zheng Li, Fuxiang Sun, Haifeng Wang, Yifan Ding, Yong Liu, and Xiang Chen. Clacer: A deep learning-based compilation error classification method for novice students’ programs. In 2021 IEEE 45th Annual Computers, Software, and Applications Conference (COMPSAC) , pages 74--83, 2021
2021
-
[163]
No need to lift a finger anymore? assessing the quality of code generation by chatgpt, 2023
Zhijie Liu, Yutian Tang, Xiapu Luo, Yuming Zhou, and Liang Feng Zhang. No need to lift a finger anymore? assessing the quality of code generation by chatgpt, 2023
2023
-
[164]
Yue Liu, Chakkrit Tantithamthavorn, Li Li, and Yepang Liu. Explainable ai for android malware detection: Towards understanding why the models perform so well? In 2022 IEEE 33rd International Symposium on Software Reliability Engineering (ISSRE) , pages 169--180, 2022
2022
-
[165]
On the reliability and explainability of language models for program generation
Yue Liu, Chakkrit Tantithamthavorn, Yonghui Liu, and Li Li. On the reliability and explainability of language models for program generation. ACM Trans. Softw. Eng. Methodol. , 33(5), June 2024
2024
-
[166]
AST - Probe : Recovering abstract syntax trees from hidden representations of pre-trained language models, September 2022
José Antonio Hernández López, Martin Weyssow, Jesús Sánchez Cuadrado, and Houari Sahraoui. AST - Probe : Recovering abstract syntax trees from hidden representations of pre-trained language models, September 2022. arXiv:2206.11719 [cs]
2022 arXiv
-
[167]
D. Li , Z. Wang , and Y. Xue . Fine-grained android malware detection based on deep learning. In 2018 IEEE Conference on Communications and Network Security (CNS) , pages 1--2, May 2018
2018
-
[168]
Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation, 2023
2023
-
[169]
Deep learning based feature envy detection
Hui Liu, Zhifeng Xu, and Yanzhen Zou. Deep learning based feature envy detection. In Proceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering , ASE 2018, pages 385--396, New York, NY, USA, 2018. ACM
2018
-
[170]
Trustworthy llms: a survey and guideline for evaluating large language models' alignment, 2024
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo, Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li. Trustworthy llms: a survey and guideline for evaluating large language models' alignment, 2024
2024
-
[171]
Do pretrained language models indeed understand software engineering tasks? IEEE Transactions on Software Engineering , 49(10):4639--4655, 2023
Yao Li, Tao Zhang, Xiapu Luo, Haipeng Cai, Sen Fang, and Dawei Yuan. Do pretrained language models indeed understand software engineering tasks? IEEE Transactions on Software Engineering , 49(10):4639--4655, 2023
2023
-
[172]
P. Liu , X. Zhang , M. Pistoia , Y. Zheng , M. Marques , and L. Zeng . Automatic text input generation for mobile testing. In 2017 IEEE/ACM 39th International Conference on Software Engineering (ICSE) , pages 643--653, May 2017
2017
-
[173]
David J. C. MacKay. Information theory, inference, and learning algorithms . Cambridge University Press, 2003
2003
-
[174]
Autopoesis and Cognition
Varela Francisco Maturana, Humberto. Autopoesis and Cognition . Springer, 1980
1980
-
[175]
Better automatic program repair by using bug reports and tests together
Manish Motwani and Yuriy Brun. Better automatic program repair by using bug reports and tests together. In 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , pages 1225--1237, May 2023
2023
-
[176]
Interpretable machine learning – a brief history, state-of-the-art and challenges
Christoph Molnar, Giuseppe Casalicchio, and Bernd Bischl. Interpretable machine learning – a brief history, state-of-the-art and challenges. In Mohammed J. Zaki, Panagiotis Papapetrou, Victor S. Sheng, Aristides Gionis, Lars Schmidt-Thieme, and Wolfgang J. Krause, editors, ECM...
2020
-
[177]
Automatic Traceability Maintenance via Machine Learning Classification
Chris Mills, Javier Escobar-Avila, and Sonia Haiduc. Automatic Traceability Maintenance via Machine Learning Classification . In Proceedings of the 34th IEEE International Conference on Software Maintenance and Evolution , ICSME'18, pages 369--380, Madrid, Spain, September 2018. ACM
2018
-
[178]
Mac \' i as-Escriv \' a , Rodolfo Haber, Raul Del Toro, and Vicente Hernandez
Frank D. Mac \' i as-Escriv \' a , Rodolfo Haber, Raul Del Toro, and Vicente Hernandez. Self-adaptive systems: A survey of current approaches, research challenges and applications . Expert Systems with Applications , 40(18):7267--7279, 2013
2013
-
[179]
Data mining static code attributes to learn defect predictors
Tim Menzies, Jeremy Greenwald, and Art Frank. Data mining static code attributes to learn defect predictors. IEEE Transactions on Software Engineering , 33(1):2--13, 2007
2007
-
[180]
Visual studio intellicode, 2025
Microsoft . Visual studio intellicode, 2025. Accessed: 2025-04-07
2025
-
[181]
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence , 267:1--38, 2019
2019
-
[182]
Causal inference and discovery in Python: unlock the secrets of modern causal machine learning with DoWhy , EconML , PyTorch and more
Aleksander Molak and Ajit Jaokar. Causal inference and discovery in Python: unlock the secrets of modern causal machine learning with DoWhy , EconML , PyTorch and more . Packt Publishing Limited, 2023
2023
-
[183]
Deepgauge: Multi-granularity testing criteria for deep learning systems
Lei Ma, Felix Juefei-Xu, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Chunyang Chen, Ting Su, Li Li, Yang Liu, Jianjun Zhao, and Yadong Wang. Deepgauge: Multi-granularity testing criteria for deep learning systems. In Proceedings of the 33rd ACM/IEEE International Conference o...
2018
-
[184]
M\" a der, P
P. M\" a der, P. L. Jones, Y. Zhang, and J. Cleland-Huang. Strategic traceability for safety-critical projects. IEEE Software , 30(3):58--66, May 2013
2013
-
[185]
A robust multi-objective approach to balance severity and importance of refactoring opportunities , 2016
Mohamed Wiem Mkaouer, Marouane Kessentini, Mel Cinn \' e ide, Shinpei Hayashi, and Kalyanmoy Deb. A robust multi-objective approach to balance severity and importance of refactoring opportunities , 2016
2016
-
[186]
Murphy, Mik Kersten, and Leah Findlater
Gail C. Murphy, Mik Kersten, and Leah Findlater. How are java software developers using the eclipse ide? IEEE Softw. , 23(4):76–83, July 2006
2006
-
[187]
Causal interpretability for machine learning -- problems, methods and evaluation, 2020
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu. Causal interpretability for machine learning -- problems, methods and evaluation, 2020
2020
-
[188]
James Murdoch, Peter J
W. James Murdoch, Peter J. Liu, and Bin Yu. Beyond word importance: Contextual decomposition to extract interactions from LSTM s. In International Conference on Learning Representations , 2018
2018
-
[189]
Andrian Marcus and Jonathan I. Maletic. Recovering documentation-to-source-code traceability links using latent semantic indexing. In Proceedings of the 25th International Conference on Software Engineering , ICSE '03, pages 125--135, Washington, DC, USA, 2003. IEEE Computer Society
2003
-
[190]
Khapra, Balaji Vasan Srinivasan, and Balaraman Ravindran
Akash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan, Mitesh M. Khapra, Balaji Vasan Srinivasan, and Balaraman Ravindran. Towards Transparent and Explainable Attention Models . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , page...
2020
-
[191]
Mahmoud, N
A. Mahmoud, N. Niu, and S. Xu. A semantic relatedness approach for traceability link recovery. In Proceedings of the 20th IEEE International Conference on Program Comprehension , ICPC'12, pages 183--192, June 2012
2012
-
[192]
Interpretable Machine Learning
Christoph Molnar. Interpretable Machine Learning . 3 edition, 2025
2025
-
[193]
User experience design, 2004
Peter Morville. User experience design, 2004. Accessed: 2025-04-07
2004
-
[194]
Reviewing the landscape: Component-based software engineering practices and challenges
Pradeep Kumar Mahapatro and Neelamadhab Padhy. Reviewing the landscape: Component-based software engineering practices and challenges. In 2024 International Conference on Emerging Systems and Intelligent Computing (ESIC) , pages 360--365, 2024
2024
-
[195]
Palacio, Carlos Bernal-C\' a rdenas, Daniel McCrystal, Denys Poshyvanyk, Chris Shenefiel, and Jeff Johnson
Kevin Moran, David N. Palacio, Carlos Bernal-C\' a rdenas, Daniel McCrystal, Denys Poshyvanyk, Chris Shenefiel, and Jeff Johnson. Improving the effectiveness of traceability link recovery using hierarchical bayesian networks. In Proceedings of the ACM/IEEE 42nd International C...
2020
-
[196]
improving the effectiveness of traceability link recovery using hierarchical bayesian networks
Kevin Moran, David N. Palacio, Carlos Bernal-Cardenas, Daniel McCrystal, Denys Poshyvanyk, Chris Shenefiel, and Jeff Johnson. Online appendix for "improving the effectiveness of traceability link recovery using hierarchical bayesian networks", 2020. Accessed: 2025-04-07
2020
-
[197]
Combining textual and structural analysis of software artifacts for traceability link recovery
Collin McMillan, Denys Poshyvanyk, and Meghan Revelle. Combining textual and structural analysis of software artifacts for traceability link recovery. In Proceedings of the ICSE Workshop on Traceability in Emerging Forms of Software Engineering , TEFSE '09, pages 41--48. IEEE ...
2009
-
[198]
Studying the usage of text-to-text transfer transformer to support code-related tasks
Antonio Mastropaolo, Simone Scalabrino, Nathan Cooper, David Nader Palacio, Denys Poshyvanyk, Rocco Oliveto, and Gabriele Bavota. Studying the usage of text-to-text transfer transformer to support code-related tasks. In Proceedings of the 43rd International Conference on Softw...
2021
-
[199]
Ahmad Haji Mohammadkhani, Chakkrit Tantithamthavorn, and Hadi Hemmatif. Explaining transformer-based code models: What do they learn? when they do not work? In 2023 IEEE 23rd International Working Conference on Source Code Analysis and Manipulation (SCAM) , pages 96--106, 2023
2023
-
[200]
Kevin P. Murphy. Machine Learning: A Probabilistic Perspective . The MIT Press, 2012
2012
-
[201]
Detecting, classifying, and tracing non-functional software requirements
Anas Mahmoud and Grant Williams. Detecting, classifying, and tracing non-functional software requirements. Requirements Engineering , 21(3):357--381, Sep 2016
2016
-
[202]
An empirical investigation into the use of image captioning for automated software documentation
Kevin Moran, Ali Yachnes, George Purnell, Junayed Mahmud, Michele Tufano, Carlos Bernal Cardenas, Denys Poshyvanyk, and Zach H'Doubler. An empirical investigation into the use of image captioning for automated software documentation. In 2022 IEEE International Conference on So...
2022
-
[203]
The general and logical theory of automata
John Von Neumann. The general and logical theory of automata. In L.A. Jeffress, editor, Cerebral Mechanisms in Behaviour—The Hixon Symposium , pages 1--41, New York, NY, USA, 1951. Wiley
1951
-
[204]
Theory of Self-Reproducing Automata
John Von Neumann. Theory of Self-Reproducing Automata . University of Illinois Press, Champaign, IL, USA, 1966
1966
-
[205]
J. Neyman. Outline of a Theory of Statistical Estimation Based on the Classical Theory of Probability . Royal Society, 1937
1937
-
[206]
Anh Tuan Nguyen and Tien N. Nguyen. Graph-based statistical language model for code. In Proceedings of the 37th International Conference on Software Engineering - Volume 1 , ICSE '15, page 858–868. IEEE Press, 2015
2015
-
[207]
Nguyen, A
T. Nguyen, A. Nguyen, and H. Nguyen. A statistical semantic language model for source code. In ESEC/FSE 2013 , 2013
2013
-
[208]
Tung Thanh Nguyen, Anh Tuan Nguyen, Hoan Anh Nguyen, and Tien N. Nguyen. A statistical semantic language model for source code. In Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering , ESEC/FSE 2013, page 532–542, New York, NY, USA, 2013. Associati...
2013
-
[209]
Codegen: An open large language model for code with multi-turn program synthesis
Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. Codegen: An open large language model for code with multi-turn program synthesis. In The Eleventh International Conference on Learning Representations , 2023
2023
-
[210]
Assessing Single-Objective Performance Convergence and Time Complexity for Refactoring Detection
David Nader-palacio, Daniel Rodr \' i guez-c \' a rdenas, and Jonatan Gomez. Assessing Single-Objective Performance Convergence and Time Complexity for Refactoring Detection . In Proceedings of the Genetic and Evolutionary Computation Conference Companion on - GECCO '18 , page...
2018
-
[211]
A sysml-based approach to traceability management and design slicing in support of safety certification: Framework, tool support, and case studies
Shiva Nejati, Mehrdad Sabetzadeh, Davide Falessi, Lionel Briand, and Thierry Coq. A sysml-based approach to traceability management and design slicing in support of safety certification: Framework, tool support, and case studies. Inf. Softw. Technol. , 54(6):569--590, June 2012
2012
-
[212]
Recovering transitive traceability links among software artifacts
Kazuki Nishikawa, Hironori Washizaki, Yoshiaki Fukazawa, Keishi Oshima, and Ryota Mibe. Recovering transitive traceability links among software artifacts. In Proceedings of the IEEE International Conference on Software Maintenance and Evolution , ICSME'15, pages 576--580. IEEE, 2015
2015
-
[213]
Khan, and Bashar Nuseibeh
Armstrong Nhlabatsi, Yijun Yu, Andrea Zisman, Thein Tun, Niamul Khan, Arosha Bandara, Khaled M. Khan, and Bashar Nuseibeh. Managing security control assumptions using causal traceability. In Proceedings of the 8th International Symposium on Software and Systems Traceability , ...
2015
-
[214]
Barr, and Justyna Petke
Leandro Oliveria de Souza, Eduardo Santana de Almeida, Paulo Anselmo da Mota Silveira Neto, Earl T. Barr, and Justyna Petke. Software product line engineering via software transplantation. ACM Trans. Softw. Eng. Methodol. , 34(2), January 2025
2025
-
[215]
Coest traceability datasets, 2025
Center of Excellence for Software & Systems Traceability (CoEST). Coest traceability datasets, 2025. Accessed: 2025-04-07
2025
-
[216]
On the equivalence of information retrieval methods for automated traceability link recovery
Rocco Oliveto, Malcom Gethers, Denys Poshyvanyk, and Andrea De Lucia. On the equivalence of information retrieval methods for automated traceability link recovery. In Proceedings of the 18th International Conference on Program Comprehension , ICPC '10, pages 68--71, 2010
2010
-
[217]
Review of system development life cycle (sdlc) models for effective application delivery
Oluwaseyi Ezekiel Olorunshola and Francisca Nonyelum Ogwueleka. Review of system development life cycle (sdlc) models for effective application delivery. In Amit Joshi, Mufti Mahmud, Roshan G. Ragel, and Nileshsingh V. Thakur, editors, Information and Communication Technology ...
2020
-
[218]
Gpt-4 technical report, 2023
OpenAI. Gpt-4 technical report, 2023
2023
-
[219]
Java language keywords, 2024
Oracle. Java language keywords, 2024. Accessed: 2025-04-07
2024
-
[220]
David N. Palacio. Information theory for interpreting software traceability. https://danaderp.github.io/danaderp/projects/1_project/, 2023
2023
-
[221]
Transportability of causal and statistical relations: A formal approach
Judea Pearl and Elias Bareinboim. Transportability of causal and statistical relations: A formal approach. In 2011 IEEE 11th International Conference on Data Mining Workshops , pages 540--547, 2011
2011
-
[222]
Palacio, Nathan Cooper, Alvaro Rodriguez, Kevin Moran, and Denys Poshyvanyk
David N. Palacio, Nathan Cooper, Alvaro Rodriguez, Kevin Moran, and Denys Poshyvanyk. Toward a Theory of Causation for Interpreting Neural Code Models , February 2023. arXiv:2302.03788 [cs, stat]
2023 arXiv
-
[223]
Deep learning & software engineering: State of research and future directions
Devanbu Prem, Matthew Dwyer, Sebastian Elbaum, Michael Lowry, Kevin Moran, Denys Poshyvanyk, Baishakhi Ray, Rishabh Singh, and Xiangyu Zhang. Deep learning & software engineering: State of research and future directions. In Proceedings of the 2019 NSF Workshop on Deep Learning...
2019
-
[224]
Causal inference in statistics: An overview
Judea Pearl. Causal inference in statistics: An overview. Statistics Surveys , 3:96--146, 2009
2009
-
[225]
Causality: Models, Reasoning and Inference
Judea Pearl. Causality: Models, Reasoning and Inference . Cambridge University Press, USA, 2nd edition, 2009
2009
-
[226]
Theoretical impediments to machine learning with seven sparks from the causal revolution
Judea Pearl. Theoretical impediments to machine learning with seven sparks from the causal revolution. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining , WSDM '18, page 3, New York, NY, USA, 2018. Association for Computing Machinery
2018
-
[227]
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu ...
2019
-
[228]
Causal Inference in Statistics, A Primer
Judea Pearl, Madelyn Glymour, and Nicholas P.Jewell. Causal Inference in Statistics, A Primer . Wiley, 2016
2016
-
[229]
Perturbation sensitivity analysis to detect unintended model biases
Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. Perturbation sensitivity analysis to detect unintended model biases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural L...
2019
-
[230]
Palacio, Dipin Khati, Daniel Rodriguez-Cardenas, Alejandro Velasco, and Denys Poshyvanyk
David N. Palacio, Dipin Khati, Daniel Rodriguez-Cardenas, Alejandro Velasco, and Denys Poshyvanyk. Codeq: On explaining (large) language models for code using global code-based explanations, 2025
2025
-
[231]
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie. The Book of Why: The New Science of Cause and Effect . Basic Books, Inc., USA, 1st edition, 2018
2018
-
[232]
Palacio, Daniel McCrystal, Kevin Moran, Carlos Bernal-Cárdenas, Denys Poshyvanyk, and Chris Shenefiel
David N. Palacio, Daniel McCrystal, Kevin Moran, Carlos Bernal-Cárdenas, Denys Poshyvanyk, and Chris Shenefiel. Learning to identify security-related issues using convolutional neural networks. In 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME)...
2019
-
[233]
Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, and Denys Poshyvanyk
David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, and Denys Poshyvanyk. Astrust: Github repository, 2024
2024
-
[234]
Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, and Denys Poshyvanyk
David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, and Denys Poshyvanyk. Towards more trustworthy and interpretable llms for code through syntax-grounded explanations, 2024
2024
-
[235]
BLEU : a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU : a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics , ACL '02, pages 311--318, USA, July 2002. Association for Compu...
2002
-
[236]
Goldstein, Jake M
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. Manipulating and Measuring Model Interpretability , August 2021. arXiv:1802.07810 [cs]
2021 arXiv
-
[237]
Pyexplainer: Explaining the predictions of just-in-time defect models
Chanathip Pornprasit, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Michael Fu, and Patanamon Thongtanunam. Pyexplainer: Explaining the predictions of just-in-time defect models. In 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE) , pages 40...
2021
-
[238]
Qualtrics experience management platform, 2025
Qualtrics. Qualtrics experience management platform, 2025. Accessed: 2025-04-10
2025
-
[239]
Software evolution and maintenance
V\' a clav Rajlich. Software evolution and maintenance. In Future of Software Engineering Proceedings , FOSE 2014, page 133–144, New York, NY, USA, 2014. Association for Computing Machinery
2014
-
[240]
On the generalizability of neural program models with respect to semantic-preserving program transformations
Md Rafiqul Islam Rabin, Nghi DQ Bui, Ke Wang, Yijun Yu, Lingxiao Jiang, and Mohammad Amin Alipour. On the generalizability of neural program models with respect to semantic-preserving program transformations. Information and Software Technology , 135:106552, 2021
2021
-
[241]
tracexplainer, 2024
RepoTrace. tracexplainer, 2024. [Accessed 11-24-2024]
2024
-
[242]
Codebleu: a method for automatic evaluation of code synthesis
Shuo Ren, Daya Guo, Shuai Lu, Long Zhou, Shujie Liu, Duyu Tang, Neel Sundaresan, Ming Zhou, Ambrosio Blanco, and Shuai Ma. Codebleu: a method for automatic evaluation of code synthesis. CoRR , abs/2009.10297, 2020
2009 arXiv
-
[243]
naturalness
Baishakhi Ray, Vincent Hellendoorn, Saheel Godhane, Zhaopeng Tu, Alberto Bacchelli, and Premkumar Devanbu. On the "naturalness" of buggy code. In Proceedings of the 38th International Conference on Software Engineering , ICSE '16, page 428–439, New York, NY, USA, 2016. Associa...
2016
-
[244]
Know what you don ' t know: Unanswerable questions for SQ u AD
Pranav Rajpurkar, Robin Jia, and Percy Liang. Know what you don ' t know: Unanswerable questions for SQ u AD . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages 784--789, Melbourne, Australia, July 2018....
2018
-
[245]
Unsupervised translation of programming languages
Baptiste Roziere, Marie-Anne Lachaux, Lowik Chanussot, and Guillaume Lample. Unsupervised translation of programming languages. In H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems , volume 33, pages 206...
2020
-
[246]
Mind the Gap : Assessing the Conformance of Software Traceability to Relevant Guidelines
Patrick Rempel, Patrick Mader, Tobias Kuschke, and Jane Cleland-Huang. Mind the Gap : Assessing the Conformance of Software Traceability to Relevant Guidelines . In Proceedings of the 36th International Conference on Software Engineering , ICSE'14, pages 943--954, Hyderabad, I...
2014
-
[247]
Traceability in the wild: automatically augmenting incomplete trace links
Michael Rath, Jacob Rendall, Jin LC Guo, Jane Cleland-Huang, and Patrick M \"a der. Traceability in the wild: automatically augmenting incomplete trace links. In Proceedings of the 40th International Conference on Software Engineering , ICSE'18, pages 834--845. ACM, 2018
2018
-
[248]
Ra \" ffa and R
H. Ra \" ffa and R. Schlaifer. Applied statistical decision theory . Studies in managerial economics. Division of Research, Graduate School of Business Adminitration, Harvard University, 1961
1961
-
[249]
COMET : A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. COMET : A neural framework for MT evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages 2685--2702, Online, November 2020. Association for Computational L...
2020
-
[250]
why should i trust you?
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. "why should i trust you?": Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , KDD '16, page 1135–1144, New York, NY, USA, ...
2016
-
[251]
A Multi-study Investigation Into Dead Code
Simone Romano, Christopher Vendome, Giuseppe Scanniello, and Denys Poshyvanyk. A Multi-study Investigation Into Dead Code . IEEE Transactions on Software Engineering , 2018
2018
-
[252]
Vechev, and Eran Yahav
Veselin Raychev, Martin T. Vechev, and Eran Yahav. Code completion with statistical language models. Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation , 2014
2014
-
[253]
Beyond accuracy: Behavioral testing of NLP models with C heck L ist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of NLP models with C heck L ist. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages 4902--4912, Online, July 2020. Associa...
2020
-
[254]
Desiderata for interpretability: Explaining decision tree predictions with counterfactuals
Kacper Sokol and Peter Flach. Desiderata for interpretability: Explaining decision tree predictions with counterfactuals. Proceedings of the AAAI Conference on Artificial Intelligence , 33:10035--10036, 07 2019
2019
-
[255]
Learning Important Features Through Propagating Activation Differences , October 2019
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. Learning Important Features Through Propagating Activation Differences , October 2019. arXiv:1704.02685 [cs]
2019 arXiv
-
[256]
Dataflow analysis-inspired deep learning for efficient vulnerability detection
Benjamin Steenhoek, Hongyang Gao, and Wei Le. Dataflow analysis-inspired deep learning for efficient vulnerability detection. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , ICSE '24, New York, NY, USA, 2024. Association for Computing Machinery
2024
-
[257]
Calibration and correctness of language models for code, 2024
Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, and Toufique Ahmed. Calibration and correctness of language models for code, 2024
2024
-
[258]
Revising rules to capture requirements traceability relations: A machine learning approach
George Spanoudakis, Artur S d'Avila Garcez, and Andrea Zisman. Revising rules to capture requirements traceability relations: A machine learning approach. In SEKE , SEKE'03, pages 570--577, 2003
2003
-
[259]
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural machine translation of rare words with subword units. arXiv preprint arXiv:1508.07909 , 2015
2015 arXiv
-
[260]
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, Zhengliang Liu, Yixin Liu, Yijue Wang, Zhikun Zhang, Bertie Vidgen, Bhavya Kailkhura, Caiming Xiong, Chaowei Xiao, Chunyuan Li, Eric Xing, Furong H...
2024
-
[261]
Do W hy: A Python package for causal inference
Amit Sharma, Emre Kiciman, et al. Do W hy: A Python package for causal inference . https://github.com/microsoft/dowhy, 2019
2019
-
[262]
Characterizing software maintenance meetings: Information shared, discussion outcomes, and information captured
Adriana Meza Soria, Taylor Lopez, Elizabeth Seero, Negin Mashhadi, Emily Evans, Janet Burge, and Andr\' e Van der Hoek. Characterizing software maintenance meetings: Information shared, discussion outcomes, and information captured. In Proceedings of the IEEE/ACM 46th Internat...
2024
-
[263]
A literature review of using machine learning in software development life cycle stages
Saad Shafiq, Atif Mashkoor, Christoph Mayr-Dorn, and Alexander Egyed. A literature review of using machine learning in software development life cycle stages. IEEE Access , 9:140896--140920, 2021
2021
-
[264]
Hiroki Sayama and Chrystopher L. Nehaniv. Self-reproduction and evolution in cellular automata: 25 years after evoloops. Artificial Life , 31(1):81–95, February 2024
2024
-
[265]
Explainable artificial intelligence (xai)
Management Solutions. Explainable artificial intelligence (xai). challenges of model interpretability. https://www.managementsolutions.com/sites/default/files/minisite/static/22959b0f-b3da-47c8-9d5c-80ec3216552b/iax/pdf/explainable-artificial-intelligence-en-04.pdf, 2022. Acce...
2022
-
[266]
What is usability? a characterization based on iso 9241-11 and iso/iec 25010, 2022
Maximilian Speicher. What is usability? a characterization based on iso 9241-11 and iso/iec 25010, 2022
2022
-
[267]
Jeffrey Svajlenko and Chanchal K. Roy. Evaluating clone detection tools with BigCloneBench . 2015 IEEE 31st International Conference on Software Maintenance and Evolution, ICSME 2015 - Proceedings , pages 131--140, 2015
2015
-
[268]
Sofia Serrano and Noah A. Smith. Is Attention Interpretable ?, June 2019. arXiv:1906.03731 [cs]
2019 arXiv
-
[269]
Dowhy: Addressing challenges in expressing and validating causal assumptions, 2021
Amit Sharma, Vasilis Syrgkanis, Cheng Zhang, and Emre Kıcıman. Dowhy: Addressing challenges in expressing and validating causal assumptions, 2021
2021
-
[270]
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep networks. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 , ICML'17, page 3319–3328. JMLR.org, 2017
2017
-
[271]
From statistical to causal learning, 2022
Bernhard Schölkopf and Julius von Kügelgen. From statistical to causal learning, 2022
2022
-
[272]
Probing Pretrained Models of Source Code , November 2022
Sergey Troshin and Nadezhda Chirkova. Probing Pretrained Models of Source Code , November 2022. arXiv:2202.08975 [cs]
2022 arXiv
-
[273]
Explainable ai for se: Challenges and future directions
Chakkrit Tantithamthavorn, Jürgen Cito, Hadi Hemmati, and Satish Chandra. Explainable ai for se: Challenges and future directions. IEEE Software , 40(3):29--33, 2023
2023
-
[274]
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4593--4601, Florence, Italy, July 2019. Association for Computational Linguistics
2019
-
[275]
Wikipedia dataset, 2025
TensorFlow Datasets . Wikipedia dataset, 2025. Accessed: 2025-04-07
2025
-
[276]
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team . Theano: A Python framework for fast computation of mathematical expressions. arXiv e-prints , abs/1605.02688, 2016
2016 arXiv
-
[277]
Understanding the interplay between trust, reliability, and human factors in the age of generative ai
Simon Thorne. Understanding the interplay between trust, reliability, and human factors in the age of generative ai. International Journal of Simulation: Systems, Science & technology , May 2024
2024
-
[278]
Using pre-trained models to boost code review automation
Rosalia Tufano, Simone Masiero, Antonio Mastropaolo, Luca Pascarella, Denys Poshyvanyk, and Gabriele Bavota. Using pre-trained models to boost code review automation. In 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE) , pages 2291--2302, 2022
2022
-
[279]
and D Poshyvanyk
Watson C Bavota G Di Penta M White M Tufano M. and D Poshyvanyk. Deep Learning Similarities from Different Representations of Source Code . In Proceedings of the 15th IEEE/ACM Conference on Mining Software Repositories (MSR’18) , MSR '18, Gothenburg, Sweden, 2018
2018
-
[280]
When and Why Your Code Starts to Smell Bad (and Whether the Smells Go Away)
Michele Tufano, Fabio Palomba, Gabriele Bavota, Rocco Oliveto, Massimiliano Di Penta, Andrea De Lucia, and Denys Poshyvanyk. When and Why Your Code Starts to Smell Bad (and Whether the Smells Go Away) . IEEE Transactions on Software Engineering , 43(11):1063--1088, 2017
2017
-
[281]
Deeptest: Automated testing of deep-neural-network-driven autonomous cars
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. Deeptest: Automated testing of deep-neural-network-driven autonomous cars. In Proceedings of the 40th International Conference on Software Engineering , ICSE '18, pages 303--314, New York, NY, USA, 2018. ACM
2018
-
[282]
Towards automating code review activities
Rosalia Tufano, Luca Pascarella, Michele Tufano, Denys Poshyvanyk, and Gabriele Bavota. Towards automating code review activities. In 43rd International Conference on Software Engineering, ICSE '21 , 2021
2021
-
[283]
On learning meaningful code changes via neural machine translation
Michele Tufano, Jevgenija Pantiuchina, Cody Watson, Gabriele Bavota, and Denys Poshyvanyk. On learning meaningful code changes via neural machine translation. In Proceedings of the 41st International Conference on Software Engineering, ICSE 2019, Montreal, QC, Canada, May 25-3...
2019
-
[284]
On the localness of software
Zhaopeng Tu, Zhendong Su, and Premkumar Devanbu. On the localness of software. In Proceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering , FSE 2014, page 269–280, New York, NY, USA, 2014. Association for Computing Machinery
2014
-
[285]
Robert R. Tucci. Introduction to judea pearl's do-calculus, 2013
2013
-
[286]
Deep learning similarities from different representations of source code
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. Deep learning similarities from different representations of source code. In 2018 IEEE/ACM 15th International Conference on Mining Software Repositories (MSR) , pages 542--...
2018
-
[287]
An empirical investigation into learning bug-fixing patches in the wild via neural machine translation
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. An empirical investigation into learning bug-fixing patches in the wild via neural machine translation. In Proceedings of the 33rd ACM/IEEE International Conference on Auto...
2018
-
[288]
Learning How to Mutate Source Code from Bug-Fixes
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. Learning How to Mutate Source Code from Bug-Fixes . ICSME 2019 , pages 301--312, 2019
2019
-
[289]
An empirical study on learning bug-fixing patches in the wild via neural machine translation
Michele Tufano, Cody Watson, Gabriele Bavota, Massimiliano Di Penta, Martin White, and Denys Poshyvanyk. An empirical study on learning bug-fixing patches in the wild via neural machine translation. ACM Trans. Softw. Eng. Methodol. , 28(4):19:1--19:29, 2019
2019
-
[290]
Vera Liao, and Jennifer Wortman Vaughan
Helena Vasconcelos, Gagan Bansal, Adam Fourney, Q. Vera Liao, and Jennifer Wortman Vaughan. Generation probabilities are not enough: Uncertainty highlighting in ai code completions. ACM Trans. Comput.-Hum. Interact. , October 2024. Just Accepted
2024
-
[291]
Blei, and Alexander M
Keyon Vafa, Yuntian Deng, David M. Blei, and Alexander M. Rush. Rationales for sequential predictions
-
[292]
Gomez, ukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems , NeurIPs'17, page 6000–6010, Re...
2017
-
[293]
Neural Machine Translation with Byte - Level Subwords , December 2019
Changhan Wang, Kyunghyun Cho, and Jiatao Gu. Neural Machine Translation with Byte - Level Subwords , December 2019. arXiv:1909.03341
2019 arXiv
-
[294]
A systematic literature review on the use of deep learning in software engineering research
Cody Watson, Nathan Cooper, David Nader - Palacio, Kevin Moran, and Denys Poshyvanyk. A systematic literature review on the use of deep learning in software engineering research. CoRR , abs/2009.06520, 2020
2009 arXiv
-
[295]
Executable code actions elicit better llm agents, 2024
Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji. Executable code actions elicit better llm agents, 2024
2024
-
[296]
Yuxiang Wei, Olivier Duchenne, Jade Copet, Quentin Carbonneaux, Lingming Zhang, Daniel Fried, Gabriel Synnaeve, Rishabh Singh, and Sida I. Wang. Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution, 2025
2025
-
[297]
Herbsleb, Alexandra Holloway, and Scott Davidoff
David Gray Widder, Laura Dabbish, James D. Herbsleb, Alexandra Holloway, and Scott Davidoff. Trust in collaborative automation in high stakes software engineering work: A case study at nasa. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems , CHI ...
2021
-
[298]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...
2020
-
[299]
Transparency: Motivations and Challenges , page 23–40
Adrian Weller. Transparency: Motivations and Challenges , page 23–40. Springer-Verlag, Berlin, Heidelberg, 2022
2022
-
[300]
Deep Representations for Software Engineering
M White. Deep Representations for Software Engineering . In Proceedings of the 37th IEEE/ACM International Conference on Software Engineering (ICSE'15) , volume 2 of ICSE '15 , pages 781--783, 5 2015
2015
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.