REVIEW 4 major objections 4 minor 68 references
Python-specific mutants expose gaps in high-coverage suites
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 07:45 UTC pith:LQ2IDGZK
load-bearing objection PyTation's Python-specific mutation operators are a real step forward, but the cross-kill metric in Table 5 contradicts its own symmetric definition, putting the headline complementarity claim on shaky ground. the 4 major comments →
Hybrid Fault-Driven Mutation Testing for Python
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that a set of seven mutation operators, each designed to reintroduce a prevalent Python anti-pattern, produces mutants that behave differently under test execution from those of general-purpose mutation tools, thereby exposing test-suite inadequacies that line coverage and generic operators miss. On 13 open-source projects averaging 93% line coverage, the tool generated an average of 309 mutants per project with a mean mutation score of 88%. On average 69% of these mutants were not dynamically subsumed by any mutant from a general-purpose baseline, the cross-kill rate between the two tools' mutants was 3.52%, and the test-overlap ratio was 9.24%—figures the autho
What carries the argument
The load-bearing machinery is the seven anti-pattern operators: remove function argument, remove conversion function, remove element from container, remove expression from condition, change used attribute, remove attribute access, and remove method call. These are applied through a hybrid analysis: static AST traversal detects syntactically identifiable targets (container literals, compound conditions), while runtime instrumentation detects context-sensitive targets (optional arguments, attribute accesses, method calls) that only exist with concrete values. Coverage data prunes unreachable candidates, and type-compatibility and trace-comparison heuristics prune likely-equivalent mutants. The
Load-bearing premise
The seven operators are assumed to encode the most frequent real-world Python faults because they come from an empirical study of bugs, but the paper does not measure how often those patterns actually occur in its 13 benchmarks, so the 'realistic Python-specific fault' claim rests on an unverified prevalence assumption.
What would settle it
Run the seven operators on a fresh sample of Python projects and count how often the targeted anti-patterns actually occur in the code; if they are rare, the 'realistic fault' premise collapses. Alternatively, run a control where mutations are chosen by the same hybrid analysis but with arbitrary attributes, arguments, or container elements rather than the anti-pattern set: if the control produces the same uniqueness and cross-kill numbers, the operator selection is not what carries the result.
If this is right
- High line coverage does not imply behavioural adequacy: mutants survived in projects with 93% average line coverage, especially in idiomatic Python patterns such as generator expressions and debug-path code.
- A test suite can pass a general-purpose mutation baseline and still miss Python-specific faults: only 3.52% of the tool's mutants were killed by the same tests as baseline mutants, and 69% were unique.
- Dynamic-analysis heuristics can suppress equivalent mutants without losing the fault signal: the measured equivalent-mutant rate was 1.61% on average, lower than published rates for general-purpose tools.
- The per-mutant runtime cost (about 25 seconds per mutant) is comparable to the general-purpose baseline, so adding a Python-specific fault model does not by itself make mutation testing impractical.
- Recurring surviving-mutant patterns point to missing behavioural assertions, not missing execution: tests reach the lines but do not check the resulting state, message, or side effect.
Where Pith is reading between the lines
- If the complementarity result generalises, the operator set could become self-updating: mine a project's own bug-fix history to derive its characteristic anti-patterns instead of relying on a fixed catalogue.
- A controlled ablation—random attribute swaps or random argument removals applied through the same hybrid analysis—would reveal whether the empirical grounding of the seven operators, rather than the dynamic plumbing, drives the observed uniqueness.
- The surviving mutants in high-coverage projects suggest that what is missing is often an assertion oracle, not more execution: tests need to assert behavioural properties such as timezone, encoding, serialisation, or filtering, not just reach lines.
- The same operator template could transfer to other dynamically-typed languages, provided the anti-pattern catalogue is re-derived from that language's own bug studies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents PyTation, a mutation testing tool for Python that introduces seven operators derived from Python anti-patterns, uses a hybrid of static AST analysis and dynamic tracing (DynaPyt) to identify and prune mutation candidates, and evaluates the approach on 13 open-source Python projects. The central claim is that PyTation's mutants complement those of the general-purpose tool Cosmic Ray, as evidenced by a high proportion of unique mutants, a low cross-kill rate, and a low test-overlap ratio, while also uncovering inadequacies in high-coverage test suites. The paper further claims that dynamic heuristics keep the equivalent-mutant rate low.
Significance. If the results hold, the paper makes a useful contribution: it targets Python-specific fault patterns that general-purpose operators miss, grounds operators in an external single-statement bug study (PySStuBs), and offers an open-source implementation with reproducibility scripts. The evaluation against Cosmic Ray on real projects is a strength, and the dynamic-analysis pruning idea is timely. However, the headline complementarity evidence is currently undermined by an internal inconsistency in the cross-kill metric, and the equivalent-mutant estimate rests on a small manual sample without agreement analysis. These issues are fixable, but the quantitative support for the main claim needs to be repaired before the paper can be accepted.
major comments (4)
- [§5.1.3, §5.2.2, Table 5] The cross-kill rate is defined symmetrically as |K_A∩K_B|/(|K_A|+|K_B|-|K_A∩K_B|). A symmetric set expression cannot yield different values for A and B, yet Table 5 reports two cross-kill columns that differ (e.g., pyjwt: 4.95 vs 10.20; pyquery: 13.30 vs 23.40). The abstract's 'low cross-kill rate' (mean 3.52%) is one of these columns. Either the implementation used an asymmetric definition or the columns are mislabeled/shifted; the mean row also shows 9.24 in a position that conflicts with the text's reported test-overlap mean of 9.24%, so column alignment needs verification. The complementarity evidence must be recomputed with the correct symmetric metric before the central claim can be assessed.
- [§5.1.2, §5.2.1] The paper claims an average equivalent-mutant rate of 1.61% (range 0–4.74%), but this is extrapolated from manual classification of only 20% of RQ1 mutants. The manuscript says two independent non-author examiners reviewed this sample, but it reports no inter-rater agreement (e.g., Cohen's kappa), no per-project/per-operator sample sizes, and no confidence bounds. Since 'few equivalent mutants' is an explicit contribution and underpins the heuristic-pruning claim, the estimate needs a reproducibility/agreement analysis or should be stated with appropriate caution.
- [§5.2.2 (RQ2.3)] The text states: 'No statistically significant association (p>0.05) was found, supporting that PyTation's mutants elicit distinct test behaviours.' This is logically inverted: p>0.05 means the test failed to reject the null hypothesis; it does not provide evidence of absence of association. This is especially problematic with small mutant counts and 2×2 tables. Please replace this statement with effect sizes and confidence intervals (e.g., Cramér's V with a confidence interval) or an equivalence test, and avoid interpreting non-significance as evidence of independence.
- [§2, Table 3] The paper grounds the seven operators in PySStuBs [29] and claims they target prevalent Python-specific faults, but the evaluation does not quantify how often the target anti-patterns occur in the 13 benchmarks. RemElCont and RemExpCond produce on average only 2 mutants per project, with several projects having zero such candidates (e.g., schedule, pyjwt, funcy, wtforms, marshmallow, graphene, praw). This sparse application makes it difficult to assess whether the operator set is representative of prevalent faults and whether the complementarity results generalize. Please report per-operator candidate counts and prevalence in the 13 projects, or explain how the operator selection still supports the 'realistic, Python-specific fault' premise.
minor comments (4)
- [Table 5] The table formatting is very hard to parse. For example, the 'schedule' row shows a 'Test Overlap' value of 178, which cannot be a percentage; several values appear to have lost decimal points or shifted columns. Please reformat the table with explicit subcolumn headers for PyTa/CoRa and verify all values.
- [Table 5] The table lists '#Invalid' as a single column, but the text says this is for Cosmic Ray, while Table 4 reports PyTation's invalid mutants. Make explicit in the caption or header which tool each invalid count refers to.
- [§5.2.2] The summary sentence 'On average, PyTation generated 248 mutants, while Cosmic Ray produced 2,537' should state that this is computed only on the 11 projects where Cosmic Ray ran, not all 13 projects.
- [References] References [23] and [24] both point to the same MutPy URL; [24] appears to be intended for Mutmut or a different tool. Please correct.
Circularity Check
No significant circularity: operators are externally grounded, comparisons use an external baseline, and no central result reduces to its inputs by definition.
full rationale
The derivation chain is self-contained. The seven operators are justified by the external PySStuBs bug-pattern study ([29], Section 2: 'They are empirically grounded in a large-scale study [29]... we selected frequent Python-specific categories'), not by PyTation's own outputs. Mutation candidates are identified via AST traversal and runtime interception, and the test-suite dependence is explicitly limited to placement, not killability (Section 3.1.3: 'ensuring executability without increasing killability'). Complementarity claims are evaluated against Cosmic Ray, an external state-of-the-art tool, using mutation scores, dynamic subsumption, cross-kill, and test-overlap metrics computed from recorded pass/fail outcomes (Section 5.1.3). No parameter is fitted to the claimed conclusion and then renamed a prediction, and no load-bearing premise rests on a self-citation. The Table 5 discrepancy between the symmetric cross-kill definition and the two reported cross-kill columns is a likely computational or reporting inconsistency and a genuine threat to the headline evidence, but it is not circularity: the claimed low cross-kill rate is not equal to the metric definition by construction. Similarly, designing operators from known bug categories and then measuring whether test suites kill such mutants is standard mutation-testing methodology, not a self-referential reduction.
Axiom & Free-Parameter Ledger
free parameters (1)
- Equivalent-mutant review sample fraction =
20% (minimum 10 mutants per operator per project)
axioms (6)
- domain assumption Python's dynamic typing, flexible argument passing, and runtime attribute resolution create fault scenarios distinct from statically typed languages.
- domain assumption The bug categories selected from PySStuBs [29] are frequent and representative of real Python faults.
- domain assumption A surviving mutant indicates test-suite inadequacy; a killed mutant indicates real fault-revealing capability.
- ad hoc to paper Manual equivalence classification on a 20% sample provides an accurate estimate of the equivalent-mutant rate.
- domain assumption DynaPyt instrumentation and coverage-guided pruning do not change program semantics in ways that bias kill outcomes.
- domain assumption Cosmic Ray is a fair and representative baseline for general-purpose Python mutation testing.
read the original abstract
Mutation testing is an effective technique for assessing the effectiveness of test suites by systematically injecting artificial faults into programs. However, existing mutation testing techniques fall short in capturing many types of common faults in dynamically typed languages like Python. In this paper, we introduce a novel set of seven mutation operators that are inspired by prevalent anti-patterns in Python programs, designed to complement the existing general-purpose operators and broaden the spectrum of simulated faults. We propose a mutation testing technique that utilizes a hybrid of static and dynamic analyses to mutate Python programs based on these operators while minimizing equivalent mutants. We implement our approach in a tool called PyTation and evaluate it on 13 open-source Python applications. Our results show that PyTation generates mutants that complement those from general-purpose tools, exhibiting distinct behaviour under test execution and uncovering inadequacies in high-coverage test suites. We further demonstrate that PyTation produces a high proportion of unique mutants, a low cross-kill rate, and a low test overlap ratio relative to baseline tools, highlighting its novel fault model. PyTation also incurs few equivalent mutants, aided by dynamic analysis heuristics.
Figures
Reference graph
Works this paper leans on
-
[1]
Sixty North AS. 2024. GitHub - sixty-north/cosmic-ray: Mutation testing for Python — github.com. https://github.com/sixty-north/cosmic-ray. [Accessed 03-02-2024]
2024
-
[2]
Michael Baer, Norbert Oster, and Michael Philippsen. 2020. MutantDistiller: Using Symbolic Execution for Automatic Detection of Equivalent Mutants and Generation of Mutant Killing Tests. In2020 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, Porto, Portugal, 294–303. doi:10.1109/ICSTW50294.2020.00055
arXiv 2020
-
[3]
Leonardo Bottaci. 2010. Type Sensitive Application of Mutation Operators for Dynamically Typed Programs. In2010 Third International Conference on Software Testing, Verification, and Validation Workshops. IEEE, Paris, France, 126–131. doi:10.1109/ICSTW.2010.56
-
[4]
1977.Users guide to the Pilot mutation system
Timothy Budd and Frederick G Sayward. 1977.Users guide to the Pilot mutation system. Technical Report 114. Department of Computer Science, Yale University
1977
-
[5]
Henry Coles, Thomas Laurent, Christopher Henard, Mike Papadakis, and An- thony Ventresque. 2016. PIT: a practical mutation testing tool for Java (demo). In Proceedings of the 25th International Symposium on Software Testing and Analysis (Saarbrücken, Germany)(ISSTA 2016). Association for Computing Machinery, New York, NY, USA, 449–452. doi:10.1145/2931037.2948707
arXiv 2016
-
[6]
1999.Mathematical methods of statistics
Harald Cramér. 1999.Mathematical methods of statistics. Vol. 9. Princeton university press
1999
-
[7]
Francisco Dalton, Márcio Ribeiro, Gustavo Pinto, Leo Fernandes, Rohit Gheyi, and Baldoino Fonseca. 2020. Is Exceptional Behavior Testing an Exception? An Empirical Assessment Using Java Automated Tests. InProceedings of the 24th International Conference on Evaluation and Assessment in Software Engineering (Trondheim, Norway)(EASE ’20). Association for Com...
arXiv 2020
-
[8]
Joost CF De Winter, Samuel D Gosling, and Jeff Potter. 2016. Comparing the Pearson and Spearman correlation coefficients across distributions and sample sizes: A tutorial using simulations and empirical data.Psychological methods21, 3 (2016), 273
2016
-
[9]
P. Delgado-Pérez, S. Segura, and I. Medina-Bulo. 2017. Assessment of C++ Object Oriented Mutation Operators: A Selective Mutation Approach.Information and Software Technology81 (2017), 169–184. doi:10.1016/j.infsof.2016.05.012
-
[10]
Pedro Delgado-Pérez, Inmaculada Medina-Bulo, Francisco Palomo-Lozano, Anto- nio García-Domínguez, and Juan José Domínguez-Jiménez. 2017. Assessment of class mutation operators for C++ with the MuCPP mutation system.Information and Software Technology81 (2017), 169–184. doi:10.1016/j.infsof.2016.07.002
-
[11]
Anna Derezinska and Karol Kowalski. 2011. Object-Oriented Mutation Applied in Common Intermediate Language Programs Originated from C#. In2011 IEEE Fourth International Conference on Software Testing, Verification and Validation Workshops. IEEE, Berlin, Germany, 342–350. doi:10.1109/ICSTW.2011.54
-
[12]
Anna Derezińska and Konrad Hałas. 2014. Analysis of Mutation Operators for the Python Language. InProceedings of the Ninth International Conference on Dependability and Complex Systems DepCoS-RELCOMEX, June 30 – July 4, 2014, Brunów, Poland, Wojciech Zamojski, Jacek Mazurkiewicz, Jarosław Sugier, Tomasz Walkowiak, and Janusz Kacprzyk (Eds.). Springer Inte...
2014
-
[13]
Kadiatou Diallo, Zizhao Chen, W. Eric Wong, and Shou-Yu Lee. 2024. An Analysis and Comparison of Mutation Testing Tools for Python. In11th International Conference on Dependable Systems and Their Applications, DSA 2024, Taicang, Suzhou, China, November 2-3, 2024. IEEE, 161–169. doi:10.1109/DSA63982.2024. 00030
arXiv 2024
-
[15]
Antonia Estero-Botaro, F Palomo-Lozano, and Inmaculada Medina-Bulo. 2008. Mutation operators for WS-BPEL 2.0. In21th International Conference on Software & Systems Engineering and their Applications. INCOSE, Paris, France
2008
-
[16]
Antonia Estero-Botaro, Francisco Palomo-Lozano, and Inmaculada Medina-Bulo
- [17]
-
[18]
Daniel Fortunato, José Campos, and Rui Abreu. 2022. QMutPy: a mutation testing tool for Quantum algorithms and applications in Qiskit. InISSTA ’22: 31st ACM SIGSOFT International Symposium on Software Testing and Analysis, Virtual Event, South Korea, July 18 - 22, 2022, Sukyoung Ryu and Yannis Smaragdakis (Eds.). ACM, 797–800. doi:10.1145/3533767.3543296
arXiv 2022
-
[19]
Milos Gligoric, Sandro Badame, and Ralph Johnson. 2011. SMutant: a tool for type-sensitive mutation testing in a dynamic language. InProceedings of the 19th ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil Alimadadi and Gharachorlu ACM SIGSOFT Symposium and the 13th European Conference on Foundations of Software Engineering(Szeged, Hungary)(ESEC/FSE ’1...
arXiv 2011
-
[20]
Guerino, Pedro H
Lucas R. Guerino, Pedro H. Kuroishi, Alexandre C. R. Paiva, and Américo M. R. Vincenzi. 2024. Static and Dynamic Comparison of Mutation Testing Tools for Python. InProceedings of the XXIII Brazilian Symposium on Software Quality. 199–209
2024
-
[21]
Lucca Renato Guerino, Pedro Henrique Kuroishi, Ana Cristina Ramada Paiva, and Auri Marcelo Rizzo Vincenzi. 2024. Static and Dynamic Comparison of Mutation Testing Tools for Python. InProceedings of the XXIII Brazilian Symposium on Software Quality, SBQS 2024, Salvador, Bahia, Brazil, November 5-8, 2024, Ivan Machado, José Carlos Maldonado, Tayana Conte, E...
arXiv 2024
-
[22]
Jan Hauke and Tomasz Kossowski. 2011. Comparison of values of Pearson’s and Spearman’s correlation coefficients on the same sets of data.Quaestiones geographicae30, 2 (2011), 87–93
2011
-
[23]
Konrad Hałas. 2024. GitHub - mutpy/mutpy: MutPy is a mutation testing tool for Python 3.x source code — github.com. https://github.com/mutpy/mutpy. [Accessed 03-02-2024]
2024
-
[24]
Anders Hovmöller. 2024. GitHub - mutpy/mutpy: MutPy is a mutation testing tool for Python 3.x source code — github.com. https://github.com/mutpy/mutpy. [Accessed 03-02-2024]
2024
-
[25]
Laura Inozemtseva and Reid Holmes. 2014. Coverage is not strongly correlated with test suite effectiveness. InProceedings of the 36th International Conference on Software Engineering(Hyderabad, India)(ICSE 2014). Association for Computing Machinery, New York, NY, USA, 435–445. doi:10.1145/2568225.2568271
arXiv 2014
-
[26]
Yue Jia and Mark Harman. 2011. An Analysis and Survey of the Development of Mutation Testing.IEEE Transactions on Software Engineering37, 5 (2011), 649–678. doi:10.1109/TSE.2010.62
-
[27]
Ernst, Reid Holmes, and Gordon Fraser
René Just, Darioush Jalali, Laura Inozemtseva, Michael D. Ernst, Reid Holmes, and Gordon Fraser. 2014. Are mutants a valid substitute for real faults in software testing?. InProceedings of the 22nd ACM SIGSOFT International Symposium on Foundations of Software Engineering(Hong Kong, China)(FSE 2014). Association for Computing Machinery, New York, NY, USA,...
-
[28]
René Just, Bob Kurtz, and Paul Ammann. 2017. Inferring mutant utility from program context. InProceedings of the 26th ACM SIGSOFT International Sym- posium on Software Testing and Analysis(Santa Barbara, CA, USA)(ISSTA 2017). Association for Computing Machinery, New York, NY, USA, 284–294. doi:10.1145/3092703.3092732
arXiv 2017
-
[29]
Kamienski, Luisa Palechor, Cor-Paul Bezemer, and Abram Hindle
Arthur V. Kamienski, Luisa Palechor, Cor-Paul Bezemer, and Abram Hindle. 2021. PySStuBs: Characterizing Single-Statement Bugs in Popular Open-Source Python Projects. In2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR). IEEE, Madrid, Spain, 520–524. doi:10.1109/MSR52588.2021. 00066
arXiv 2021
-
[30]
Shuhei Kimura, Keisuke Hotta, Yoshiki Higo, Hiroshi Igaki, and Shinji Kusumoto
-
[31]
Marinos Kintis, Mike Papadakis, Andreas Papadopoulos, Evangelos Valvis, Nicos Malevris, and Yves Le Traon. 2017. How effective are mutation testing tools? An empirical analysis of Java mutation testing tools with manual analysis and real faults.Empirical Software Engineering23 (2017), 2426 – 2463. https://api. semanticscholar.org/CorpusID:13723714
2017
-
[32]
Bob Kurtz, Paul Ammann, Márcio Eduardo Delamaro, Jeff Offutt, and Lin Deng
-
[33]
Bob Kurtz, Paul Ammann, and Jeff Offutt. 2015. Static analysis of mutant subsump- tion. InEighth IEEE International Conference on Software Testing, Verification and Validation, ICST 2015 Workshops, Graz, Austria, April 13-17, 2015. IEEE Computer Society, 1–10. doi:10.1109/ICSTW.2015.7107454
arXiv 2015
-
[34]
Kaufman, Ryan Featherman, Harrison Potter, Amir Madadi, and René Just
Brittany Kushigian, Steven J. Kaufman, Ryan Featherman, Harrison Potter, Amir Madadi, and René Just. 2024. Equivalent Mutants in the Wild: Identifying and Efficiently Suppressing Equivalent Mutants for Java Programs. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA). ACM, 654–665
2024
-
[35]
Mutant Subsumption Graphs. InSeventh IEEE International Conference on Software Testing, Verification and Validation, ICST 2014 Workshops Proceedings, March 31 - April 4, 2014, Cleveland, Ohio, USA. IEEE Computer Society, 176–185. doi:10.1109/ICSTW.2014.20
-
[36]
Yu-Seung Ma and Sang-Woon Kim. 2016. Mutation testing cost reduction by clustering overlapped mutants.Journal of Systems and Software115 (2016), 18–30
2016
-
[37]
Shabnam Mirshokraie, Ali Mesbah, and Karthik Pattabiraman. 2013. Efficient JavaScript Mutation Testing. In2013 IEEE Sixth International Conference on Soft- ware Testing, Verification and Validation. IEEE, Luxembourg City, Luxembourg, 74–83. doi:10.1109/ICST.2013.23
-
[38]
Jukka Lehtosalo. 2024. GitHub - python/mypy: Optional static typing for Python. https://github.com/python/mypy. [Accessed 03-02-2024]
2024
-
[39]
Jefferson Offutt, Yu-Seung Ma, and Yong Rae Kwon
A. Jefferson Offutt, Yu-Seung Ma, and Yong Rae Kwon. 2004. An experimental mutation system for Java.ACM SIGSOFT Softw. Eng. Notes29 (2004), 1–4
2004
-
[40]
A Jefferson Offutt and Jie Pan. 1997. Automatically detecting equivalent mutants and infeasible paths.Software testing, verification and reliability7, 3 (1997), 165–192
1997
-
[41]
Akbar Siami Namin and James H. Andrews. 2009. The influence of size and coverage on test suite effectiveness. InProceedings of the Eighteenth International Symposium on Software Testing and Analysis(Chicago, IL, USA)(ISSTA ’09). Association for Computing Machinery, New York, NY, USA, 57–68. doi:10.1145/ 1572272.1572280
arXiv 2009
-
[42]
Jefferson Offutt and Ronald H
A. Jefferson Offutt and Ronald H. Untch. 2001.Mutation 2000: Uniting the Orthog- onal. Kluwer Academic Publishers, USA, 34–44
2001
-
[43]
Wonseok Oh and Hakjoo Oh. 2022. PyTER: effective program repair for Python type errors. InProceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering(Singapore, Singapore)(ESEC/FSE 2022). Association for Computing Machinery, New York, NY, USA, 922–934. doi:10.1145/3540250.3549130
arXiv 2022
-
[44]
Jefferson Offutt and Jie Pan
A. Jefferson Offutt and Jie Pan. 1997. Automatically detecting equivalent mutants and infeasible paths.Software Testing, Verification and Reliability7, 3 (1997), 165–
1997
-
[45]
Mike Papadakis, Marcio Delamaro, and Yves Le Traon. 2014. Mitigating the effects of equivalent mutants with mutant classification strategies.Science of Computer Programming95 (2014), 298–319. doi:10.1016/j.scico.2014.05.012 Special Section: ACM SAC-SVT 2013 + Bytecode 2013
-
[46]
Mike Papadakis, Marinos Kintis, Jie Zhang, Yue Jia, Yves Le Traon, and Mark Harman. 2019. Chapter Six - Mutation Testing Advances: An Analysis and Survey. InAdvances in Computers, Atif M. Memon (Ed.). Advances in Computers, Vol. 112. Elsevier, Amsterdam, Netherlands, 275–378. doi:10.1016/bs.adcom.2018.03.015
-
[47]
Mike Papadakis and Yves Le Traon. 2014. Effective fault localization via muta- tion analysis: a selective mutation approach. InProceedings of the 29th Annual ACM Symposium on Applied Computing(Gyeongju, Republic of Korea)(SAC ’14). Association for Computing Machinery, New York, NY, USA, 1293–1300. doi:10.1145/2554850.2554978
arXiv 2014
-
[48]
Kai Pan, Sunghun Kim, and E. James Whitehead Jr. 2009. Toward an un- derstanding of bug fix patterns.Empir. Softw. Eng.14, 3 (2009), 286–315. doi:10.1007/S10664-008-9077-5
-
[49]
Papadakis, J
M. Papadakis, J. Zhang, and Y. Le Traon. 2014. Minimizing Mutation Testing Cost via Test Case Prioritization.Proceedings of the 2014 International Conference on Software Engineering1 (2014), 1190–1201
2014
-
[51]
Goran Petrovic, Marko Ivankovic, Bob Kurtz, Paul Ammann, and René Just. 2018. An industrial application of mutation testing: Lessons, challenges, and research directions. In2018 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, 47–53
2018
-
[52]
Mike Papadakis, Nicos Malevris, and Marinos Kintis. 2010. Mutation Testing Strategies - A Collateral Approach.. InICSOFT 2010 - Proceedings of the 5th International Conference on Software and Data Technologies, Vol. 2. SciTePress, Athens, Greece, 325–328
2010
-
[53]
Cedric Richter and Heike Wehrheim. 2022. Learning Realistic Mutations: Bug Creation for Neural Bug Detectors. In2022 IEEE Conference on Software Testing, Verification and Validation (ICST). IEEE, Valencia, Spain, 162–173. doi:10.1109/ ICST53961.2022.00027
arXiv 2022
-
[54]
Vitalis Salis, Thodoris Sotiropoulos, Panos Louridas, Diomidis Spinellis, and Dimitris Mitropoulos. 2021. PyCG: Practical Call Graph Generation in Python. InProceedings of the 43rd International Conference on Software Engineering (ICSE ’21). IEEE Press, Madrid, Spain, 1646–1657. doi:10.1109/ICSE43902.2021.00146
arXiv 2021
-
[55]
A. B. Sanchez, J. A. Parejo, S. Segura, A. Duran, and M. Papadakis. 2024. Muta- tion Testing in Practice: Insights from Open-Source Software Developers.IEEE Transactions on Software Engineering50, 01 (mar 2024), 1–15. doi:10.1109/TSE. 2024.3377378
arXiv 2024
-
[56]
Alessandro Viola Pizzoleto, Fabiano Cutigi Ferrari, Jeff Offutt, Leo Fernandes, and Márcio Ribeiro. 2020. A Systematic Literature Review of Techniques and Metrics to Reduce the Cost of Mutation Testing.Journal of Systems and Software 157 (2020), 110398. doi:10.1016/j.jss.2019.110398
arXiv 2020
-
[57]
David Schuler and Andreas Zeller. 2013. Covering and Uncovering Equivalent Mutants.Software Testing, Verification & Reliability23, 5 (2013), 353–374. doi:10. 1002/stvr.1473 Hybrid Fault-Driven Mutation Testing for Python ICSE ’26, April 12–18, 2026, Rio de Janeiro, Brazil
2013
-
[58]
Florian Schwander, Rahul Gopinath, and Andreas Zeller. 2021. Inducing Subtle Mutations with Program Repair. In2021 IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW). IEEE, Porto, Portugal, 25–34. doi:10.1109/ICSTW52544.2021.00018
arXiv 2021
-
[59]
2004.Enhancing file system integrity through checksums
Gopalan Sivathanu, Charles P Wright, and Erez Zadok. 2004.Enhancing file system integrity through checksums. Technical Report. Citeseer
2004
-
[60]
David Schuler and Andreas Zeller. 2009. Javalanche: efficient mutation testing for Java. InProceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering(Amsterdam, The Netherlands)(ESEC/FSE ’09). Association for Com- puting Machinery, New York, NY, USA, 297–298...
arXiv 2009
-
[61]
Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok, Haodi Qi, Jack Phan, Qijin Tay, Constance Tan, Fiona Wee, Jodie Ethelda Tan, Yuheng Yieh, Brian Goh, Ferdian Thung, Hong Jin Kang, Thong Hoang, David Lo, and Eng Lieh Ouh. 2020. BugsInPy: a database of existing bugs in Python programs to enable controlled testing and debugging studies. InProceedings of the...
2020
-
[62]
Ratnadira Widyasari, Sheng Qin Sim, Camellia Lok, Haodi Qi, Jack Phan, Qijin Tay, Constance Tan, Fiona Wee, Jodie Ethelda Tan, Yuheng Yieh, Brian Goh, Ferdian Thung, Hong Jin Kang, Thong Hoang, David Lo, and Eng Lieh Ouh. 2024. BugsInPy: A Database of Existing Bugs in Python Programs to Enable Controlled Testing and Debugging Studies.CoRRabs/2401.15481 (2...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2401.15481 2024
-
[63]
Xiangjuan Yao, Mark Harman, and Yue Jia. 2014. A study of equivalent and stubborn mutation operators using human analysis of equivalence. InProceedings of the 36th international conference on software engineering. 919–930
2014
-
[64]
Philipp Straubinger, Marvin Kreis, Stephan Lukasczyk, and Gordon Fraser. 2025. Mutation Testing via Iterative Large Language Model-Driven Scientific Debug- ging. InIEEE International Conference on Software Testing, Verification and Valida- tion, ICST 2025 - Workshops, Naples, Italy, March 31 - April 4, 2025. IEEE, 358–367. doi:10.1109/ICSTW64639.2025.10962485
arXiv 2025
-
[65]
Zejun Zhang, Zhenchang Xing, Dehai Zhao, Qinghua Lu, Xiwei Xu, and Liming Zhu. 2024. Hard to Read and Understand Pythonic Idioms? DeIdiom and Explain Them in Non-Idiomatic Equivalent Code. InProceedings of the IEEE/ACM 46th International Conference on Software Engineering(Lisbon, Portugal)(ICSE ’24). Association for Computing Machinery, New York, NY, USA,...
arXiv 2024
-
[66]
Hong Zhu, Patrick A. V. Hall, and John H. R. May. 1997. Software unit test coverage and adequacy.Comput. Surveys29, 4 (dec 1997), 366–427. doi:10.1145/ 267580.267590
arXiv 1997
-
[68]
Lingming Zhang, Milos Gligoric, Darko Marinov, and Sarfraz Khurshid. 2013. Operator-based and random mutant selection: Better together. In2013 28th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, Palo Alto, CA, USA, 92–102. doi:10.1109/ASE.2013.6693070
arXiv 2013
-
[192]
doi:10.1002/(SICI)1099-1689(199709)7:3<165::AID-STVR143>3.0.CO;2-U
-
[2010]
In2010 Third International Conference on Software Testing, Verification, and Validation Workshops
Quantitative Evaluation of Mutation Operators for WS-BPEL Composi- tions. In2010 Third International Conference on Software Testing, Verification, and Validation Workshops. IEEE, Paris, France, 142–150. doi:10.1109/ICSTW.2010.36
-
[2014]
Does return null matter?. In2014 Software Evolution Week - IEEE Conference on Software Maintenance, Reengineering, and Reverse Engineering, CSMR-WCRE 2014, Antwerp, Belgium, February 3-6, 2014, Serge Demeyer, Dave W. Binkley, and Filippo Ricca (Eds.). IEEE Computer Society, 244–253. doi:10.1109/CSMR- WCRE.2014.6747176
arXiv 2014
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.