REVIEW 4 major objections 5 minor 60 references
Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Flaky tests frequently fail in clusters that share root causes, so one repair can fix many tests at once.
desk verdict Solid novel empirical study of flaky-test co-occurrence; the 75% prevalence is an upper bound, but the phenomenon is real and deserves serious peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Jaccard distance over failure sets. For flaky tests $A$ and $B$, let $A$ and $B$ also denote the sets of test-suite run IDs in which each test fails; then $J(A,B) = 1 - |A \cap B| / |A \cup B|$, where 0 means the two tests always fail in the same runs and 1 means they never co-occur. Agglomerative clustering over these distances, with the clustering cutoff chosen by maximizing mean silhouette score subject to a minimum of 0.6, converts raw co-occurrence into named clusters. Two supporting mechanisms carry the rest: a prediction pipeline in which character- and set-based string distances on test names and tokenized source, plus a hierarchy distance on package/class paths, feed tree-ensemble models that predict $J(A,B)$ and cluster membership; and a manual inspection protocol in which sampled stack traces and error messages are reviewed by multiple inspectors and resolved through negotiated agreement to assign cause themes to clusters. Together these carry the argument from "these tests fail in the same runs" to "these tests fail for a shared, repairable reason."
What would settle it
For each of the 45 clusters, count the member tests whose failure stack traces end in the same root exception; if most clusters lack a single shared root exception, the claim that co-occurrence implies a shared fixable cause fails. A stronger test would stub out the suspected network path or external dependency for a cluster and check whether all member tests stop failing together.
Extended reading notes
Core claim
The paper's central discovery is that flaky-test failures co-occur in structured clusters, not as isolated events. On a dataset of 10,000 test-suite runs for each of 24 Java projects, of which 22 contain at least one flaky test and 810 flaky tests in total, the authors compute the Jaccard distance between every pair of flaky tests' sets of failing run IDs and cluster the tests agglomeratively, choosing the distance threshold that maximizes the mean silhouette score. This produces 45 non-singleton clusters in 10 projects; 606 of the 810 flaky tests (75%) fall into a cluster, the mean cluster size is 13.5 tests, and clusters span 2.9 test classes on average. The paper argues that these co-occurrence clusters correspond to shared root causes, and manual inspection of stack traces identifies intermittent networking issues and instabilities in external dependencies as the dominant causes. It then shows that an extra-trees model trained only on static distance measures between test names and tokenized source code predicts pairwise Jaccard distances with mean $R^2 = 0.74$ and cluster membership with mean Matthews correlation coefficient 0.74, so much of the cluster structure is recoverable without exhaustive reruns. The discovery, if correct, overturns the default assumption that flaky failures are isolated and makes systemic flakiness a first-class target for repair and tooling.
Load-bearing premise
The load-bearing premise is that two flaky tests failing in the same test-suite run share one repairable root cause; the paper itself acknowledges that random chance, a common external outage, or an artifact of the run environment could also produce co-occurrence.
Editorial extensions
If this is right
- One repair action aimed at a cluster's shared cause can fix, on average, 13.5 flaky tests at once, directly attacking the measured cost of flaky-test repair.
- Impact studies of flaky tests on fault localization, mutation testing, and automated program repair should be repeated with co-occurring clusters; the paper argues that ignoring them misrepresents the real effect.
- Previous root-cause taxonomies that rank concurrency and asynchronous issues highest for individual flaky tests may be skewed, because the dominant cluster-level causes are networking and external-dependency instability.
- Static analysis alone, using test names, code tokens, and package/class hierarchy, can predict a substantial share of the cluster structure, so teams can triage systemic flakiness without 10,000 reruns.
- Because clusters span about 2.9 test classes on average, developers should treat test-class decoupling and isolation from environmental variability as protective measures.
Reading between the lines
- Editorial extension: the dataset rebooted the machine between test-suite runs to isolate runs; a CI pipeline that reuses processes, caches, and network connections may show weaker or different clustering, so the 75% prevalence should be re-estimated in production-like settings.
- Editorial extension: if the hierarchy distance is the strongest predictor, then tests in the same package or class subtree tend to share failure conditions; a cheap, proactive risk signal would flag same-hierarchy tests that touch the network or external services before any rerun is done.
- Editorial extension: the cluster-then-fix view suggests a measurable triage workflow: group newly failing tests by co-occurrence in a sliding window of CI runs, fix the cluster's shared cause, and record how many tests are repaired per action; that yield can be compared against per-test debugging.
- Editorial extension: because the dominant causes are environmental, effective fixes may be infrastructure changes rather than test-code changes; adding a precondition that checks a server or directory before a test runs is a concrete mitigation that could be tested A/B in a repository.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the concept of 'systemic flakiness': flaky tests whose failures co-occur in the same test-suite runs and are assumed to share a root cause. Using the existing FlakeFlagger dataset of 10,000 runs across 24 Java projects (810 flaky tests), the authors apply agglomerative clustering to Jaccard distances between failing-run sets, report that 75% of flaky tests fall into 45 non-singleton clusters, train tree-based models to predict pairwise Jaccard distances and cluster membership from static test-code distance measures, and manually inspect stack traces to attribute cluster causes. The central claims are that systemic flakiness is widespread, that it can be predicted without full reruns, and that developers can fix multiple flaky tests by addressing shared root causes.
Significance. If the prevalence claim withstands scrutiny, the paper identifies a genuinely under-explored phenomenon with practical implications for flaky-test repair and for simulating flakiness in software-engineering experiments. Strengths include the use of a well-established, expensive 10,000-run dataset; the public replication package; the mixed-method design combining clustering, prediction, and manual inspection; and the negotiated-agreement protocol for qualitative analysis. The manual inspection gives credible evidence that some clusters share environmental causes, such as networking failures. However, the RQ1 prevalence figure is not tested against a null model of independent failures, and RQ2 lacks the baselines needed to interpret its predictive results; both issues are load-bearing for the paper's headline conclusions rather than cosmetic.
major comments (4)
- [Section 2.2, Table 1] The headline 75% prevalence figure is not tested against a null model in which each flaky test's failing runs are independent of every other test. The pipeline selects, per project, the distance threshold that maximizes mean silhouette and then discards projects whose best silhouette is below 0.6; with 810 flaky tests spread over 10,000 runs and low per-test failure counts, non-singleton high-silhouette clusters can arise by chance. I request a permutation baseline that preserves each flaky test's failure count and, ideally, each run's failure count, reruns the same clustering pipeline, and reports the distribution of the percentage of flaky tests in clusters. Without such a baseline, the statement that 75% of flaky tests 'belong to a cluster' cannot support the conclusion that systemic flakiness is widespread, because some clustering is expected under independence.
- [Section 2.5, Tables 4 and 5] The inference from co-occurrence to a shared fixable root cause is load-bearing for the practical payoff, but the evidence supports a weaker conclusion. The construct-validity section explicitly lists random chance as a possible contributor, and RQ3's dominant themes are networking (25/45 clusters) and external dependency (14/45) — common environmental conditions rather than code-level defects that a single developer action repairs. Table 5 reports Unknown for 14 clusters and lists mitigations such as Avoid Networking and Better Error Checking, which are not single-fix repairs. The authors should either provide evidence that a meaningful fraction of clusters trace to a specific fixable defect in the project's own code, or reframe the contribution as characterizing co-occurrence and its environmental causes rather than as 'fix one root cause, repair many tests.'
- [Section 2.3, Table 2] RQ2's evaluation lacks the baselines needed to interpret the reported R2 and MCC values. The regression target is a pairwise Jaccard distance, which is strongly related by construction to the name- and code-based distance features, several of which are themselves Jaccard distances over static tokens; the reported R2 is relative only to a constant-mean predictor. The classification labels come from the same silhouette-optimized clustering, so the experiments evaluate in-sample predictability of the clustering output rather than prediction of independently validated systemic flakiness. I ask the authors to add at least (i) a baseline using only the hierarchy distance or a single best feature, (ii) a permutation or random-feature baseline, and (iii) ideally a cross-project evaluation, and to temper the RQ2 conclusion accordingly.
- [Section 2.2] The clustering procedure is incompletely specified. The text does not state the linkage criterion used by SciPy's agglomerative clustering (for example, single, complete, average, or Ward) or the range and step size of the distance thresholds searched when maximizing mean silhouette. These choices materially change cluster membership and therefore the 606/810 count; they need to be reported either in this section or in the replication package for the analysis to be reproducible and for the threshold sensitivity of the headline result to be assessed.
minor comments (5)
- [Abstract] The sentence 'We call this phenomenonsystemic flakiness' is missing a space; it should read 'We call this phenomenon systemic flakiness.'
- [Section 2.2] The text 'it is does not require us to prespecify the number of clusters' contains a duplicated verb; it should read 'it does not require us to prespecify.'
- [Section 2.3] The text 'a value of 0 indicates they they differ in their first component' has a doubled 'they'; one occurrence should be removed.
- [Table 3 caption] The caption reads '21 static test case distances measures'; this should be '21 static test case distance measures.'
- [Section 2.4] The number of stack traces sampled per cluster is not reported, and the Levenshtein-diversity sampling procedure is described only briefly; adding these details would improve the reproducibility of RQ3.
Circularity Check
Partial circularity: RQ1's 75% prevalence follows from defining clusters by co-occurrence; shared root cause is assumed, not derived, though RQ3's manual inspection provides independent support.
-
self definitional
[Section 1 (Introduction), Section 2.2 (Methodology for RQ1), Section 2.5 (Threats to Validity, Construct Validity)]
""We discovered that flaky tests often exist in clusters, with co-occurring failures that share the same root causes, which we call systemic flakiness." ... "In the latter case [Jaccard distance 0], this would imply that the two flaky tests likely share the same root cause and are therefore a manifestation of systemic flakiness." ... "This study assumes that co-occurrence of flaky test failures is a symptom of systemic flakiness.""
The target construct, systemic flakiness, is defined to include shared root causes, but the RQ1 clustering operationalizes it solely through co-occurrence of failing run IDs (Jaccard distance). The reported result that 75% of flaky tests belong to a cluster is therefore a direct consequence of the cluster definition and the chosen silhouette-based threshold, not an independent measurement of shared root causes. The shared-root-cause component is explicitly assumed in Section 2.5 rather than derived from the clustering. RQ3's manual stack-trace inspection partially repairs this by providing qualitative evidence of shared causes, so the circularity is partial rather than total.
full rationale
The paper's central quantitative claim is operationalized by agglomerative clustering over Jaccard distances between sets of failing runs. Since the definition of systemic flakiness already includes shared root causes, the RQ1 prevalence figure does not independently test that notion; it restates the cluster definition together with a threshold chosen to maximize silhouette. The paper acknowledges this gap in Section 2.5, where it states that co-occurrence is assumed to be a symptom of systemic flakiness. This is a self-definitional element rather than a fitted-input prediction: no parameter is fitted to a subset and then used to 'predict' a closely related quantity, and the RQ2 machine learning models use static test-case distance features that are independent of the failure co-occurrence labels. The RQ3 manual inspection, conducted by four authors with negotiated agreement, provides independent qualitative evidence that many co-occurring failures share root causes such as networking and external dependencies, so the central claim is not wholly circular. There is no load-bearing self-citation chain or author-imported uniqueness theorem; the authors' prior work is cited mainly for definitions and existing flaky-test taxonomies, not as the foundation of the cluster analysis. The absence of a null model for chance co-occurrence is a conclusion-validity concern about how strongly the 75% figure supports the root-cause claim, but it does not add a separate circular step beyond the definitional operationalization already identified.
Assumptions & free parameters
free parameters (2)
- Silhouette cutoff for declaring clusters meaningful =
0.6
- Per-project Jaccard distance threshold =
0.00 to 0.52 depending on project (Table 1)
assumptions (3)
- domain assumption Co-occurrence of flaky test failures indicates shared root causes.
- domain assumption The FlakeFlagger dataset's 10,000 runs are a representative sample of flaky test behavior.
- domain assumption Static test case distance measures are informative about failure co-occurrence.
invented entities (1)
-
Systemic flakiness
Cite this review
Pith. "Pith review of Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures." pith.science (2026). https://pith.science/paper/UIQL5Z24
@misc{pith2026250416777,
author = {Pith},
title = {Pith review of: Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures},
year = {2026},
howpublished = {\url{https://pith.science/paper/UIQL5Z24}},
note = {Machine review of arXiv:2504.16777}
}
abstract
Flaky tests produce inconsistent outcomes without code changes, creating major challenges for software developers. An industrial case study reported that developers spend 1.28% of their time repairing flaky tests at a monthly cost of $2,250. We discovered that flaky tests often exist in clusters, with co-occurring failures that share the same root causes, which we call systemic flakiness. This suggests that developers can reduce repair costs by addressing shared root causes, enabling them to fix multiple flaky tests at once rather than tackling them individually. This study represents an inflection point by challenging the deep-seated assumption that flaky test failures are isolated occurrences. We used an established dataset of 10,000 test suite runs from 24 Java projects on GitHub, spanning domains from data orchestration to job scheduling. It contains 810 flaky tests, which we levered to perform a mixed-method empirical analysis of co-occurring flaky test failures. Systemic flakiness is significant and widespread. We performed agglomerative clustering of flaky tests based on their failure co-occurrence, finding that 75% of flaky tests across all projects belong to a cluster, with a mean cluster size of 13.5 flaky tests. Instead of requiring 10,000 test suite runs to identify systemic flakiness, we demonstrated a lightweight alternative by training machine learning models based on static test case distance measures. Through manual inspection of stack traces, conducted independently by four authors and resolved through negotiated agreement, we identified intermittent networking issues and instabilities in external dependencies as the predominant causes of systemic flakiness.
Figures
Reference graph
Works this paper leans on
-
[1]
Replication Package, https://doi.org/10.5281/zenodo.15267575
2025. Replication Package, https://doi.org/10.5281/zenodo.15267575
-
[2]
2025. scikit-learn: machine learning in Python — scikit-learn 1.5.2 documentation, https://scikit-learn.org/1.5/index.html
work page 2025
-
[3]
SciPy documentation — SciPy v1.14.1 Manual, https://docs.scipy.org/doc/s cipy-1.14.1/index.html
2025. SciPy documentation — SciPy v1.14.1 Manual, https://docs.scipy.org/doc/s cipy-1.14.1/index.html
work page 2025
-
[4]
M. R. Ackermann, J. Blömer, D. Kuntze, and C. Sohler. 2014. Analysis of Agglom- erative Clustering. Algorithmica 1, 2 (2014), 184–215
work page 2014
-
[5]
A. Afeltra, A. Cannavale, F. Pecorelli, V. Pontillo, and F. Palomba. 2024. A Large- Scale Empirical Investigation Into Cross-Project Flaky Test Prediction. IEEE Access 12 (2024), 131255–131265. Systemic Flakiness: An Empirical Analysis of Co-Occurring Flaky Test Failures
work page 2024
- [6]
-
[7]
A. Alshammari, P. Ammann, M. Hilton, and J. Bell. 2024. A Study of Flaky Failure De-Duplication to Identify Unreliably Killed Mutants. In Proceedings of the International Conference on Software Testing, Verification and Validation Workshops (ICSTW). 257–262
work page 2024
-
[8]
A. Alshammari, C. Morris, M. Hilton, and J. Bell. 2021. FlakeFlagger: Predicting Flakiness Without Rerunning Tests. InProceedings of the International Conference on Software Engineering (ICSE)
work page 2021
Show all 60 references
-
[9]
G. An, J. Yoon, J. Sohn, j. Hong, D. Hwang, and S. Yoo. 2022. Automatically Identifying Shared Root Causes of Test Breakages in SAP HANA. In Proceedings of the International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). 65–74
2022
-
[10]
J. Bell, O. Legunsen, M. Hilton, L. Eloussi, T. Yung, and D. Marinov. 2018. De- Flaker: Automatically Detecting Flaky Tests. In Proceedings of the International Conference on Software Engineering (ICSE) . 433–444
2018
-
[11]
L. Breiman. 2001. Random Forests. Machine Learning 45, 1 (2001), 5–32
2001
-
[12]
Camara, M
B. Camara, M. Silva, A. Endo, and Vergilio S. 2021. What is the Vocabulary of Flaky Tests? An Extended Replication. In Proceedings of the International Conference on Program Comprehension (ICPC) . 444–454
2021
-
[13]
Chicco, D
G. Chicco, D. Jurman. 2020. The Advantages of the Matthews Correlation Coeffi- cient (MCC) Over F1 Score and Accuracy in Binary Classification Evaluation. BMC Genomics 21, 6 (2020), 1471–2164
2020
-
[14]
Cordy, R
M. Cordy, R. Rwemalika, A. Franci, M. Papadakis, and M. Harman. 2022. FlakiMe: Laboratory-Controlled Test Flakiness Impact Assessment. In Proceedings of the International Conference on Software Engineering (ICSE) . 982–994
2022
-
[15]
Durieux, C
T. Durieux, C. L. Goues, M. Hilton, and R. Abreu. 2020. Empirical Study of Restarted and Flaky Builds on Travis CI. In Proceedings of the International Conference on Mining Software Repositories (MSR) . 254–264
2020
-
[16]
M. Eck, F. Palomba, M. Castelluccio, and A. Bacchelli. 2019. Understanding Flaky Tests: The Developer’s Perspective. In Proceedings of the Joint Meeting of the European Software Engineering Conference and the Symposium on the Foundations of Software Engineering (ESEC/FSE) . 830–840
2019
-
[17]
Elgendy, R
I. Elgendy, R. Hierons, and P. McMinn. 2024. Evaluating String Distance Met- rics for Reducing Automatically Generated Test Suites. In Proceedings of the International Conference on Automation of Software Test (AST) . 171–181
2024
-
[18]
Elgendy, R
I. Elgendy, R. Hierons, and P. McMinn. 2025. A Systematic Mapping Study of the Metrics, Uses and Subjects of Diversity-Based Testing Techniques. Software Testing, Verification and Reliability 35, 2 (2025), e1914
2025
-
[19]
Fatima, T
S. Fatima, T. Ghaleb, and L. Briand. 2022. Flakify: A Black-Box, Language Model- based Predictor for Flaky Tests. Transactions on Software Engineering (2022), 1–17
2022
-
[20]
J. H. Friedman. 2001. Greedy Function Approximation: A Gradient Boosting Machine. The Annals of Statistics 29, 5 (2001), 1189–1232
2001
-
[21]
Garousi and B
V. Garousi and B. Küçük. 2018. Smells in Software Test Code: A Survey of Knowledge in Industry and Academia. Journal of Systems and Software 138 (2018), 52–81
2018
-
[22]
Ernst, and L
P Geurts, D. Ernst, and L. Wehenkel. 2006. Extremely Randomized Trees.Machine Learning 63, 1 (2006), 3–42
2006
-
[23]
Golagha, C
M. Golagha, C. Lehnhoff, A. Pretschner, and H. Ilmberger. 2019. Failure Clustering Without Coverage. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA). 134–145
2019
-
[24]
Gruber and G
M. Gruber and G. Fraser. 2022. A Survey on How Test Flakiness Affects Develop- ers and What Support They Need to Address It. InProceedings of the International Conference on Software Testing, Verification and Validation (ICST)
2022
-
[25]
Gruber, M
M. Gruber, M. Heine, N. Oster, M. Philippsen, and G. Fraser. 2023. Practical Flaky Test Prediction using Common Code Evolution and Test History Data. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST)
2023
-
[26]
Gruber, S
M. Gruber, S. Lukasczyk, F. Kroiß, and G. Fraser. 2021. An Empirical Study of Flaky Tests in Python. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST)
2021
-
[27]
Habchi, G
S. Habchi, G. Haben, M. Papadakis, M. Cordy, and Y. Le Traon. 2022. A Qualitative Study on the Sources, Impacts, and Mitigation Strategies of Flaky Tests. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST)
2022
-
[28]
Harman and P
M. Harman and P. O’Hearn. 2018. From Start-ups to Scale-ups: Opportunities and Open Problems for Static and Dynamic Program Analysis. In Proceedings of the International Working Conference on Source Code Analysis and Manipulation (SCAM). 1–23
2018
-
[29]
Hashemi, A
N. Hashemi, A. Tahir, and S. Rasheed. 2022. An Empirical Study of Flaky Tests in JavaScript. In International Conference on Software Maintenance and Evolution (ICSME). 24–34
2022
-
[30]
Hilton, N
M. Hilton, N. Nelson, T. Tunnell, D. Marinov, and D. Dig. 2017. Trade-Offs in Continuous Integration: Assurance, Security, and Flexibility. In Proceedings of the Symposium on the Foundations of Software Engineering (FSE) . 197–207
2017
-
[31]
G. M. Kapfhammer. 2004. Software Testing. In The Computer Science Handbook
2004
-
[32]
Kaufman and P
L. Kaufman and P. J. Rousseeuw. 1990. Finding Groups in Data: An Introduction to Cluster Analysis. John Wiley & Sons
1990
-
[33]
W. Lam, P. Godefroid, S. Nath, A. Santhiar, and S. Thummalapenta. 2019. Root Causing Flaky Tests in a Large-Scale Industrial Setting. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA) . 204–215
2019
-
[34]
W. Lam, K. Muşlu, H. Sajnani, and S. Thummalapenta. 2020. A Study on the Lifecycle of Flaky Tests. InProceedings of the International Conference on Software Engineering (ICSE). 1471–1482
2020
-
[35]
W. Lam, R. Oei, A. Shi, D. Marinov, and T. Xie. 2019. IDFlakies: A Framework for Detecting and Partially Classifying Flaky Tests. InProceedings of the International Conference on Software Testing, Verification and Validation (ICST) . 312–322
2019
-
[36]
Leinen, D
F. Leinen, D. Elsner, A. Pretschner, A. Stahlbauer, M. Sailer, and E. Jurgens. 2024. Cost of Flaky Tests in Continuous Integration: An Industrial Case Study. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST). 329–340
2024
-
[37]
C. Li, C. Zhu, and A. Wang, W. Shi. 2022. Repairing Order-Dependent Flaky Tests via Test Generation. In Proceedings of the International Conference on Software Engineering (ICSE). 1881–1892
2022
-
[38]
Li and A
S. Li and A. Shi. 2022. Evolution-Aware Detection of Order-Dependent Flaky Tests. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA). 114–125
2022
-
[39]
S. M. Lundberg, G. Erion, H. Chen, A. DeGrave, J. M. Prutkin, B. Nair, R. Katz, J. Himmelfarb, N. Bansal, and S. Lee. 2020. From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence 2, 1 (2020), 2522–5839
2020
-
[40]
Q. Luo, F. Hariri, L. Eloussi, and D. Marinov. 2014. An Empirical Analysis of Flaky Tests. In Proceedings of the Symposium on the Foundations of Software Engineering (FSE). 643–653
2014
-
[41]
Memon, Z
A. Memon, Z. Gao, B. Nguyen, S. Dhanda, E. Nickell, R. Siemborski, and J. Micco
-
[42]
Parry, G
O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn. 2021. A Survey of Flaky Tests. Transactions on Software Engineering and Methodology 31, 1 (2021), 1–74
2021
-
[43]
Parry, G
O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn. 2022. Evaluating Features for Machine Learning Detection of Order- and Non-Order-Dependent Flaky Tests. In Proceedings of the International Conference on Software Testing, Verification and Validation (ICST). 93–104
2022
-
[44]
Parry, G
O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn. 2022. Surveying the Developer Experience of Flaky Tests. InProceedings of the International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . 253–262
2022
-
[45]
Parry, G
O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn. 2022. What Do Developer-Repaired Flaky Tests Tell Us About the Effectiveness of Automated Flaky Test Detection?. In Proceedings of the International Conference on Automa- tion of Software Test (AST). 160–164
2022
-
[46]
Parry, G
O. Parry, G. M. Kapfhammer, M. Hilton, and P. McMinn. 2023. Empirically Evaluating Flaky Test Detection Techniques Combining Test Case Rerunning and Machine Learning Models. Empirical Software Engineering 28, 72 (2023)
2023
-
[47]
Pinto, B
G. Pinto, B. Miranda, S. Dissanayake, M. D. Amorim, C. Treude, A. Bertolino, and M. D’amorim. 2020. What is the Vocabulary of Flaky Tests?. In Proceedings of the International Conference on Mining Software Repositories (MSR) . 492–502
2020
-
[48]
Pontillo, F
V. Pontillo, F. Palomba, and F. Ferrucci. 2022. Static Test Flakiness Prediction: How Far Can We Go? Empirical Software Engineering (2022), 325–327
2022
-
[49]
Y. Qin, S. Wang, K. Liu, B. Lin, H. Wu, L. Li, and X. Mao. 2022. PEELER: Learning to Effectively Predict Flakiness without Running Tests. In International Conference on Software Maintenance and Evolution (ICSME) . 257–268
2022
-
[50]
K. R. Shahapure and C. Nicholas. 2020. Cluster Quality Analysis Using Silhouette Score. In Proceedings of the International Conference on Data Science and Advanced Analytics (DSAA). 747–748
2020
-
[51]
A. Shi, W. Lam, R. Oei, T. Xie, and D. Marinov. 2019. iFixFlakies: A Framework for Automatically Fixing Order-dependent Flaky Tests. InProceedings of the Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FS...
2019
-
[52]
Spadini, M
D. Spadini, M. Aniche, M. Bruntink, and A. Bacchelli. 2017. To Mock or Not to Mock? An Empirical Study on Mocking Practices. In Proceedings of the Interna- tional Conference on Mining Software Repositories (MSR) . 402–412
2017
-
[53]
Uddin and H
S. Uddin and H. Lu. 2024. Confirming the Statistically Significant Superiority of Tree-Based Machine Learning Algorithms Over Their Counterparts for Tabular Data. Plos One 19, 4 (2024), e0301541
2024
-
[54]
Vahabzadeh, A
A. Vahabzadeh, A. A. Fard, and A. Mesbah. 2015. An Empirical Study of Bugs in Test Code. In Proceedings of the International Conference on Software Maintenance and Evolution (ICSME). 101–110
2015
-
[55]
Vancsics, T
B. Vancsics, T. Gergely, and A. Beszédes. 2020. Simulating the Effect of Test Flakiness on Fault Localization Effectiveness. In Proceedings of the International Workshop on Validation, Analysis and Evolution of Software Tests (VST) . 28–35
2020
-
[56]
Verdecchia, E
R. Verdecchia, E. Cruciani, B. Miranda, and A. Bertolino. 2021. Know Your Neighbor: Fast Static Prediction of Test Flakiness. IEEE Access 9 (2021), 76119– 76134. Issue 4. Owain Parry, Gregory M. Kapfhammer, Michael Hilton, and Phil McMinn
2021
-
[57]
J. Wang, Y. Lei, M. Li, G. Ren, H. Xie, S. Jin, J. Li, and J. Hu. 2024. Flakyrank: Predicting Flaky Tests Using Augmented Learning to Rank. In Proceedings of the International Conference on Software Analysis, Evolution and Reengineering (SANER). 872–883
2024
-
[58]
R. Wang, Y. Chen, and W. Lam. 2022. iPFlakies: A Framework for Detecting and Fixing Python Order-Dependent Flaky Tests. In Proceedings of the International Conference on Software Engineering Companion (ICSE Companion) . 120–124
2022
-
[59]
Zhang, D
S. Zhang, D. Jalali, J. Wuttke, K. Muşlu, W. Lam, M. D. Ernst, and D. Notkin. 2014. Empirically Revisiting the Test Independence Assumption. In Proceedings of the International Symposium on Software Testing and Analysis (ISSTA) . 385–396
2014
-
[2017]
InProceedings of the International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP)
Taming Google-Scale Continuous Testing. InProceedings of the International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . 233–242
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.