{"id":"a1296725-9996-4f5a-a38a-bc13398fd6c3","arxiv_id":"1908.05399","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 142 empirical studies shows that object-oriented code refactoring does not consistently improve software quality, with outcomes varying by quality attribute, refactoring type, and study setting.","lead":"This paper reviews 142 past experiments on whether cleaning up object-oriented code, known as refactoring, improves software quality. It finds that refactoring helps in some cases but not others, and that results differ between academic and industrial settings.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Vote-counting synthesis is not robust: multiplicative vote weights and selective exclusion of non-significant votes can flip which quality attributes are classified as improved, degraded, unchanged, or inconsistent; the authors' own §5.5 example shows this.","rationale":"This is a carefully conducted and transparent systematic mapping study, and the qualitative conclusion that refactoring does not always improve software quality is well supported. The search protocol, study selection, quality assessment, and reporting are strengths, and the authors openly discuss many limitations. My concern is narrower: the precise attribute-level findings in the abstract and in RQ5 are not robust to the vote-counting design. The multiplicative vote weighting and the exclusion of non-significant votes are both inherited from Dallal and Abdin, but inheritance does not establish validity. The paper's own Section 5.5 provides an example where including non-significant votes changes the classification, so the sensitivity is acknowledged rather than hypothetical. Because the 'except' list is defined by the 50% threshold, small changes in vote counting are load-bearing. The academic-versus-industrial comparison is additionally fragile because it compares 125 academic PSs with 17 industrial PSs and is confounded with dataset type, as the authors themselves note. The proposed recomputation is feasible with the online dataset and would settle whether the findings are artifacts. I would condition acceptance on reporting this sensitivity analysis or on rewording the attribute-level conclusions appropriately; this does not require rejection because the mapping study's descriptive contributions and open-challenge enumeration remain valuable.","tokens_in":52452,"tokens_out":6416,"duration_ms":62707,"concrete_test":"Recompute the final impact classifications in Figures 20–22 and the academic-versus-industrial percentages in Tables 11–12 using one vote per primary study per quality attribute, with each study's direction determined by its own reported conclusion or by majority of its measure/activity/dataset outcomes, while keeping the 50% threshold. Also run a second variant that includes non-significant +/−/= votes alongside significant votes, as the paper's own §5.5 example suggests. If any of the five 'except' attributes changes category, or if any academic/industrial comparison reverses, then the headline conclusion is an artifact of the multiplicative weighting and the paper should either report this sensitivity analysis or soften the attribute-level claims.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that refactoring improves or degrades all quality attributes except cohesion, complexity, inheritance, fault-proneness, and power consumption—depends on the vote-counting procedure described in Section 4.2.5. That procedure has two fragile choices. First, each primary study contributes votes equal to (#quality measures) × (#refactoring activities) × (#datasets); for example, [S25] contributes 12 × 9 × 6 = 648 votes for cohesion. These votes are not independent evidence, and this weighting can let one study dominate dozens of others without empirical justification. Second, whenever any statistically significant result exists for an attribute, all non-significant votes are discarded, and only 26 of 142 PSs report statistical significance. For attributes such as cohesion, fault-proneness, and power consumption, the final classification can rest on a handful of significant votes from one or two studies. The authors acknowledge in Section 5.5 that including non-significant votes can change the final impact classification, and they give a concrete example where the classification flips. Because the 'except' list is exactly the set of attributes near the 50% threshold, this is not a minor caveat: the headline attribute-level conclusions and the academic-versus-industrial contrast are direct outputs of these counting choices.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic mapping study (SMS) of 142 primary studies, published up to December 2017, on the effect of object-oriented code refactoring activities on software quality. The authors follow Kitchenham and Charters and Petersen et al., document automatic and manual search procedures, report Cohen's kappa agreement across screening stages, apply a ten-question quality assessment, and classify studies along seven facets. They answer five research questions, using vote-counting for RQ5 to synthesize outcomes on internal and external quality attributes, academic versus industrial settings, and overall versus individual refactoring activities. The headline conclusions are that refactoring improves or degrades most quality attributes except cohesion, complexity, inheritance, fault-proneness, and power consumption, and that academic studies report more positive impact than industrial studies.","tokens_in":52778,"tokens_out":7884,"duration_ms":76432,"significance":"If the RQ5 synthesis were robust, this would be the most comprehensive secondary study to date on the impact of refactoring on object-oriented quality, nearly doubling the primary-study pool of the closest prior SLR and adding an academic/industrial comparison plus tool and dataset classifications. The procedural strengths are real: search strings and screening counts are reported, inter-rater agreement is quantified, quality criteria are explicit, and the online appendix provides extraction data. The central limitation is that the headline attribute-level conclusions are driven by a vote-counting procedure whose weighting and significance-exclusion rules are fragile; because the authors themselves document in Section 5.5 an example where the classification flips, the current framing overstates confidence in both the 'except' list and the academic-versus-industrial contrast. The manuscript is a useful contribution, but the synthesis claims need substantial reframing or sensitivity analysis before publication.","major_comments":[{"comment":"Section 4.2.5 defines each primary study's vote contribution for a quality attribute as the product of its number of quality measures, refactoring activities, and datasets, so a single study such as [S25] is reported as contributing 648 votes for cohesion. This multiplicative weighting has no empirical justification and makes vote counts a poor summary of evidence: one study can dominate dozens of others in Tables 11 and 12. Because the final classification depends on a 50% threshold of these weighted votes, the paper should report at least one alternative synthesis at the level of primary studies (for example, one vote per study per attribute) and a sensitivity analysis across plausible weighting schemes before presenting attribute-level conclusions.","section":"4.2.5"},{"comment":"The vote-counting procedure discards all non-significant votes for any quality attribute that has at least one statistically significant vote, and Table 6 (QA7) shows that only 18.3% of the 142 primary studies report statistical significance. Consequently, the classifications for cohesion, complexity, size, inheritance, and fault-proneness can rest on a small number of significant votes drawn from very few studies, and the 'except' list is exactly the set of attributes closest to the 50% threshold. The authors acknowledge this fragility in Section 5.5, where the Extract Class/cohesion example flips from negative under significant-only voting to positive when non-significant votes are included; this is a load-bearing caveat, not a minor one. The paper should present both significant-only and all-vote results consistently, and should phrase conclusions as provisional or sensitive to synthesis assumptions.","section":"4.2.5 / Section 5.5"},{"comment":"The academic-versus-industrial comparison rests on 125 academic versus 17 industrial primary studies, and Section 5.5 itself lists dataset type and quality-measure selection as factors confounded with study setting. The abstract's statement that academic studies 'found more positive impact' of refactoring therefore goes beyond what this mapping design can establish. At minimum, the claim should be reported as an observed association with the confounds named in the same sentence; ideally, the authors should re-analyze a matched subset of studies that controls for dataset type and quality measures, or explicitly state why such a control is impossible.","section":"4.2.5.1 / 4.2.5.2 / 5.5"},{"comment":"Figure 22 and the surrounding text classify 71 individual refactoring-activity/quality-attribute pairs, but Section 4.2.5.3 reports that 590 of the 661 total pairs are investigated by only one or two primary studies, and many displayed pairs are based on a single study. Presenting these as the 'impact' of a refactoring activity without showing the number of supporting studies or confidence information is misleading; Figure 22 should indicate the evidence count per pair, and the text should avoid categorical statements such as 'most refactoring activities have either negative or neutral impact' when the underlying evidence is so sparse.","section":"4.2.5.3 / Figure 22"}],"minor_comments":[{"comment":"The numerical example about Extract Class attributes contains a typo: the counts are introduced for cohesion, but the following sentence refers to inheritance; this should be corrected.","section":"5.5"},{"comment":"The abstract says refactoring caused all quality attributes to improve or degrade except the listed ones, but Figures 20 and 21 include an explicit 'inconsistent' category; the wording should acknowledge that some attributes are classified as inconsistent rather than simply improved or degraded.","section":"Abstract / Section 4.2.5"},{"comment":"The caption of Figure 5f labels the facet as 'Focus (RQ6)', but the study defines only RQ1 through RQ5; this should be corrected to RQ5 or renamed consistently.","section":"Figure 5f"},{"comment":"In the Fault-Proneness row, the negative-vote study list contains a formatting error — '[S14], [[S38], S68]' has an extra opening bracket before [S38].","section":"Table 12"},{"comment":"Figure 16 is a very dense matrix of unlabeled numbers and is difficult to interpret; a clearer layout with a legend, or a split into smaller figures, would substantially improve readability.","section":"Figure 16"}],"recommendation":"major_revision","confidential_remarks":"The vote-counting fragility is the main reason for the major-revision recommendation. I would ask the authors to re-run or at least clearly present study-level sensitivity analyses; if the online appendix is accessible, the vote tables (Tables 11 and 12) should be verified against it during revision. The scope fit is appropriate for a software engineering journal, and I see no novelty or citation-pattern concerns beyond the synthesis-methodology issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Kaur and Singh have produced a workmanlike systematic mapping study that deserves attention from anyone working on refactoring. The genuinely new parts are the expanded corpus—142 primary studies versus 76 in Al Dallal and Abdin—the explicit academic/industrial split, and the inclusion of process, people, performance, and tool-related work. The protocol is well documented: Kitchenham and Petersen guidelines, kappa agreement, quality assessment, online data, snowballing. That is real, reproducible effort, and it shows.\n\nThe headline claim—that refactoring does not consistently improve software quality—is supported by the assembled evidence and is properly hedged. But the stress-test concern is on target: the vote-counting synthesis in Section 4.2.5 is fragile. Voting weights are #measures × #refactorings × #datasets, so a single study like [S25] can contribute hundreds of votes for one attribute. And whenever any statistically significant vote exists for an attribute, all non-significant votes are discarded. Since only 26 of 142 primary studies report significance, the attribute-level classifications rest on a thin base. The authors acknowledge this in Section 5.5 and even give an example where the classification flips, so the issue is not concealed. The consequence is that the specific “improved/degraded/unchanged” labels for cohesion, complexity, inheritance, fault-proneness, and power consumption are not robust; change the counting rule and they could move. The high-level conclusion—that evidence is mixed and often methodologically weak—would survive, though.\n\nThe academic-versus-industrial contrast is also softer than it looks: only 17 industrial studies, and those differ from academic ones in dataset type and measures. I would treat that as an observation about the literature rather than a measured effect.\n\nNone of this is fatal. The paper’s value is organizational: a comprehensive, current map with explicit research gaps. Readers should use the map and the challenges, not the exact vote counts. It deserves a serious referee; I would send it out with a request for a sensitivity analysis (e.g., one-study-one-vote) or at least a stronger caveat that the attribute-level results are counting-dependent. For my own work over the next year, I likely won’t cite it, but it is a legitimate reference for refactoring researchers and a reasonable entry point for practitioners.","headline":"A solid, transparent systematic map of 142 refactoring studies; the map and the mixed-evidence conclusion hold up, but the attribute-level vote-counting labels are counting-dependent and shouldn't be read as robust.","tokens_in":53197,"tokens_out":2880,"would_cite":false,"duration_ms":27432,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This systematic mapping of 142 empirical studies argues that code refactoring improves or degrades every software quality attribute examined except cohesion, complexity, inheritance, fault-proneness, and power consumption, and that…","keywords":["software quality","object-oriented software","refactoring activity","quality measures","systematic mapping study","code smells","vote-counting","refactoring impact"],"falsifier":"Recompute the vote-counting results for cohesion, complexity, inheritance, fault-proneness, and power consumption using all votes—significant and non-significant—rather than significant-only votes; the paper itself notes that including non-significant votes changes the classification of Extract Class's effect on inheritance from negative to positive. If for any of the five attributes the 50% threshold flips, the claim that these are the only attributes without a directional effect fails. A second check would re-run the tally with each study contributing one vote per attribute instead of a measure-activity-dataset product and see whether the academic-versus-industrial gap survives.","tokens_in":52213,"feed_emoji":"🔧","tokens_out":5814,"duration_ms":58087,"temperature":0.7,"pith_summary":"The paper sets out to settle a long-running dispute: does refactoring object-oriented code actually improve software quality? It maps 142 empirical studies published through December 2017 and combines their results by vote-counting. The central finding is that refactoring does not always help: across all quality attributes combined, every attribute either improved or degraded except cohesion, complexity, inheritance, fault-proneness, and power consumption, which stayed unsettled. The paper also finds a persistent split between settings: studies run in academia report mostly positive effects, while industrial studies mostly report no detectable effect. This matters because practitioners are reluctant to refactor without evidence, and the paper shows the evidence is weaker and more context-dependent than the textbook promise.","feed_headline":"Refactoring gains are real in academia, weak in industry","feed_subtitle":"Across 142 empirical papers, every quality attribute moves except five—and industrial settings see little benefit.","key_machinery":"The vote-counting synthesis is the machinery. Each primary study contributes votes for a quality attribute equal to the product of the number of quality measures, refactoring activities, and datasets it used; significant votes outweigh non-significant ones, and a 50% threshold classifies the attribute's net effect as positive, negative, neutral, or inconsistent. It is the device that turns 142 heterogeneous studies into the paper's headline conclusions.","core_discovery":"The paper claims that when the 142 primary studies are tallied by vote-counting, overall refactoring is beneficial for coupling, size, encapsulation, messaging, polymorphism, composition, and combinations of attributes, and for most external attributes such as understandability, maintainability, and reusability. It is harmful for adaptability, time behaviour, developers' coordination, productivity, and analysability. Cohesion and inheritance remain unchanged, fault-proneness and power consumption show insignificant change, and complexity is inconsistent, so these five attributes are the exceptions to a directional effect. The paper further claims that academic settings yield positive effects more often than industrial settings, and that individual refactoring activities have variable effects: the same activity can improve one quality attribute while degrading another, and can even improve the same attribute in one study and degrade it in another. Refactoring is therefore not a guaranteed quality improvement, and its measured benefits depend on setting, measures, and datasets.","pith_inferences":["If the vote-counting weighting is sensitive to the measure-activity-dataset product, the five exception attributes may be an artifact of that weighting rather than a stable finding; a sensitivity analysis using unweighted study counts could flip the classification.","The academic-versus-industrial gap may partly reflect dataset type rather than setting itself, since academic studies mostly used open-source systems while industrial studies used commercial ones; the paper does not control for this confound.","Contradictions on the same dataset, such as opposite defect findings for ArgoUML, suggest that re-running key studies on JHotDraw, GanttProject, and Apache Ant with a shared metric framework could resolve disputes cheaply.","A standard unified metric framework, which the paper recommends, would allow direct comparison across studies and might turn the vote-counting synthesis into a quantitative meta-analysis."],"forward_implications":["Practitioners should not assume refactoring is uniformly beneficial; the same activity can improve one quality attribute while degrading another, so refactorings should be chosen with a specific target attribute in mind.","Academic results may overstate the benefits of refactoring, so industrial validation with commercial datasets and experienced developers is needed before relying on refactoring for quality improvement.","The five unsettled attributes—cohesion, complexity, inheritance, fault-proneness, and power consumption—are the places where the evidence is weakest and where new studies would have the most leverage.","The near-total absence of statistical significance testing in the primary studies limits what any synthesis can conclude; more studies reporting significance levels would make the field's evidence base firmer.","Only a handful of tools predict or assess refactoring impact, and most are outdated or restricted to Java, leaving a clear gap for tool builders."],"supporting_citations":[{"why":"Supplies the vote-counting approach, the six outcome categories, and the 50% threshold that the paper adapts for its synthesis.","marker":"[21]"},{"why":"Defines the Fowler refactoring catalogue used to decide which refactoring activities target code smells and qualify for inclusion.","marker":"[8]"},{"why":"Provides the systematic literature review guidelines that shape the search, screening, and quality assessment protocol.","marker":"[31]"},{"why":"Supplies the internal versus external quality measure classification used to organise the extracted data.","marker":"[65]"},{"why":"Defines the MOOSE/CK metric suite that dominates the quality measures used in the primary studies.","marker":"[69-70]"},{"why":"Describes Ref-Finder, the tool many primary studies use to extract refactoring activities from software repositories.","marker":"[67]"},{"why":"The ISO/IEC 25010 standard used to map external quality attribute subtypes to the most appropriate external quality attributes.","marker":"[77]"},{"why":"Cited as the basis for the paper's own caution that vote-counting may be erroneous.","marker":"[89]"}],"fun_headline_variants":["Refactoring quality gains: strong in academia, weak in industry","Refactoring benefits are real but only in academic settings","Five quality attributes resist refactoring's influence","Refactoring doesn't always improve quality—context matters","Academic refactoring looks great; industry sees little benefit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The synthesis stands or falls on the vote-counting rules: that each study's contribution should be weighted by the product of its quality measures, refactoring activities, and datasets, that non-significant findings can be discarded whenever any significant result exists, and that the 50% threshold cleanly separates positive, negative, neutral, and inconsistent effects.","fun_headline_variants_meta":{"raw":{"variants":["Refactoring quality gains: strong in academia, weak in industry","Refactoring benefits are real but only in academic settings","Five quality attributes resist refactoring's influence","Refactoring doesn't always improve quality—context matters","Academic refactoring looks great; industry sees little benefit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000426,"raw_usage":{"total_tokens":2195,"prompt_tokens":973,"completion_tokens":1222,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":589,"completion_tokens_details":{"reasoning_tokens":1141}},"tokens_in":589,"tokens_out":1222,"duration_ms":10077,"temperature":1.0,"reasoning_tokens":1141,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:14:20.078337+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the vote-counting results for cohesion, complexity, inheritance, fault-proneness, and power consumption using all votes—significant and non-significant—rather than significant-only votes; the paper itself notes that including non-significant votes changes the classification of Extract Class's effect on inheritance from negative to positive. If for any of the five attributes the 50% threshold flips, the claim that these are the only attributes without a directional effect fails. A second check would re-run the tally with each study contributing one vote per attribute instead of a measure-activity-dataset product and see whether the academic-versus-industrial gap survives.","supporting_citations":[{"cited_title":"Empirical evaluation of the impact of object-oriented code refactoring on quality attributes: A systematic literature review,","cited_arxiv_id":null,"evidence_quote":"Supplies the vote-counting approach, the six outcome categories, and the 50% threshold that the paper adapts for its synthesis."},{"cited_title":"Guidelines for performing systematic literature reviews in software engineering ,","cited_arxiv_id":null,"evidence_quote":"Provides the systematic literature review guidelines that shape the search, screening, and quality assessment protocol."},{"cited_title":"Software metrics a rigorous & practical approach,","cited_arxiv_id":null,"evidence_quote":"Supplies the internal versus external quality measure classification used to organise the extracted data."},{"cited_title":"Template- based reconstruction of complex refactorings,","cited_arxiv_id":null,"evidence_quote":"Describes Ref-Finder, the tool many primary studies use to extract refactoring activities from software repositories."},{"cited_title":"Systems and software engineering – Systems and software Quality Requirements and Evaluation (SQuaRE) – System and software quality models,","cited_arxiv_id":null,"evidence_quote":"The ISO/IEC 25010 standard used to map external quality attribute subtypes to the most appropriate external quality attributes."},{"cited_title":"Aggregation of empirical evidence,","cited_arxiv_id":null,"evidence_quote":"Cited as the basis for the paper's own caution that vote-counting may be erroneous."}],"review_version":1}