{"id":"8d596227-c87c-496d-b758-9d786f9d506e","arxiv_id":"2507.14547","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A multivocal review of 108 studies finds architectural degradation is a socio-technical process whose measurement is well supported but whose continuous remediation is largely missing.","lead":"After reviewing 108 studies, this paper maps how software architectures decay, why it happens, and which tools measure the damage. It is worth reading because it shows that degradation is as much an organizational problem as a technical one, and that current tools detect decay but rarely fix it.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Process-debt percentages in the text ('knowledge debt = 50%') contradict the paper's own Table 13 (≈32%), undermining the socio-technical reframing.","rationale":"Read in good faith, the paper is a useful MLR that organizes a fragmented literature and identifies a plausible gap between detection and continuous remediation. The qualitative synthesis, the Sankey diagram, and the tool/method tables provide independent value. I do not object to the overall conditional verdict. However, the most load-bearing quantitative assertion — that knowledge debt constitutes half of process debt and hence that 'architecture suffers most when architectural understanding is lost' — is contradicted by the paper's own Table 13. This is not a matter of external consensus or subjective preference; it is an internal inconsistency that can be checked directly against the published replication data. The reader flagged coding reliability and search completeness as threats; the specific 50% discrepancy is even more damaging because it does not depend on re-coding or broader searching, only on arithmetic consistency with the authors' own extraction. If the table is right, the 'predominance' claim is overstated; if the text is right, the taxonomy table is wrong. Either outcome requires a correction before the taxonomy is used as a baseline. The abstract's 'organizational' emphasis is not baseless — 22 process-debt SPs and qualitative evidence support it — but the specific quantitative support should be bounded. Therefore I recommend leaving the reader's verdict unchanged: CONDITIONAL, with the requirement that the authors reconcile the process-debt percentages and re-verify the knowledge-debt 50% claim.","tokens_in":39805,"tokens_out":7732,"duration_ms":84386,"concrete_test":"Using the replication package (https://doi.org/10.5281/zenodo.15848510), re-tabulate the motivations extracted from the 108 primary studies and reproduce the Process Debt classification: for each SP, list the motivation(s) and the debt category assigned (Architectural, Code, Process, combined). Sum the Process Debt SPs and the Knowledge-related sub-categories exactly as defined in Table 13. Then compute the percentage of Process Debt that is knowledge-related. Compare this with the text's 50% claim and with the 20.8%/29.2%/50% distribution in Section 4.2. If the reproduction yields approximately 32% rather than 50%, the claim is contradicted by the authors' own data; if it yields 50%, Table 13 must be corrected and the discrepancy explained. This single check settles whether the quantitative backbone of the socio-technical conclusion is sound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion that degradation is 'both technical and organizational' and that knowledge loss is a primary driver rests on the claim that knowledge-related process debt accounts for 50% of Process Debt (Section 5, bullet 'Socio-technical thinking'). This number appears in the abstract, results (Section 4.2), and discussion, yet it is not supported by the paper's own data tables. Table 13 reports Process Debt with 22 SPs (20.4%) and the Knowledge sub-category with 7 SPs (6.5%), i.e., 7/22 ≈ 31.8%, not 50%. A generous reading that adds the organizational-best-practices-unfollowed row gives 8/22 ≈ 36.4%. The text's other process-debt percentages (Development Practices 20.8%, Governance 29.2%, Knowledge 50%) sum to 100% but do not match Table 13's counts (5, 3, and 7 respectively; even combining Effort and Governance gives 13, not 7). Because the 'predominance of knowledge debt' is the key quantitative justification for the socio-technical reframing, this internal inconsistency is load-bearing. The paper provides a replication package, so the discrepancy is checkable; if the tables are correct, the 'half' claim is false and the discussion overstates organizational drivers; if the text is correct, the tables are wrong and the review's quantitative foundation is unreliable. Either way the conditional verdict stands, but the specific quantitative support for the central claim must be corrected or bounded.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a multivocal literature review (MLR) of 108 primary studies (1992–2024) on architectural degradation, defined broadly to include erosion, decay, and aging. The authors extract definitions, motivations (categorized as architectural, code, combined, and process debt), metrics, measurement approaches, tools, and remediation strategies. The main claims are that the concept has shifted from a purely technical to a socio-technical phenomenon, that process-related causes (especially knowledge loss) are underappreciated, that existing metrics and tools are mostly structural and disconnected from remediation, and that continuous, preventive remediation is lacking. The paper proposes a unified definition and a research agenda. A replication package is provided at a Zenodo DOI.","tokens_in":40117,"tokens_out":8314,"duration_ms":82786,"significance":"If the quantitative inconsistencies are corrected, this would be a valuable synthesis for the software architecture community. The multivocal scope (white and gray literature), the explicit taxonomy of debt types, and the spotlight on the disconnect between detection and remediation are useful contributions. The paper also ships a replication package, which is a strength. The qualitative narrative—that degradation is socio-technical and that remediation is underdeveloped—is plausible and consistent with the broader literature. However, the quantitative support for some central claims is currently unreliable, and the significance of the contribution depends on fixing these issues.","major_comments":[{"comment":"The claim that knowledge debt accounts for 50% of Process Debt is not supported by the paper's own data. Table 13 reports Process Debt as 22 primary studies (20.4%) and the Knowledge sub-category as 7 primary studies (6.5%), i.e., 7/22 ≈ 31.8%. Figure 7 lists Knowledge as (10, 9.3%), which would give at most 8/22 ≈ 36.4% if the organizational-best-practices-unfollowed row is included. The text in the Abstract, Section 4.2 (key insight #3), and Section 5 (first discussion bullet) repeats the 50% figure, and the text's other process-debt percentages (Development Practices 20.8%, Governance 29.2%, Knowledge 50%) sum to 100% but do not match the counts in Table 13 (5, 3/10, and 7 respectively). Because the predominance of knowledge debt is a load-bearing quantitative justification for the socio-technical reframing, this internal inconsistency must be resolved: either the tables are corrected to match the text, or the text must be revised to state the percentages actually derivable from the data.","section":"§4.2, §5, Table 13, Figure 7, Abstract"},{"comment":"The metric counts in RQ3 are internally inconsistent. The text states that 54 metrics were identified, and Table 14 lists 24 architectural-debt metrics and 30 code-debt metrics, which sum to 54. However, the table's own header columns report 'Architectural debt (13, 12%)' and 'Code debt (15, 13.9%)', which sum to only 28 metrics and 25.9%. Figure 8 repeats these conflicting labels (e.g., Architectural Debt shown as 13, 12% with a separate '24 Metrics' tag). A reader cannot determine the actual distribution of metrics across the two categories from the paper as written. Since the abstract and conclusions explicitly cite '54 metrics', the authors should correct the table and figure headers so that the category totals match the listed items.","section":"§4.3, Table 14, Figure 8"},{"comment":"Several percentage labels in the motivation figures and tables do not match each other or the text. Specific examples: Figure 6 shows Code Debt as (48, 44.4%) while its child 'Implementation & Code Quality' is (52, 48.1%), so the child count exceeds its parent; Table 11 reports this child as (50, 46.3%), a third value. Figure 7 reports Knowledge as (10, 9.3%) while Table 13 reports Knowledge as (7, 6.5%). The text in §4.2 says the Technological Evolution category is 1.9% for both Architectural Debt and Code Debt, but Table 10 and Table 11 each report it as (1, 0.9%). These discrepancies make the quantitative synthesis of motivations unreliable as presented. The authors should audit all percentages in the results sections against the appendix tables and ensure that each percentage is computed on a stated, consistent denominator.","section":"§4.2, Figures 5–7, Tables 10–13"}],"minor_comments":[{"comment":"The search string restricts to 'software architec*' in title/abstract together with degradation/aging/erosion/decay terms. Although snowballing is used, relevant studies using other terminology (e.g., 'design erosion' without 'software architecture') may be missed. Please make this limitation explicit in the threats-to-validity section, beyond the general conclusion-validity sentence.","section":"§3.2.1, §6"},{"comment":"The axial coding of motivations was performed by the three most senior authors, and the inter-rater agreement for this step is not reported (the paper reports agreement only for study selection). Please report a kappa statistic or equivalent for the classification step, or at least describe the reconciliation process in more detail.","section":"§3.4"},{"comment":"There are several typographical errors and inconsistent labels throughout: 'colored threes' should be 'colored trees' (§4 opening), 'hone hand' should be 'one hand' (§4.5), 'sepraterly' should be 'separately' (§4.2), 'gwoing' should be 'growing' (§5), 'Idnetifying' is misspelled in Figure 11, and 'Statistic analysis' in Figure 11 should be 'Static analysis'.","section":"§4 and figures"},{"comment":"The abstract and results mention '31 measurement techniques', but the total is not clearly derived from the categories in Figure 9 or Table 15. Please state explicitly how the count of 31 is obtained, and ensure that the counts in Figure 9 sum to that number.","section":"§4.4, Figure 9"},{"comment":"The conclusion-validity paragraph says the snowballing process 'resulted in one additional relevant paper', while Table 4 and the text in §3.2.4 report +19 papers from snowballing. This appears to be a typo and should be corrected to '19 additional papers' or rephrased if a different step is meant.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":"The paper's replication package is a strong point, and the qualitative synthesis is likely salvageable. However, the internal numerical inconsistencies are not merely cosmetic: the 50% knowledge-debt claim is used as the quantitative anchor for the socio-technical reframing, and the metric-category counts are contradicted by the same table that lists them. Before sending this for revision, I would ask the authors to re-derive every percentage from their raw data and to state the denominator used in each case. If the tables are corrected, the paper could become a solid reference for the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take: this is a useful, well-scoped multivocal review that updates the prior surveys and gives the field a common vocabulary. The best parts are the corpus (108 studies incl. gray lit), the taxonomy of degradation motivations built on Li et al.'s technical debt classification, and the Sankey diagram showing how measurement flows rarely reach remediation. The claim that detection is well-studied but continuous remediation is lacking is credible and supported by the synthesis.\n\nThe soft spot is real: the process-debt percentages in the text contradict the paper's own tables. The abstract and discussion say knowledge debt is 50% of process debt; Table 13 counts 7 of 22 (about 32%) or 8 with the best-practices row (36%). The text's 20.8/29.2/50 split also doesn't match the table's 5/3/7 counts. This inconsistency is load-bearing because the socio-technical reframing leans on the predominance of knowledge debt. The fix is straightforward: recalculate and present the actual numbers, or bound the claim to what the qualitative coding shows. I wouldn't reject the paper over this, but it needs correction.\n\nOther concerns are minor. The search string is narrow (title/abstract with 'software architec*' plus degradation/aging/erosion/decay), though snowballing helps. The axial coding is done by the three senior authors, but they report inter-rater agreement and provide a replication package, so that's a reasonable — not bulletproof — basis.\n\nWho is this for? Researchers in software architecture and technical debt will find it a useful reference and a baseline for future work. Practitioners might browse the tool and metric lists, but it's not a hands-on guide. It deserves peer review; the synthesis is substantial and the gap statement is actionable. I'd send it with a request to fix the quantitative inconsistencies and bound the search limitations.","headline":"Largely useful multivocal synthesis with a credible detection-vs-remediation gap; the 50% knowledge-debt claim contradicts the paper's own Table 13 and must be fixed before the taxonomy is used as a baseline.","tokens_in":40654,"tokens_out":5484,"would_cite":true,"duration_ms":51784,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This multivocal review of 108 studies claims architectural degradation is both technical and organizational, and that while detection is well-studied, continuous remediation is missing.","keywords":["software architecture","architectural degradation","architectural erosion","architectural decay","technical debt","multivocal literature review","software metrics","remediation approaches"],"falsifier":"Run an independent replication of the coding: have a fresh team classify the same 108 primary studies' motivations using the paper's debt categories, then compare the resulting percentages (for example, knowledge debt as 50% of process debt); if the shares move substantially or inter-rater agreement is poor, the taxonomy's numerical claims are unstable.","tokens_in":39625,"feed_emoji":"🏗️","tokens_out":5907,"duration_ms":60340,"temperature":0.7,"pith_summary":"This paper is a systematic, multivocal review of 108 studies (1992–2024) aiming to unify how software architectural degradation—erosion, decay, or aging—is defined, measured, and remediated. It claims that degradation is not only a code-level or design-level problem but a socio-technical one: poor documentation, hasty fixes, time pressure, and knowledge loss all drive the same underlying drift between intended and actual architecture. The review finds that researchers and practitioners have produced 54 metrics and 31 measurement techniques, mostly for smells, coupling and cohesion, and evolution, but that these are rarely connected to tools that repair or prevent degradation. The central conclusion is that detection is well-studied while continuous, preventive remediation is missing, and future effort should build integrated pathways from metrics to tools to repair logic.","feed_headline":"Software decays from lost knowledge, not just bad code","feed_subtitle":"A 108-study review finds detection is strong but continuous repair is absent.","key_machinery":"The central machinery is the debt-layer taxonomy (architectural, code, and process debt) derived from axial coding of the primary studies, used to classify definitions, motivations, metrics, measurement approaches, tools, and remediation strategies. The other load-bearing instrument is the Sankey-style flow diagram that traces studies from measurement approaches through metrics to tools and remediation, making visible the “No Metrics”, “No Tool”, and “No Remediation” bottlenecks. These two devices together carry the argument that degradation is socio-technical and that remediation is disconnected from detection.","core_discovery":"On its own terms, the paper establishes a taxonomy of degradation causes split into architectural debt, code debt, combined debt, and process debt, and proposes a unified definition: architectural degradation is the progressive divergence between implemented and intended architecture, caused by repeated violations of rules and design principles and by cumulative code-level changes, leading to loss of key properties and rising complexity. The discovery is the gap structure: the field is mature at recognizing degradation but weak at acting on it. A flow analysis that traces studies from measurement approaches through metrics to tools and remediation shows that most approaches either stop before defining concrete metrics or lead to “No Tool” and “No Remediation” dead ends. The paper interprets this as evidence that metrics, tools, and repair logic are not integrated, and that degradation is both technical and organizational.","pith_inferences":["Beyond the paper: the knowledge-debt finding suggests a testable hypothesis that repositories with higher developer turnover should show faster architectural decay even when code quality metrics are controlled.","Beyond the paper: the taxonomy provides categories for building a degradation risk score that combines process indicators (time pressure, turnover, documentation freshness) with structural metrics (coupling, smells), though the paper itself does not propose such a formula.","Beyond the paper: because the paper finds no tools that measure knowledge loss, a natural next step is to mine developer communication and onboarding documents for process-debt signals.","Beyond the paper: if the “No Metrics” bottleneck is real, reflection-model and smell-detection studies could be re-analyzed to attach existing metrics to their qualitative findings, turning descriptive studies into quantitative baselines."],"forward_implications":["If degradation is socio-technical, then process debt—especially knowledge loss, turnover, and time pressure—must be measured and managed alongside structural metrics.","The 54 metrics and 31 measurement techniques catalogued here give researchers and practitioners a baseline for choosing indicators rather than inventing new ones.","Because most tools detect smells or violations but do not connect to repair, CI/CD pipelines will need tools that trigger remediation recommendations from metric thresholds.","A unified definition of architectural degradation could let empirical studies compare results across terminologies like erosion, decay, and aging."],"supporting_citations":[{"why":"Supplies the multivocal literature review guidelines that justify including gray literature and the quality criteria used for gray sources.","marker":"Garousi et al., 2019"},{"why":"Provides the systematic literature review protocol and the two-reviewer data extraction method.","marker":"Kitchenham and Charters, 2007"},{"why":"Defines the snowballing procedure used to recover 19 additional primary studies.","marker":"Wohlin, 2014"},{"why":"Provides the technical debt classification that the paper extends into architectural, code, and process debt.","marker":"Li et al., 2015"},{"why":"Grounds the axial coding approach the three senior authors used to build the taxonomy.","marker":"Corbin and Strauss, 2014"},{"why":"The prior erosion-control survey whose minimize, prevent, and repair categories frame the paper's remediation analysis.","marker":"de Silva and Balasubramaniam, 2012"},{"why":"The prior mapping study of architectural erosion metrics that the paper's metric catalog extends and updates.","marker":"Baabad et al., 2022"},{"why":"Supplies practitioner perspectives on erosion as a structural issue with technical and non-technical causes, supporting the socio-technical conclusion.","marker":"Li et al., 2021"}],"fun_headline_variants":["Software decay: we detect but rarely repair, 108-study review","Architectural erosion: detection strong, ongoing repair missing","Software aging: metrics and tools exist, but no continuous fix","Decay detectors plentiful, repair logic scarce: 108 studies","Software decay is organizational, not just technical"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative claims depend on the three senior authors' coding of primary studies into architectural, code, and process debt, so a different coding team could shift the percentages that drive the gap analysis.","fun_headline_variants_meta":{"raw":{"variants":["Software decay: we detect but rarely repair, 108-study review","Architectural erosion: detection strong, ongoing repair missing","Software aging: metrics and tools exist, but no continuous fix","Decay detectors plentiful, repair logic scarce: 108 studies","Software decay is organizational, not just technical"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1458,"prompt_tokens":926,"completion_tokens":532,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":450}},"tokens_in":542,"tokens_out":532,"duration_ms":6480,"temperature":1.0,"reasoning_tokens":450,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:53:30.639781+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an independent replication of the coding: have a fresh team classify the same 108 primary studies' motivations using the paper's debt categories, then compare the resulting percentages (for example, knowledge debt as 50% of process debt); if the shares move substantially or inter-rater agreement is poor, the taxonomy's numerical claims are unstable.","supporting_citations":[{"cited_title":", author Felderer, M","cited_arxiv_id":null,"evidence_quote":"Supplies the multivocal literature review guidelines that justify including gray literature and the quality criteria used for gray sources."},{"cited_title":", author Charters, S","cited_arxiv_id":null,"evidence_quote":"Provides the systematic literature review protocol and the two-reviewer data extraction method."},{"cited_title":", year 2014","cited_arxiv_id":null,"evidence_quote":"Defines the snowballing procedure used to recover 19 additional primary studies."},{"cited_title":", author Strauss, A","cited_arxiv_id":null,"evidence_quote":"Grounds the axial coding approach the three senior authors used to build the taxonomy."},{"cited_title":", author Balasubramaniam, D","cited_arxiv_id":null,"evidence_quote":"The prior erosion-control survey whose minimize, prevent, and repair categories frame the paper's remediation analysis."},{"cited_title":", author Zulzalil, H.B","cited_arxiv_id":null,"evidence_quote":"The prior mapping study of architectural erosion metrics that the paper's metric catalog extends and updates."}],"review_version":1}