{"id":"add5fc2c-c67f-4a7d-b2ce-2738eda6af97","arxiv_id":"2501.17004","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The authors present an improved Sustainability Impact Score with risk-based prioritization and a benchmark comparison, and apply it to two architectural approaches in an energy-sector case study.","lead":"This paper refines a score for measuring how software design choices affect money, nature, and people, and tests it on an energy planning system. The score could help companies report sustainability impacts and catch hidden side effects of technical choices early.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The normalization in Eq. 4 is anchored to the observed minimum SIS per dimension pair, not to an absolute lower bound, so the claimed cross-dimension comparability is not established.","rationale":"The reader's weakest assumption (equal +1/−1 magnitudes) is valid, but the normalization flaw in Eq. 4 is more directly tied to the paper's stated contribution. The abstract and Section 2.3 promise a SIS that is comparable across dimension pairs; the case study's headline conclusions rely on the normalized percentages in Table 3. Because min(SIS) is the observed minimum among the alternatives for each pair, the normalized scores are arbitrarily re-based per pair, and Section 6 explicitly concedes that no lower bound exists. This is an internal inconsistency, not merely a consensus disagreement. A reproduction with a common absolute lower bound would settle whether the reported cross-dimension comparisons survive. Given that the central improvement is currently unsupported, the paper should not be accepted as is; a major revision or rejection is warranted. The DMatrix trade-off observations (e.g., reproducibility negatively affecting resource utilization) may still be valuable, but the quantitative SIS comparison as presented is not trustworthy.","tokens_in":11408,"tokens_out":4566,"duration_ms":42056,"concrete_test":"Recompute the Normalized SIS values in Table 3 using a single absolute lower bound for all dimension pairs—e.g., the theoretical minimum SIS given the priority values in Table 2 and all effects set to −1—and compare the resulting percentages across pairs. If the qualitative conclusions (which dimension pair underperforms, which approach is best) change relative to the published Table 3, the claim that Eq. 4 enables cross-dimension comparison is refuted.","verdict_should_be":"REJECT","load_bearing_attack":"The paper's principal improvement over prior work is that the new SIS is comparable across different dimension pairs (Section 2.3, limitation ii). Equation 4 defines Normalized SIS (%) = (SIS − min(SIS))/(TO(SIS) − min(SIS)) × 100. In Table 3, min(SIS) is evidently the minimum observed SIS among the two alternatives for that pair: for T−Ec, min = −0.425 (Single Model); for T−En, min = −1.05 (Multi Model). Consequently, the 0% point differs between pairs: 0% for T−Ec corresponds to −0.425 raw SIS, while 0% for T−En corresponds to −1.05. Section 6 explicitly states 'we do not have a lower bound for a SIS value,' yet Eq. 4 requires one. As a result, the normalized percentages are not on a common scale and cannot be compared across dimension pairs in the way the paper claims (e.g., 'both approaches underperform across the T-En pair' vs. 'Multi model approach outperforms for economic dimension'). This is a technical error, independent of the equal-magnitude assumption. Even with exact effect magnitudes, the normalization is not invariant.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an improved Sustainability Impact Score (SIS) for software architecture evaluation. The method extends prior work by computing SIS for any pair of sustainability dimensions, deriving QA priorities from stakeholder-provided importance and risk levels via a weighted sum, and normalizing scores relative to a theoretical-optimal DMatrix. The approach is demonstrated on a single industrial case study (MMvIB) comparing a single-model and a multi-model architecture. The paper reports inter-QA trade-offs and normalized SIS percentages for the T-Ec, T-En, T-S, and S-Ec dimension pairs, concluding that the multi-model approach is better for the economic and social dimensions and that both approaches underperform on the environmental dimension.","tokens_in":11675,"tokens_out":4980,"duration_ms":44706,"significance":"If the method were sound, it would provide a lightweight, practitioner-oriented tool for early sustainability assessment in architecture evaluation, complementing heavier scenario-based methods such as ATAM. The risk- and importance-based prioritization is a reasonable adaptation of ATAM's utility tree concept, and the authors ship a replication package, which is a concrete strength. The central limitation claims of the prior SIS (Section 2.3) are addressed in principle. However, the significance is conditional on fixing a load-bearing normalization problem and on treating the single-participant case study as an illustration rather than as evidence of generalizable empirical findings. The paper is honest in acknowledging several threats to validity, but the main contribution as stated, cross-dimension comparability of SIS values, is not currently established.","major_comments":[{"comment":"The normalization in Eq. (4) uses the observed minimum SIS per dimension pair as the lower bound, not a common absolute lower bound. In Table 3, for T-Ec the minimum is -0.425 (Single Model), while for T-En the minimum is -1.05 (Multi Model). Consequently, a normalized value of 0% corresponds to different raw SIS values in different rows and columns, so the normalized percentages are not on a common scale. This directly contradicts the claim in Section 5.3 that 'the normalization allowed us to compare the SIS values across dimensions,' and it fails to overcome limitation (ii) of Section 2.3. The issue is independent of the equal-magnitude assumption: even with exact effect magnitudes, the percentages cannot be compared across pairs because the anchor point varies. Section 6 even states that 'we do not have a lower bound for a SIS value,' so the min(SIS) used in Eq. (4) is only an artifact of the alternatives considered. The authors should either define a theoretical lower bound (e.g., a worst-case DMatrix) that is common across all pairs, or explicitly restrict all comparative claims to within-pair comparisons.","section":"Section 5.3"},{"comment":"The assumption that all QA effects have the same magnitude (+1 or -1) is load-bearing for the reported scores. The authors acknowledge this explicitly: 'one QA may affect another QA more positively or negatively than another QA.' Because all effects are equal-magnitude, the SIS values reduce to weighted counts of positive and negative edges, and the differences between alternatives (e.g., 76.51% vs 28.41% for T-Ec) depend entirely on which edges are present. No sensitivity analysis is provided to show how the rankings would change if magnitudes were differentiated or if the effect signs were assigned differently. At a minimum, the scores should be described as ordinal and the reported percentages should be accompanied by a sensitivity analysis over plausible effect-weight variations.","section":"Section 5.3"},{"comment":"The central empirical findings rest on inputs from a single study participant, who supplied the QA set, the effect signs, the importance and risk levels, and the theoretical-optimal DMatrix. The weights wI and wR are fixed to 0.5 without exploration, and no inter-rater comparison or independent validation is reported. The conclusions such as 'the multi model approach performs better across economic and social dimensions' are deterministic transformations of those inputs, so they cannot be read as robust empirical findings. The paper should either add a sensitivity analysis that varies the inputs within plausible ranges, or explicitly frame the case study as a feasibility demonstration rather than as evidence that these particular trade-offs hold for the MMvIB system generally.","section":"Sections 4 and 5"}],"minor_comments":[{"comment":"There is a typo: 'Is it theocratically possible' should be 'Is it theoretically possible.'","section":"Section 5.3"},{"comment":"The phrase 'containerless provided sustainability support' appears to be a typo; the intended word is likely 'containerization.'","section":"Section 3"},{"comment":"The summation notation 'n,mX' is difficult to read; please use standard double-sum notation such as \\sum_{i=1}^n \\sum_{j=1}^m.","section":"Equations (1) and (2)"},{"comment":"The 'x' marks in the sustainability-dimension columns are not explained; a legend is needed to clarify that they indicate the dimension to which each QA belongs.","section":"Table 2"},{"comment":"The example 'T-Ec with E-Ec' likely contains a typo; it should probably be 'T-Ec with En-Ec' or 'T-Ec with S-Ec,' since no 'E-Ec' dimension appears elsewhere in the paper.","section":"Section 2.3"},{"comment":"The statement 'we do not have a lower bound for a SIS value' conflicts with Eq. (4), which uses min(SIS) as a lower bound. Please state explicitly that min(SIS) is the observed minimum among the compared alternatives, not an absolute lower bound.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The normalization issue in Eq. (4) is the main technical blocker and should be resolved before the paper can be accepted. If it cannot be resolved, the cross-dimension comparability claim, which is the paper's core improvement over prior work, must be withdrawn and the paper reframed as a within-pair comparison tool. I also recommend asking the authors to soften the empirical 'reveals' language given the single-participant basis, or to add a sensitivity analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things before reading this one. First, the paper actually improves on the authors' prior SIS: it generalizes the formula to all dimension pairs, replaces ranking-based priority with a risk/importance-weighted function, and adds a theoretical-optimal benchmark. Second, the main new claim — that normalized SIS values are comparable across dimension pairs — does not hold as presented. The stress-test note is right. Equation 4 uses min(SIS) from the observed alternatives in each pair, so the 0% point means different things across pairs: for T–Ec it is −0.425, for T–En it is −1.05. Section 6 even admits there is no lower bound for SIS, yet Eq. 4 needs one. So the percentages in Table 3 are not on a common scale, and statements like \"both approaches underperform across the T-En pair\" are not supported as cross-dimension comparisons.\n\nWhat the paper does well: the method is simple and reproducible. The authors provide a replication package, the calculations are transparent, and they are upfront about the equal-magnitude assumption (+1/−1 for all effects) and the single-participant case study. The related work is reasonable and the motivation via CSRD reporting is timely. The industrial context — a multi-model energy decision-support system — is genuinely interesting, and the DMap visualization seems useful for surfacing indirect impacts.\n\nThe soft spots beyond normalization: the headline finding that \"technical quality concerns have significant, often unrecognized impacts across sustainability dimensions\" is essentially baked into the modeling process. The SIS is a deterministic function of the same participant-supplied priorities, impacts, and effects that define it, with no external validation or independent benchmark. That doesn't make the work worthless, but it means it is a method demonstration, not a discovery about the case. The equal-magnitude limitation is acknowledged but load-bearing; if effect magnitudes vary, rankings could shift. The paper is honest about these threats, but the abstract overclaims.\n\nWho is this for? Researchers and practitioners in sustainable software engineering who want a lightweight, early-stage architecture evaluation tool. It deserves a serious peer review, but with a required fix: either revise the normalization to use a true absolute lower bound or a defensible reference point, or drop the cross-dimension comparability claim and present the scores as within-pair relative comparisons only. I would send it to review, expecting major revision.\n\nMy verdict: engage, but read the tables critically.","headline":"The proposed SIS extension is a real step forward for sustainability-aware architecture evaluation, but the paper's headline claim of cross-dimension comparability is undermined by a flawed normalization that anchors to observed minima, not an absolute lower bound.","tokens_in":12202,"tokens_out":1339,"would_cite":false,"duration_ms":13901,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Technical quality choices carry measurable sustainability consequences that a weighted impact score can expose.","keywords":["software architecture evaluation","sustainability impact score","quality attribute trade-offs","sustainability dimensions","architecture trade-off analysis","multi-model systems","decision maps","sustainability reporting"],"falsifier":"Recompute the SIS for the single-model and multi-model approaches using effect magnitudes estimated from actual system measurements (for example, the monetary cost or energy use associated with each dependency) instead of uniform $\\pm 1$; if any dimension-pair ranking between the two approaches reverses, the equal-magnitude assumption is not sufficient for the reported comparison.","tokens_in":11225,"feed_emoji":"♻️","tokens_out":8371,"duration_ms":74389,"temperature":0.7,"pith_summary":"This paper claims that sustainability impact at the software architecture level can be quantified with an improved Sustainability Impact Score (SIS), a weighted sum of quality-attribute effects across sustainability dimensions. The authors apply the score to a real energy-sector decision-support system and report that technical quality attributes have significant, often hidden impacts on economic, environmental, and social concerns. If correct, the score gives architects a way to compare design options by their sustainability consequences before implementation, and gives organizations a documented basis for sustainability reporting such as the CSRD regulation. The central pay-off is early visibility: trade-offs that usually surface only as side effects can be made explicit and managed during architecture evaluation.","feed_headline":"A weighted sustainability score exposes hidden software trade-offs","feed_subtitle":"Case study of an energy decision system shows technical quality choices ripple across economic, environmental, and social dimensions.","key_machinery":"The central mechanism is the Sustainability Impact Score (SIS): a weighted sum over dependency matrices that pair quality attributes of two sustainability dimensions, with each cross-effect recorded as $+1$, $-1$, or $0$. Priorities in the sum are derived from a utility matrix that maps scenarios to quality attributes and scores each attribute by importance and risk, normalized to $[0.1,1]$. To make scores comparable across dimension pairs, the paper normalizes each score against a theoretical-optimal dependency matrix in which negative effects are minimized, converting the result to a percentage. This machinery turns subjective architecture reasoning—which quality attributes matter and how they interact—into a single numeric comparison per dimension pair.","core_discovery":"On its own terms, the paper argues that a risk- and importance-weighted Sustainability Impact Score can turn architecture trade-offs into comparable numbers. The score is computed for pairs of sustainability dimensions as $\\mathrm{SIS}_{\\mathrm{dim1},\\mathrm{dim2}} = \\sum_{i,j}(\\mathrm{Priority}_{\\mathrm{dim1},i} + \\mathrm{Priority}_{\\mathrm{dim2},j}) \\times \\mathrm{Impact}_{ij}$, with effects taken as $+1$, $-1$, or $0$ from dependency matrices; priorities come from a weighted combination of each quality attribute's importance and risk. Compared with a theoretical-optimal dependency matrix, normalized scores express how close an architecture comes to fully supporting a dimension. In the MMvIB case, the multi-model approach scores highest for social and economic support while both approaches underperform on the environmental dimension, and the analysis exposes chains such as traceability improving reproducibility, transparency, stakeholder stake, and ultimately monetary cost. The paper's central claim is that these cross-dimension effects are real, often unrecognized, and can be surfaced by the SIS procedure.","pith_inferences":["Beyond the paper, the equal-magnitude effect assumption could be stress-tested by replacing uniform $\\pm 1$ values with measured strengths and re-running the MMvIB comparison; the ranking would hold only if it survives that perturbation.","Beyond the paper, the normalized SIS could be inverted into a design objective, letting architects search over workflow configurations for the highest sustainability support across all dimension pairs.","Beyond the paper, if external benchmarks such as carbon budgets become available, the relative percentages could be anchored to absolute targets, turning the score from a comparison device into a compliance metric.","Beyond the paper, multi-stakeholder projects would need an explicit consensus procedure for risk, importance, and effect values; the utility-matrix structure suggests a group-weighted aggregation as a natural next step."],"forward_implications":["Architects can compare alternative designs by normalized SIS per sustainability dimension pair, not just within one pair.","Enabling and systemic sustainability effects, such as reproducibility requirements increasing resource utilization, become visible at evaluation time instead of after deployment.","High-priority technical attributes that harm another dimension can be explicitly weighed against the consequences of changing them, as in the reproducibility-resource utilization trade-off.","Organizations gain a repeatable, documented path toward sustainability reporting obligations, including preparing for regulations like CSRD.","The relative comparison to a theoretical optimal provides a benchmark for how much sustainability support an architecture is leaving on the table."],"supporting_citations":[{"why":"Earlier SIS design that this paper extends by adding risk- and importance-based prioritization and cross-dimension comparability.","marker":"[14]"},{"why":"Dependency Matrix representation used to record positive, negative, and neutral effects between quality attributes.","marker":"[7]"},{"why":"Architecture trade-off analysis method whose importance and risk levels drive the new priority formulas.","marker":"[19]"},{"why":"Decision Map and Sustainability Quality Model used to elicit quality concerns and map them to sustainability dimensions.","marker":"[20]"},{"why":"Weighted sum model that provides the functional form of the SIS calculation.","marker":"[3]"},{"why":"Defines sustainability as a software quality property with economic, environmental, social, and technical dimensions, the conceptual basis for the score.","marker":"[21]"},{"why":"MMvIB multi-model architecture documentation serving as the industrial case study object.","marker":"[23]"}],"fun_headline_variants":["Sustainability score reveals hidden software architecture trade-offs","Risk-weighted metric quantifies architecture sustainability impacts","New score exposes unseen sustainability effects in software design","How a weighted score turns architecture quality into sustainability numbers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing simplification is that every quality attribute's effect on another has the same magnitude, recorded as $+1$, $-1$, or $0$; if real effects differ in strength, the scores and the ranking of architectures could change.","fun_headline_variants_meta":{"raw":{"variants":["Sustainability score reveals hidden software architecture trade-offs","Risk-weighted metric quantifies architecture sustainability impacts","New score exposes unseen sustainability effects in software design","How a weighted score turns architecture quality into sustainability numbers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000299,"raw_usage":{"total_tokens":1733,"prompt_tokens":957,"completion_tokens":776,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":718}},"tokens_in":573,"tokens_out":776,"duration_ms":6135,"temperature":1.0,"reasoning_tokens":718,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:14:23.580402+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the SIS for the single-model and multi-model approaches using effect magnitudes estimated from actual system measurements (for example, the monetary cost or energy use associated with each dependency) instead of uniform $\\pm 1$; if any dimension-pair ranking between the two approaches reverses, the equal-magnitude assumption is not sufficient for the reported comparison.","supporting_citations":[{"cited_title":"Software Architecture Assessment for Sustainability: A Case Study","cited_arxiv_id":null,"evidence_quote":"Earlier SIS design that this paper extends by adding risk- and importance-based prioritization and cross-dimension comparability."},{"cited_title":"Defining Interdi- mensional Dependencies of the Sustainability-Quality Model","cited_arxiv_id":null,"evidence_quote":"Dependency Matrix representation used to record positive, negative, and neutral effects between quality attributes."},{"cited_title":"Jeromy Carri`ere, and Steven G","cited_arxiv_id":null,"evidence_quote":"Architecture trade-off analysis method whose importance and risk levels drive the new priority formulas."},{"cited_title":"A Handbook on Multi-Attribute Decision-Making Methods","cited_arxiv_id":null,"evidence_quote":"Weighted sum model that provides the functional form of the SIS calculation."},{"cited_title":"Framing Sustainability as a Property of Software Quality","cited_arxiv_id":null,"evidence_quote":"Defines sustainability as a software quality property with economic, environmental, social, and technical dimensions, the conceptual basis for the score."},{"cited_title":"Multimodelling documentation, 2023","cited_arxiv_id":null,"evidence_quote":"MMvIB multi-model architecture documentation serving as the industrial case study object."}],"review_version":1}