{"id":"cb2a7c30-7d89-44f9-8a09-b9d634c5e661","arxiv_id":"1908.04070","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Using survey data from Slovenian CIOs and the OrdEval algorithm, the paper classifies software development disciplines into Kano quality categories linked to IT project net benefits.","lead":"This study of 113 Slovenian IT projects finds that how companies apply software development methods in each discipline relates to project benefits, with testing and deployment acting as must-have quality factors and requirements acquisition as an attractive one. It offers a way to prioritize process improvements using Kano's model of quality.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Kano classification in Table 1 rests on an unformalized visual reading of OrdEval plots; the paper's own Section 6 admits the classification is subjective, leaving the must-be/attractive distinction unsupported.","rationale":"The reader's weakest assumption identifies exactly the fragile step: Kano categories are assigned by visually interpreting OrdEval output without a formal rule or external validation. A careful reading of the manuscript supports this concern and adds specific evidence: the authors explicitly list 'subjective classification' as a limitation and propose future quantitative Kano models to address it. This is not a disagreement with the field's consensus but an internal gap between the data analysis and the conclusions drawn from it. The ReliefF ranking also lacks confidence intervals, so the 'testing strongest' claim is weaker than stated, but the Kano classification is more load-bearing because it is the paper's distinctive contribution and the basis for its prescriptive statements. A simulation-based recovery test would settle whether the OrdEval visualization can be interpreted reliably at this sample size. Since the reader already returned CONDITIONAL with high confidence, my assessment does not change the verdict; it strengthens the condition under which the paper would be acceptable, namely that the authors provide a formal classification rule or external validation.","tokens_in":11202,"tokens_out":3398,"duration_ms":38168,"concrete_test":"Obtain the raw survey data or generate synthetic datasets with n=113, the same 7-point Likert scales, and known Kano categories per discipline; run OrdEval with the same R package and parameters; then classify each output using a pre-registered decision rule (e.g., must-be iff significant downward reinforcement occurs only at low values; one-dimensional iff significant reinforcements appear across mid-range values in both directions; attractive iff significant upward reinforcement occurs only at high values). Compare the recovered categories to the known ground truth; if recovery accuracy is below 80% or the rule does not reproduce Table 1, the visual Kano mapping in Figs. 4-9 is not a reliable basis for the central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is the assignment of software development disciplines to Kano quality categories in Table 1: testing and deployment as must-be, coding and integration as one-dimensional, requirements acquisition as attractive. This assignment is read directly from the OrdEval bar charts in Figs. 4-9 without any stated decision rule, without inter-rater reliability checks, and without validation against the original Kano questionnaire. The text describes which bars are 'statistically significant' but does not define the mapping from patterns of significant reinforcement bars to Kano categories. The paper itself concedes in Section 6 that 'subjective classification' is present in the Kano model and suggests future integration with quantitative Kano models, which is an in-text acknowledgment that the current classification is not objectively grounded. With n=113 and multiple attributes and value transitions tested, the selected 'significant' bars could also be influenced by multiple comparisons; no correction or pre-registered rule is reported. If the visual mapping is unreliable, the must-be versus attractive distinction, and the prescriptive advice about caution in altering testing/deployment versus experimenting with requirements acquisition, is not supported by the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes survey data from 113 CIOs of the largest Slovenian enterprises to relate the perceived quality of software development methodology (SDM) application in individual development disciplines to the net benefits of IT projects. Using the ReliefF algorithm for attribute ranking and the OrdEval algorithm to characterize nonlinear, asymmetric effects of individual attribute values on the outcome, the authors report that Testing has the strongest positive association with net benefits, followed by Coding and integration, while Project management shows no positive association. Based on visual interpretation of OrdEval reinforcement plots, they assign disciplines to Kano quality categories: Testing and Deployment as must-be, Coding and integration as one-dimensional, Requirements acquisition as attractive, with System design and architecture inconclusive and Project management inconclusive on average but attractive for a subgroup. The paper draws prescriptive conclusions, e.g., that enterprises should be cautious when altering testing and deployment but may experiment with requirements acquisition.","tokens_in":11382,"tokens_out":2834,"duration_ms":29270,"significance":"If the findings hold, the paper offers a practical approach to evaluating software development disciplines through a nonlinear quality lens, which is a genuine contribution to IT project management and Kano-model research. The study uses well-established algorithms (ReliefF, OrdEval) implemented in a public R package, and the authors explicitly acknowledge the non-random sample and the retrospective survey design as limitations. The core correlational claim—that SDM application in some disciplines is associated with project net benefits—is plausible and consistent with prior work. However, the more distinctive contribution, the Kano classification in Table 1, currently rests on an unformalized visual reading of OrdEval plots, with no decision rule, no inter-rater reliability check, and no validation against an independent Kano questionnaire. Given that the prescriptive advice depends directly on this classification, the significance is conditional on addressing that methodological gap.","major_comments":[{"comment":"The assignment of disciplines to Kano quality categories is based on a visual interpretation of statistically significant reinforcement bars, but no formal decision rule is stated that maps a pattern of significant upward/downward reinforcements to a Kano category (e.g., must-be vs. attractive). The text in Section 6 concedes that subjective classification is present in the Kano model and suggests future integration with quantitative Kano models. Because the prescriptive conclusions about caution in altering testing/deployment versus experimenting with requirements acquisition depend directly on this classification, the manuscript needs a reproducible rule, an inter-rater reliability check, or a validation against the original Kano questionnaire on a subsample. Without this, the must-be/attractive distinction is not supported beyond the authors' visual impression.","section":"Section 4, Table 1, Figs. 4-9"},{"comment":"The claim that 'Testing discipline has the strongest association' with net benefits, with Coding and integration second, is reported without any uncertainty intervals or significance tests on the ReliefF scores. With n = 113 and six disciplines, the differences between adjacent ranks may be within sampling noise. Please provide bootstrap confidence intervals or a permutation test for the ReliefF values, and address the multiple-comparisons issue when interpreting the ranking and the OrdEval significance bars.","section":"Section 4, Fig. 3"},{"comment":"The OrdEval hyperparameters (e.g., number of nearest neighbors, number of iterations, missing-value handling parameters) are not reported. Since the OrdEval visualization and the resulting Kano classification depend on these settings, the reproducibility of the central classification requires reporting the exact parameter values and, ideally, a sensitivity analysis showing that the classification is stable across reasonable hyperparameter choices.","section":"Section 3"}],"minor_comments":[{"comment":"The conclusion uses causal language such as 'Testing providing the highest net benefits' and 'increasing its quality will result in a stable growth of benefits,' but the study is cross-sectional and observational. Please temper these statements to reflect association rather than causation.","section":"Section 6, Conclusion"},{"comment":"The caption states an assumption that the scales are 7-point Likert, which is a modeling assumption for the idealized curves; clarify that this is an illustrative assumption and not a property of the collected data.","section":"Introduction, Fig. 1 caption"},{"comment":"The text describes reinforcement bars as 'statistically significant' without stating that significance is determined by bars extending beyond the box-and-whiskers whiskers at the 95% level. Please state this criterion explicitly in the results section or in the figure captions.","section":"Section 4, Figs. 4-9"},{"comment":"The paper alternates between 'CIOs' satisfaction with the application of SDM in a discipline' and 'net benefits of IT projects' as the response variable. Clarify which variable is used as the OrdEval/ReliefF response in the analysis, as this is essential for interpreting the results.","section":"Section 2, OrdEval description"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be a preprint of a published article in Business & Information Systems Engineering, which may affect the framing of novelty in the cover letter. The main technical concern—the unsupported visual Kano classification—is correctable in principle, so I recommend major revision rather than rejection, provided the authors can formalize the classification procedure or substantially hedge the prescriptive claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the empirical result is plausible but the headline classification is soft. The genuinely new bit is applying OrdEval to CIO survey data to sort software development disciplines into Kano categories, and the specific findings for Slovenian enterprises. ReliefF ranking and the basic association pattern are plausible. The authors explain OrdEval clearly, and they are honest about the survey's limitations.\n\nThe core problem is exactly where the stress-test puts it. Table 1 is produced by eyeballing the OrdEval bar charts in Figs. 4-9. No formal rule maps significant reinforcement bars to Kano categories, no validation against the original Kano questionnaire, no inter-rater check. The paper itself admits in Section 6 that subjective classification is present. With n=113 and many possible value transitions, the few bars that pass the bootstrap confidence intervals could easily be noise; there is no multiple-comparison correction. So the must-be versus attractive distinction is not actually supported.\n\nThat said, the paper is not careless. The claims are phrased as indications, and the limitations section is upfront. The ReliefF ranking lacks error bars, but that is a minor issue because the ranking is used qualitatively. The bigger issue is the Kano mapping. If the authors shared data and OrdEval parameters and defined an explicit decision rule, the classification could be checked. Without that, I would treat the categories as hypotheses, not results.\n\nBottom line: this is a legitimate exploratory study that a serious referee should engage with. The method is interesting, the data are what they are, and the central claim about association is fine. The Kano labels need to be downgraded. I would bring it to a reading group looking at survey methods in SE, but I would not cite the specific Kano assignments as established. Accept with major revision if the mapping is made reproducible.","headline":"Plausible empirical study; the Kano classification is read off plots by eye and needs a formal rule to be credible.","tokens_in":11922,"tokens_out":2467,"would_cite":false,"duration_ms":27346,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that software development disciplines fall into distinct Kano quality categories, with testing most strongly tied to IT project net benefits.","keywords":["Kano model","software development disciplines","IT project net benefits","OrdEval","ReliefF","CIO survey","software process quality"],"falsifier":"Re-survey the same type of respondents with Kano's original paired functional and dysfunctional questions for each discipline, or apply an explicit scoring rule to the OrdEval reinforcement bars, and check whether testing and deployment still classify as must-be, coding as one-dimensional, and requirements acquisition as attractive.","tokens_in":10976,"feed_emoji":"📊","tokens_out":5081,"duration_ms":48410,"temperature":0.7,"pith_summary":"This paper claims that how thoroughly a software development method is applied in individual disciplines—requirements acquisition, design, coding and integration, testing, deployment, and project management—is measurably tied to the net benefits of IT projects. Using survey responses from 113 CIOs of large Slovenian enterprises, it ranks disciplines by strength of association and then classifies each discipline by Kano's quality categories. The central result is that testing shows the strongest association with net benefits and is a must-be quality, deployment is also must-be, coding and integration is one-dimensional, and requirements acquisition is attractive. If true, managers can treat changes to testing and deployment routines as risky disruptions, while experimenting with requirements acquisition is comparatively safe.","feed_headline":"Testing discipline most strongly tied to IT project benefits","feed_subtitle":"CIO survey ranks process disciplines by Kano quality: testing and deployment are must-be, requirements acquisition is attractive.","key_machinery":"The central machinery is the pair of attribute-evaluation algorithms ReliefF and OrdEval. ReliefF ranks disciplines by the strength of their association with net benefits while taking attribute interdependencies into account. OrdEval, an ordinal attribute evaluator built on ReliefF, computes upward and downward reinforcement factors for each value of a discipline's satisfaction score; the pattern of statistically significant bars in its visual output indicates whether a discipline behaves as must-be, one-dimensional, attractive, indifferent, or reverse Kano quality. The Kano model supplies the interpretive categories, and the paper's key analytic step is reading OrdEval bar plots to assign each discipline to a Kano category.","core_discovery":"The paper's central claim is that software development disciplines differ in Kano quality type according to CIOs' perception of how well the discipline's methodology is applied. Based on OrdEval reinforcement patterns, testing and deployment are must-be qualities, coding and integration is one-dimensional, requirements acquisition is attractive, and system design and architecture as well as project management are inconclusive on average. The paper also claims that ReliefF ranking shows testing has the strongest positive association with project net benefits, coding and integration ranks second, and project management shows no positive association on average.","pith_inferences":["Going beyond the paper, the same OrdEval-Kano mapping could be validated by administering the original Kano questionnaire alongside the short survey; if the two classifications agree, the short instrument could replace the longer one in process-quality audits.","Going beyond the paper, the must-be versus attractive split suggests a resource-allocation rule: protect the must-be disciplines' baseline routines, then invest surplus improvement effort in attractive or one-dimensional disciplines.","Going beyond the paper, the ranking is based on CIO perception in large Slovenian enterprises; a direct test would repeat the survey in mid-size firms or other countries to see whether the Kano types and the testing-first ranking persist."],"forward_implications":["Testing is the discipline whose application quality is most strongly tied to IT project net benefits, so improvement efforts aimed at testing are the likeliest to pay off.","Testing and deployment are must-be quality: cutting back on their established methodology risks significant dissatisfaction and lost benefits, while pushing them beyond basic levels yields little extra satisfaction.","Coding and integration is one-dimensional: raising its methodology quality should produce a steady, proportional increase in benefits.","Requirements acquisition is attractive: adopting improved techniques can raise satisfaction and benefits substantially, with less risk of disrupting routines.","Project management shows no average positive association with net benefits, although a subgroup of CIOs perceives it as attractive."],"supporting_citations":[{"why":"Supplies the Kano model and its five quality categories that the paper uses to classify software development disciplines.","marker":"Kano et al. 1984"},{"why":"Introduces the OrdEval algorithm whose reinforcement-factor visualizations are the basis for assigning each discipline to a Kano category.","marker":"Robnik-Šikonja & Vanhoof 2007"},{"why":"Introduces ReliefF, the attribute-evaluation measure that produces the discipline ranking headed by testing.","marker":"Robnik-Šikonja & Kononenko 2003"},{"why":"Provides the precedent that OrdEval can recover Kano categories from a shorter one-question-per-attribute survey.","marker":"Čufar et al. 2015"},{"why":"Defines the Rational Unified Process disciplines used to structure the survey into requirements acquisition, design, coding, testing, and deployment.","marker":"Kruchten 2009"},{"why":"Reviews two decades of Kano applications and documents the complexity of the original Kano questionnaire that motivates the short-survey approach.","marker":"Lofgren & Witell 2008"},{"why":"Supports the interpretation that project management methodology is perceived as helpful by only some managers, explaining the inconclusive average result.","marker":"Wells 2012"}],"fun_headline_variants":["Test discipline drives IT project benefits most","Must-be disciplines: testing and deployment, CIOs say","Coding second, testing first in IT benefit rank","CIOs: testing is must-have for project payoff","Testing tops CIO ranking for IT project benefits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the visual pattern of statistically significant bars in the OrdEval plots maps unambiguously onto Kano's theoretical categories; the paper provides no formal rule or independent validation for that mapping.","fun_headline_variants_meta":{"raw":{"variants":["Test discipline drives IT project benefits most","Must-be disciplines: testing and deployment, CIOs say","Coding second, testing first in IT benefit rank","CIOs: testing is must-have for project payoff","Testing tops CIO ranking for IT project benefits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000336,"raw_usage":{"total_tokens":1803,"prompt_tokens":829,"completion_tokens":974,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":902}},"tokens_in":445,"tokens_out":974,"duration_ms":8339,"temperature":1.0,"reasoning_tokens":902,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:51:49.106798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-survey the same type of respondents with Kano's original paired functional and dysfunctional questions for each discipline, or apply an explicit scoring rule to the OrdEval reinforcement bars, and check whether testing and deployment still classify as must-be, coding as one-dimensional, and requirements acquisition as attractive.","supporting_citations":[],"review_version":1}