{"id":"04cba1fe-e726-4678-b4d9-4c15786e87a5","arxiv_id":"1908.06350","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"An empirical test of a modified IS Success Model for GCC e-commerce smartphone apps reports support for all hypotheses, but its central claim about service quality contradicts its own path coefficients.","lead":"This paper adapts a well-known information systems success model to smartphone shopping apps in the Gulf region and tests it with 803 survey responses. It claims customer service features matter most, but the statistical tables in the paper do not support that headline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim is contradicted by the paper's own path coefficients: Table 8 shows IQ→US=0.472 vs SQU→US=0.428 and IQ→IU=0.383 vs SQU→IU=0.339, so Service Quality is not ranked above Information Quality.","rationale":"The reader's weakest_assumption concerns the inadequate fit of the Information Quality measurement model in Table 7. That is a real problem: if IQ is misspecified, its large path coefficients may be inflated. However, the single most load-bearing concern is more direct: even taking Table 8 at face value, the paper's central claim that Service Quality is more significant than Information Quality and System Quality is contradicted by its own estimates. The IQ fit problem reinforces the rejection but is not necessary for it. The reader's rationale does note the contradiction between the headline and the path coefficients, so there is partial agreement, but the weakest_assumption selected is different. The appropriate disposition remains rejection: the paper's headline overstates what its evidence shows, and a correction of the headline alone would not resolve the other reporting and sample-generalization issues. No judgment about author intent is made; the concern is purely about the relation between the reported results and the central claim.","tokens_in":19603,"tokens_out":4179,"duration_ms":41823,"concrete_test":"Construct a comparison of direct and total standardized effects from Table 8 for the three quality constructs on US and IU. For each quality construct, compute total effect on IU as direct IQ→IU (or SQU→IU) plus (path to US × 0.313). If IQ total ≥ SQU total for both outcomes, as the reported coefficients imply, the abstract's claim that Service Quality is more important than Information Quality is unsupported. Alternatively, ask the authors to identify the specific reported coefficient that ranks Service Quality above Information Quality and System Quality at the construct level; any such coefficient must be pointed out explicitly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"For the abstract's 'significance of Service Quality ... over Information Quality and System Quality' to hold, the standardized construct-level effects of Service Quality (SQU) on User Satisfaction (US) and Intention to Use (IU) would need to exceed those of Information Quality (IQ) and System Quality (SQ). The paper's own Table 8 reports the opposite ordering: IQ→US=0.472 vs SQU→US=0.428; IQ→IU=0.383 vs SQU→IU=0.339; SQ→US=0.328 and SQ→IU=0.363. Table 6 shows the same pattern in correlations (IQ-US=0.472 vs SQU-US=0.428; IQ-IU=0.283 vs SQU-IU=0.239). Adding the indirect path through US→IU (0.313) does not reverse it: total IQ→IU is 0.383 + 0.472*0.313 = 0.531, while total SQU→IU is 0.339 + 0.428*0.313 = 0.473. The only table entries favoring Service Quality are first-order loadings of service sub-constructs on SQU (e.g., CF_HT→SQU=0.581), but those are not construct-level comparisons and do not establish that Service Quality ranks over the other quality constructs. This internal contradiction is independent of the Information Quality fit problems in Table 7, and it directly undermines the paper's stated headline.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript adapts the DeLone and McLean IS Success Model to e-commerce smartphone applications in the Gulf Cooperation Council (GCC) region. It decomposes System Quality, Information Quality, and Service Quality into sub-constructs derived from prior literature, and links these to User Satisfaction, Intention to Use, and Net Benefits. The model is tested with survey responses from 803 participants from Saudi Arabia, the UAE, and Qatar, using exploratory factor analysis, confirmatory factor analysis, and structural path estimation. The abstract and conclusion claim that Service Quality is the most significant quality dimension and that User Satisfaction is the central mediating construct. All fifteen hypotheses are reported as accepted on the basis of Table 8.","tokens_in":19942,"tokens_out":3240,"duration_ms":33707,"significance":"If the empirical claims were reliable, the paper would offer a region-specific, theory-grounded instrument for designers and researchers concerned with e-commerce smartphone applications in the GCC. The study has genuine strengths: a comparatively large sample, explicit questionnaire arbitration, and a measurement model that is reported in sufficient detail for the reader to check the claimed fit indices. However, the headline finding is contradicted by the paper's own path coefficients, and the construct with the largest estimated effects (Information Quality) fails the paper's stated measurement model fit criteria. These problems are load-bearing: they concern the central claim rather than presentation. The paper also has no reproducibility artifacts, so the evidence base is limited to what is reported in the tables.","major_comments":[{"comment":"The abstract states that Service Quality is significant 'over Information Quality and System Quality,' but Table 8 reports standardized path coefficients IQ→US = 0.472 versus SQU→US = 0.428, and IQ→IU = 0.383 versus SQU→IU = 0.339; System Quality is lower still. Even after adding the indirect path through US→IU (0.313), the total effect of IQ on IU is 0.383 + 0.472×0.313 = 0.531, while the total effect of SQU on IU is 0.339 + 0.428×0.313 = 0.473. The paper's own results therefore rank Information Quality above Service Quality, contradicting the abstract, the discussion, and the conclusion. This is not a minor wording issue; it reverses the stated headline finding.","section":"Abstract and Section 4.4 (Table 8)"},{"comment":"The Information Quality measurement model fails most of the fit criteria that the paper itself adopts in Table 5: X²/df = 4.112 exceeds the 3.0 threshold, and GFI (0.873), TLI (0.833), NFI (0.816), CFI (0.853), and IFI (0.854) are all below 0.90. The text nevertheless states that 'the results show proportional model at recommended values.' Because Information Quality is the construct with the largest structural path coefficients, its measurement model misspecification directly undermines the reliability of the path estimates used in the paper's ranking of factors. The authors need to re-specify or justify this measurement model before any claims about relative importance can be accepted.","section":"Section 4.3, Table 7"},{"comment":"The same dataset is used first to refine the measurement model through EFA (including the elimination of 17 items and the formation of sub-constructs) and then to evaluate that refined model through CFA and hypothesis testing. This procedure capitalizes on sample-specific chance variation, and the reported fit indices are therefore not an independent confirmation of the model. A split-sample analysis, cross-validation, or an explicit acknowledgment of the exploratory nature of the results is needed before the measurement model can be treated as validated.","section":"Sections 4.2 and 4.3"},{"comment":"The paper claims to study GCC consumers, but the reported sample consists only of respondents from Saudi Arabia (48%), the UAE (31%), and Qatar (21%). Bahrain, Kuwait, and Oman are absent. As a result, the findings at best describe three GCC countries, and the generalization to 'the GCC' in the title, abstract, and conclusion is not supported by the presented sampling information.","section":"Section 3 and Table 3"}],"minor_comments":[{"comment":"The Likert scale is described as ranging from 1 (strongly agree) to 5 (strongly disagree), yet all reported means are above 4 for most constructs. If the scale direction is as stated, the means imply that respondents disagreed with positive statements, which is inconsistent with the positive factor loadings and the interpretation of high satisfaction. The authors should correct the scale anchor or clarify that items were reverse-scored.","section":"Section 3, survey instrument"},{"comment":"The hypotheses H6a through H6d are introduced as 'the following information quality hypotheses,' but they concern Service Quality sub-constructs, not Information Quality. This labeling error should be fixed.","section":"Section 2, hypotheses"},{"comment":"The header row reads 'RMR ≤ 0.8' and 'RMSEA ≤ 0.8'; the standard criteria are RMR/RMSEA below 0.08, as also implied by Table 5. The typo should be corrected to avoid confusion.","section":"Table 7"},{"comment":"The table headers 'Indicators Removed' and 'Factors Extracted' are unclear, and the surrounding text says that item loadings of 0.5 or more were retained while also stating that seventeen items were eliminated. A clearer reconciliation of the number of indicators before and after EFA is needed.","section":"Table 4 and Section 4.2"},{"comment":"The discussion refers to 'twenty functions' belonging to the six main groups, while Section 2 describes fifteen sub-constructs and Table 4 lists different counts. The relationship between indicators, sub-constructs, and functions is not explained and should be clarified.","section":"Section 5, Discussion"},{"comment":"The manuscript would benefit from a Limitations section, particularly regarding the non-representative sample coverage within the GCC, the absence of split-sample validation, and the borderline measurement model fit for Information Quality.","section":"General"}],"recommendation":"reject","confidential_remarks":"The paper's central claim is internally contradicted by its own reported path coefficients, and the measurement model of the construct that would be ranked first is not adequate under the paper's criteria. These are not cosmetic issues; they affect the main conclusion. I also note the unusually high number of references to the authors' own previous work (e.g., [4], [9], [10], [12], [37], [85]), which contributes little to the construct validation. If the authors resubmit with a re-analysis, the editor may wish to require independent data or at least a split-sample validation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the thing: the paper's own numbers contradict its abstract. Table 8 gives IQ→US = .472, SQU→US = .428; IQ→IU = .383, SQU→IU = .339. Even adding the indirect path through US→IU doesn't reverse the ordering. So the claim that Service Quality is 'significant over Information Quality and System Quality' is false as stated. That is not a debatable interpretive nuance; it's the headline.\n\nWhat is genuinely useful: the authors assemble a broad indicator set for smartphone e-commerce apps and map it onto DeLone and McLean's IS success model, with separate sub-constructs for appearance, organization, content, interaction, assurance, and customer focus. For GCC practitioners, that checklist could be a reasonable starting point. The sample of 803 is large enough for the EFA/CFA machinery they use, and the path coefficients and fit indices are reported in enough detail to check their work—which is exactly how I found the contradiction.\n\nSoft spots, in order of severity.\n\nFirst, the abstract overclaims. If IQ has the strongest paths to both US and IU, the paper should say that. The conclusion does note that factors like technical support, credibility build, and adaptability have higher path coefficients, but those are first-order loadings on latent constructs, not construct-level effects. That conflation carries the abstract's ranking and it doesn't survive Table 8.\n\nSecond, the Information Quality measurement model does not fit by the paper's own criteria. Table 7 reports X²/df = 4.112 (cutoff <3), GFI = .873, TLI = .833, NFI = .816, CFI = .853—all below .90—yet Section 4.3 claims the model is at recommended values. Since IQ has the largest structural paths, this is load-bearing. The authors should respectify the IQ model or report the poor fit as a limitation.\n\nThird, the sample: 92% bachelor's degree or higher, drawn from only three of six GCC states, recruited through social networks. It is a convenience sample, and the title's GCC-wide generalization is not supported. That is a standard limitation, but it should be stated plainly.\n\nThe heavy self-citation pattern is not itself damning; several earlier papers are direct precursors on the same e-commerce context. A reviewer might still ask whether the indicator list was generated primarily from the authors' own prior work rather than an independent literature synthesis.\n\nWho is this for? A GCC app-development team wanting a feature checklist could get some value; an IS researcher will mostly see a re-estimation of known relationships with a reporting problem. I would send it to referees rather than desk-reject—there is real data and a clear model—but the current version's headline must be corrected and the IQ fit issue honestly addressed before publication. I would not cite the abstract's ranking as it stands.","headline":"The GCC e-commerce adaptation of DeLone and McLean has a genuinely useful indicator inventory, but the abstract's central ranking claim is contradicted by the paper's own Table 8.","tokens_in":20525,"tokens_out":2495,"would_cite":false,"duration_ms":26324,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"In GCC e-commerce apps, service quality—loyalty, chat, support, credibility—is the decisive success factor, with user satisfaction as the mediating hub.","keywords":["IS Success Model","m-commerce","e-commerce smartphone apps","GCC","Service Quality","User Satisfaction","customer loyalty","confirmatory factor analysis"],"falsifier":"Re-estimate the structural model after freeing or dropping the Information Quality items that produced the weak fit in Table 7; if the Information Quality → User Satisfaction coefficient (0.472) falls below the Service Quality coefficient (0.428), or the relative ranking reverses, the paper's headline ranking collapses. Independently, track actual repeat purchases and support-ticket behavior in a GCC e-commerce app and test whether service-quality perceptions predict retention better than information or system quality.","tokens_in":19347,"feed_emoji":"📱","tokens_out":9774,"duration_ms":94122,"temperature":0.7,"pith_summary":"The paper asks what makes an e-commerce smartphone application succeed for customers in the Gulf Cooperation Council (GCC) countries. It re-specifies the IS Success Model—a standard framework that links system, information, and service quality to satisfaction, use, and net benefits—to the smartphone shopping context, expanding the three quality dimensions into fifteen measurable sub-constructs. The model is then tested on 803 smartphone shoppers from Saudi Arabia, the United Arab Emirates, and Qatar, with the data analyzed through exploratory and confirmatory factor analysis and path estimation. The paper's central claim is that Service Quality, defined through mobile-loyalty building, customer chat and feedback, help and technical support, and credibility and reliability building, is more consequential than Information Quality and System Quality, and that User Satisfaction is the construct that turns quality into intention to use and net benefits. If this is right, retailers entering the GCC market should treat in-app customer service and trust features as core investments, not peripheral additions.","feed_headline":"GCC shoppers rank customer service above app polish","feed_subtitle":"803 shoppers in Saudi Arabia, UAE, and Qatar point to help, chat, and credibility as the loyalty drivers.","key_machinery":"The load-bearing object is a modified IS Success Model, a standard framework for explaining information-system success through quality, satisfaction, use, and net benefits, here re-specified for GCC e-commerce smartphone applications. The model's detail is a set of fifteen sub-constructs under the three quality dimensions: System Quality (attractive appearance and balancing, color and text usage, planning and consistency, navigation links), Information Quality (updating content and relevant information, accurate and relevant data, content display, multimedia adoption, adaptability, customer advisor, assurance), and Service Quality (mobile-loyalty building, customer chat and feedback, help and technical support, credibility and reliability build). The mechanism that carries the conclusion is a 97-item survey instrument assessed with exploratory factor analysis, confirmatory factor analysis, composite reliability, average variance extracted, and discriminant-validity checks, followed by a path model running from the three quality constructs to User Satisfaction and then to Intention to Use and Net Benefits.","core_discovery":"On its own terms, the paper's discovery is an empirical ranking inside the IS Success Model for GCC e-commerce smartphone apps: the three quality dimensions do not weigh equally, and Service Quality is the dimension the authors conclude is most significant. The paper builds Service Quality from four customer-focus sub-constructs—mobile-loyalty building, customer chat and feedback, help and technical support, and credibility and reliability building—and reports these service features feeding User Satisfaction, which then drives Intention to Use and Net Benefits. Help and Technical Support and Credibility and Reliability Build are singled out as the strongest measured indicators in the study. The authors present User Satisfaction as the pivotal mediating construct, and they support all twenty-four hypotheses in the modified model.","pith_inferences":["The sample covers three of the six GCC states and skews toward university-educated, monthly online shoppers, so the framework is most directly supported for Saudi Arabia, the UAE, and Qatar; testing in the remaining GCC states and in less-educated or less-frequent shopper segments would show how far the ranking generalizes.","A natural causal extension is an A/B experiment inside a live GCC shopping app: switching on live chat, help-desk access, and loyalty rewards while holding information and system features constant should raise satisfaction and repeat-purchase intent if the model is right.","The paper's evidence is self-reported satisfaction and intention; linking the same constructs to behavioural data such as retention, purchase frequency, and support-ticket resolution would turn the ranking into a predictive tool.","Re-specifying the Information Quality measurement model to meet the paper's own fit thresholds could change its path coefficients, so the practical message is most secure for service features while the exact coefficient ordering should be treated as provisional."],"forward_implications":["Developers of GCC e-commerce apps should put service features first: loyalty programmes, in-app chat and feedback, help and technical support, and visible credibility signals.","User Satisfaction should be tracked as the key intermediate outcome, because the model routes the quality dimensions through it before intention to use and net benefits.","The fifteen-sub-construct instrument gives retailers and researchers a ready-made questionnaire for evaluating an e-commerce app in this region.","Net Benefits are predicted by both User Satisfaction and Intention to Use, so success measures should include repeat-purchase intent and perceived value, not just download or visit counts."],"supporting_citations":[{"why":"Supplies the base IS Success Model with the six dimensions that the paper modifies for m-commerce.","marker":"[18]"},{"why":"Provides the dimensions, measures, and interrelationships used to operationalize the IS success constructs.","marker":"[19]"},{"why":"Gives the e-commerce quality dimensions and the quality-satisfaction relationship that the study re-examines.","marker":"[45]"},{"why":"Supports the one-directional causal relationship from User Satisfaction to Intention to Use used in the model.","marker":"[60]"},{"why":"Supplies the exploratory and confirmatory factor analysis criteria and model-fit thresholds used throughout the analysis.","marker":"[92]"},{"why":"Supplies the recommended cutoff values for model-fit indices and correlation thresholds used to accept the measurement model.","marker":"[94]"},{"why":"Provides the Fornell-Larcker criteria for convergent and discriminant validity that validate the constructs.","marker":"[100]"}],"fun_headline_variants":["Service quality beats app polish for GCC shoppers","Help and credibility drive GCC app loyalty","GCC shoppers want support, not just slick apps","Quality service outranks app design in GCC","User satisfaction hinges on service in GCC apps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the questionnaire items grouped under 'Information Quality' really form one coherent quality concept, even though Table 7 reports fit statistics for that construct below the paper's own thresholds (X²/df = 4.112, GFI = 0.873, TLI = 0.833, NFI = 0.816, CFI = 0.853) — if that measurement is misspecified, the path coefficients in Table 8 and the paper's ranking of Service Quality over Information Quality cannot be trusted.","fun_headline_variants_meta":{"raw":{"variants":["Service quality beats app polish for GCC shoppers","Help and credibility drive GCC app loyalty","GCC shoppers want support, not just slick apps","Quality service outranks app design in GCC","User satisfaction hinges on service in GCC apps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1271,"prompt_tokens":877,"completion_tokens":394,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":493,"completion_tokens_details":{"reasoning_tokens":326}},"tokens_in":493,"tokens_out":394,"duration_ms":4268,"temperature":1.0,"reasoning_tokens":326,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:48:27.737122+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the structural model after freeing or dropping the Information Quality items that produced the weak fit in Table 7; if the Information Quality → User Satisfaction coefficient (0.472) falls below the Service Quality coefficient (0.428), or the relative ranking reverses, the paper's headline ranking collapses. Independently, track actual repeat purchases and support-ticket behavior in a GCC e-commerce app and test whether service-quality perceptions predict retention better than information or system quality.","supporting_citations":[{"cited_title":"H., & Mclean, E","cited_arxiv_id":null,"evidence_quote":"Gives the e-commerce quality dimensions and the quality-satisfaction relationship that the study re-examines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the one-directional causal relationship from User Satisfaction to Intention to Use used in the model."},{"cited_title":"F., Black, W., Babin, B., & Anderson, R","cited_arxiv_id":null,"evidence_quote":"Supplies the exploratory and confirmatory factor analysis criteria and model-fit thresholds used throughout the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the recommended cutoff values for model-fit indices and correlation thresholds used to accept the measurement model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Fornell-Larcker criteria for convergent and discriminant validity that validate the constructs."}],"review_version":1}