{"id":"0fbfde68-3049-458c-84ab-685585c0817e","arxiv_id":"2607.29089","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Enterprise AI value is gated by deployment friction, not model capability; the paper introduces the Deployment Wall, the Seam Index (0-12), and Deployment Debt to locate, score, and cost that friction.","lead":"This paper argues that most enterprise AI pilots fail because of deployment friction — integration, governance, identity, workflow change — not because the models are too weak, and offers a 0-12 \"Seam Index\" for scoring AI platforms by how much of that friction they remove. If the diagnosis is right, it redirects how enterprises choose AI platforms and where vendors compete.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Seam Index lacks construct validity: change-management and governance seams are adopter-side, not platform-inheritable, and no anchor for a '2' is defined — P1 is untestable without this fix.","rationale":"The paper is a well-structured design-science contribution with transparent limitations; I read it as a proposal, not an empirical demonstration. The central assertion is that enterprise GenAI failure is caused by deployment friction and that the Seam Index operationalizes this. The weakest point is the index's construct validity: the instrument assigns equal, unit-weight scores to six seams, but the seams are not all platform attributes. Change management is an organizational process; governance has both platform and organizational components. The paper's own evidence (Sec. 5) says change management is the most value-correlated seam, yet §8.1 provides no anchor for a platform 'inheriting' it. Thus P1 — the flagship falsifiable claim — is not yet testable. The absence of any psychometric validation, inter-rater reliability, or competing-factor analysis compounds this. However, the paper explicitly proposes a research agenda (Sec. 11) and discloses its limitations (Sec. 12), so a conditional verdict (fix the instrument's specification and validate) is appropriate. My concern does not move the verdict from the reader's CONDITIONAL.","tokens_in":12145,"tokens_out":5957,"duration_ms":68422,"concrete_test":"Run a construct-specification check: attempt to apply the §8.1 protocol to the change-management seam for a real platform (e.g., Microsoft 365 Copilot or a generic LLM gateway). Require three independent raters to produce an artifact-based score of 0, 1, or 2. If no concrete, pre-registered anchor for a 'native change-management' score exists, raters will be unable to justify a 2 and will rely on ad hoc judgment; this would falsify the reproducibility claim and invalidate P1 as currently defined. Optional quantitative follow-up: content-validity panel classifying each seam as platform-removable vs adopter-resolvable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing issue is not just missing validation data; it is that the Seam Index, as specified, may not measure what it claims. Section 8 scores every seam 0–2 according to whether the platform 'removes' friction natively, but at least two seams — change management and (arguably) governance — are primarily adopter-side organizational capabilities, not platform attributes. The paper itself states change management is 'the seam most correlated with realized value' (Sec. 5) yet supplies no evidence anchor for what a native score of 2 means for it (Sec. 8.1 lists anchors for data, identity, governance, but not change management). Consequently, the Index conflates platform architecture with organizational readiness; P1's predictor is not well-defined. In addition, the equal weighting of six seams with no factorial or reliability support, and the 'mechanical reproduction' of the 95% survival rate via hand-set attrition rates (Sec. 4, Fig. 2), mean the central promise — making the bottleneck measurable — is currently supported by assertion, not by a validated instrument.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the dominant cause of enterprise generative-AI failure is deployment friction rather than model capability. It introduces three linked constructs: the Deployment Wall (a six-stage survival funnel), the Seam Index (a 0–12 platform-scoring instrument across six 'seams'), and Deployment Debt (a compounding liability reframing). It grounds these in a synthesis of field studies and the author's own engagements, explicitly labels figures as illustrative, and proposes six falsifiable propositions. The paper is presented as a design-oriented conceptual contribution, not an empirical validation, and its Limitations section (Sec. 12) acknowledges the untested status of the framework.","tokens_in":12357,"tokens_out":2920,"duration_ms":35450,"significance":"If the framework held up empirically, it would be genuinely useful: it converts platform procurement from benchmark comparison to architecture comparison, and its propositions (P1–P6) are stated in falsifiable form. The paper is unusually honest, with explicit boundary conditions, attributed single-source figures, and a planned research agenda. However, the central instrument—the Seam Index—currently lacks construct-validity evidence, and the DeployWall's 'mechanical reproduction' of the 95% survival rate is a local calibration rather than an independent test. The paper's practical value therefore depends on validation work it does not yet provide.","major_comments":[{"comment":"The scoring protocol promises evidence anchors for each seam, but only data, identity, and governance are exemplified; change management—explicitly called 'the seam most correlated with realized value' in Sec. 5—has no anchor defining what a native score of 2 means. Since change management is an adopter-side organizational capability, not a platform-inheritable attribute, the Seam Index conflates platform architecture with organizational readiness. Without an anchor and a rule for who owns this seam, P1's predictor is not well-defined and the instrument is not reproducible for the seam that matters most.","section":"Sec. 8.1, Table 4"},{"comment":"The 'mechanical reproduction' of the observed ~5% survival rate is achieved by hand-setting per-stage attrition rates so that the funnel lands on the target outcome. The paper acknowledges the rates are illustrative (Sec. 12), but the abstract and Sec. 4 use 'mechanically reproduces' and 'reproduces, mechanically' as if the model were independently confirmed. This is a fit-to-outcome calibration, not a reproduction. Either stage-level survival data are reported, or the wording should be changed to 'consistent with field survival rates'—otherwise the Wall's apparent precision is misleading.","section":"Sec. 4, Fig. 2"},{"comment":"The claim that the six seams are 'the' set of decision-relevant friction points is asserted, not demonstrated. No competing-factor analysis rules out omitted variables such as task selection, sponsorship, or business-case quality, and no inter-rater reliability or factorial validity evidence is provided. If an unlisted friction explains most variance in production success, then even a correctly scored Seam Index will not predict P1. This is load-bearing: P1, P2, and P5 all presuppose that the six seams are exhaustive and dominant. The paper should either provide evidence for exhaustiveness or weaken the claim to 'a recurring set' and add an explicit validity agenda.","section":"Sec. 5, Table 2"},{"comment":"The Seam Index sums six equally weighted 0–2 scores. No justification is given for equal weighting or for the assumption that seams are roughly independent. If, for example, the data seam dominates effort or if governance and security interact, the 0–12 total will systematically misrank platforms. The worked example in Sec. 8.2 depends entirely on this additive form. Before P1 can be tested, the paper should justify the equal-weight assumption, report sensitivity to alternative weights, or present the Index as a profile rather than a single sum.","section":"Sec. 8, Table 4"}],"minor_comments":[{"comment":"The headline 95% figure and the 'roughly twice as often' buy-versus-build result both trace to a single source [28]. The paper does attribute this, which is good, but given the weight these numbers carry, a sentence noting the absence of independent corroboration for each would help calibrate readers.","section":"Sec. 3, Table 1"},{"comment":"The section title 'External Validation' overstates what lab hiring demonstrates. The labs' behavior is consistent with a deployment bottleneck, but it is also consistent with standard go-to-market strategy for a commoditizing product. Suggest retitling to 'External Evidence' or 'Consistent Behavior' to match the paper's otherwise careful epistemic hedging.","section":"Sec. 6"},{"comment":"The first paragraph and Sec. 10.2 both make the same point about agentic systems amplifying seam burden. Consider condensing to avoid repetition.","section":"Sec. 10"},{"comment":"The phrase 'and so on' after listing anchors for data, identity, and governance is insufficient for an instrument that claims reproducibility. At minimum, the supplementary material should specify anchors for all six seams, including change management and cost.","section":"Sec. 8.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is well above the typical standard for conceptual AI-management pieces in its honesty and falsifiability. The main risk is that the Seam Index, as specified, is not yet a valid instrument; the authors' own planned research agenda acknowledges this, but the current manuscript presents P1 as a testable proposition when the predictor's construct validity is still open. Major revision with a request for a validity plan or relaxed claims seems appropriate. I would not reject: the central diagnosis is grounded in external literature and the contributions are clearly scoped."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful parts. The paper names a real phenomenon — enterprise GenAI pilots failing on deployment friction rather than model capability — and supports it with a structured synthesis of external field studies. The Seam Index is genuinely new: a 0–12 rubric that scores a platform by how many of six friction seams it removes natively, sitting one level above the ML Test Score. The Deployment Wall and Deployment Debt are memorable constructs that make the problem concrete. The paper also follows good design-science practice: reproducible scoring protocol, evidence anchors, a worked example, six falsifiable propositions, and a limitations section that states the propositions are untested and the field observations non-random.\n\nThe soft spots are real and, in my view, located exactly where the stress-test note puts them. The Seam Index may not measure what it claims. Change management and, to a lesser extent, governance are adopter-side organizational capabilities, not platform attributes you can inherit. The paper itself calls change management the seam most correlated with value, but the scoring anchors in Section 8.1 don't define what a native score of 2 means for it. So the Index conflates platform architecture with organizational readiness, and P1's predictor is not well-defined until that's fixed. The equal weighting of six seams is asserted, not supported by factorial or reliability evidence. And the Deployment Wall's \"mechanical reproduction\" of ~5% survival is a calibration of hand-set attrition rates to the target outcome — Fig. 2 is a fit, not a finding. The 95% headline is single-sourced.\n\nNone of this is deceptive; the paper discloses most of it. But the central promise — making the bottleneck measurable — currently rests on assertion. The fix is the research the paper itself proposes: inter-rater reliability and a P1-style predictive study.\n\nBottom line: read it for the diagnosis, cite it for the constructs, and don't use the Seam Index for procurement until it's validated. It deserves a serious referee, one who'll push on construct validity and the calibration issue.","headline":"A candid, well-structured framework that names a real bottleneck, but the Seam Index's construct validity needs serious work before it can be used as a predictor.","tokens_in":12925,"tokens_out":2819,"would_cite":true,"duration_ms":28464,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Enterprise AI fails from deployment friction, not model intelligence; a 0–12 Seam Index makes the bottleneck measurable at purchase time.","keywords":["generative AI","enterprise AI deployment","Deployment Wall","Seam Index","Deployment Debt","technical debt","AI adoption","measurement instrument"],"falsifier":"A longitudinal study that records a platform's Seam Index and the model's benchmark rank at the moment of purchase, then follows initiatives to production: if benchmark rank predicts production success better than the Seam Index, the paper's central claim fails. Even simpler: if any unlisted friction—task selection, sponsorship, business-case quality—accounts for most outcome variance in existing pilot data, the six-seam set is incomplete as the load-bearing decomposition.","tokens_in":11915,"feed_emoji":"🧱","tokens_out":5547,"duration_ms":50876,"temperature":0.7,"pith_summary":"The paper tries to establish that the reason about 95% of enterprise generative-AI pilots fail to produce measurable profit-and-loss impact is not that models are too weak, but that organizational and architectural friction prevents capable models from reaching production. It makes this diagnosis operational with three linked constructs: the Deployment Wall (a six-stage value-leak model), the Seam Index (a 0–12 instrument scoring how many of six recurring friction seams a platform removes natively), and Deployment Debt (a compounding liability from unresolved seams). A sympathetic reader should care because the framework turns an eight-figure platform decision from a benchmark comparison into an architecture comparison, with six falsifiable propositions offered for empirical validation. If correct, the paper's claim implies that patience—waiting for smarter models—is the most expensive strategy, and that investment in deployment capability compounds.","feed_headline":"Deployment friction, not model IQ, decides AI success","feed_subtitle":"A 0-12 Seam Index scores platforms by the friction they remove, turning eight-figure bets into architecture comparisons","key_machinery":"The Deployment Wall is a six-stage survival funnel—model selection, integration, governance, workflow redesign, enterprise adoption, realized business value—through which every initiative must pass; moderate attrition at each gate mechanically yields the single-digit survival rates field studies observe. The Seam Index is the measuring instrument: six recurring friction seams, each scored 0 (adopter builds it), 1 (partial support), or 2 (inherited natively), summed to a 0–12 score. Deployment Debt renames unresolved seams as a compounding financial, organizational, and competitive liability. Together these constructs turn a platform choice into an architecture comparison: the buyer counts ho","core_discovery":"The central claim is that enterprise generative-AI investments fail because of organizational and architectural friction at six recurring boundaries—data, identity and access, security and compliance, governance, change management, and cost—not because models are insufficiently capable. The paper models this as a Deployment Wall with six cumulative stages that mechanically produces the observed low survival rate, and introduces the Seam Index, a reproducible 0–12 score rating how many seams a platform removes natively rather than leaving to the adopter. It argues this score, measured at purchase time, should predict production success better than benchmark rank, and that durable competitive","pith_inferences":["An immediate extension: apply the Seam Index retrospectively to already-completed pilots; if scores at kickoff separate later production survivors from failures, the instrument's predictive claim gains support without waiting for new deployments.","The framework implies that 'open' or best-benchmark platforms are systematically over-valued in current procurement, and that procurement teams should weight seam inheritance at least as heavily as capability—a change in practice the paper advocates but does not quantify.","If deployment debt compounds as described, the cost of delaying deployment-capability building should grow super-linearly; a longitudinal accounting of remediation costs across stalled initiatives would test this compounding function.","The agentic-AI wave offers a natural experiment: if seam burden amplifies for autonomous agents, organizations that close the six seams for assisted use first should fail less catastrophically than those that grant autonomy early—something the paper predicts but does not test."],"forward_implications":["A platform's Seam Index at selection should predict production success more strongly than its model's benchmark rank (P1).","Higher Seam Index scores should come with lower total cost of ownership over the initiative lifecycle (P2).","Remediating a seam later in the deployment sequence should cost strictly more than remediating it earlier, because deployment debt compounds (P3).","Workflow redesign, not model quality, should drive measurable value; wrapping workflows around a model should underperform (P4).","Platforms that leave seams to the adopter should require heavier external consulting, and durable advantage should sit with providers that own identity, data, and governance seams rather than the top model (P5, P6)."],"fun_headline_variants":["AI pilots fail from friction, not weak models","Seam Index: why 95% of AI pilots fail","Deployment debt, not model IQ, sinks AI projects","Six seams decide enterprise AI success, not benchmarks","Stop blaming the model: enterprise AI fails at deployment"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework assumes the six listed seams—data, identity and access, security and compliance, governance, change management, and cost—are the right, complete, and roughly independent set of places where enterprise-AI value leaks.","fun_headline_variants_meta":{"raw":{"variants":["AI pilots fail from friction, not weak models","Seam Index: why 95% of AI pilots fail","Deployment debt, not model IQ, sinks AI projects","Six seams decide enterprise AI success, not benchmarks","Stop blaming the model: enterprise AI fails at deployment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000376,"raw_usage":{"total_tokens":1850,"prompt_tokens":761,"completion_tokens":1089,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":1012}},"tokens_in":505,"tokens_out":1089,"duration_ms":7888,"temperature":1.0,"reasoning_tokens":1012,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T13:51:38.739575+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal study that records a platform's Seam Index and the model's benchmark rank at the moment of purchase, then follows initiatives to production: if benchmark rank predicts production success better than the Seam Index, the paper's central claim fails. Even simpler: if any unlisted friction—task selection, sponsorship, business-case quality—accounts for most outcome variance in existing pilot data, the six-seam set is incomplete as the load-bearing decomposition.","supporting_citations":[],"review_version":1}