REVIEW 4 major objections 4 minor 41 references
Enterprise AI fails from deployment friction, not model intelligence; a 0–12 Seam Index makes the bottleneck measurable at purchase time.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 13:51 UTC pith:T62H53UY
load-bearing objection A candid, well-structured framework that names a real bottleneck, but the Seam Index's construct validity needs serious work before it can be used as a predictor. the 4 major comments →
The Deployment Wall: A Diagnostic Framework and Instrument for Enterprise AI in the Deployment Era
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that enterprise generative-AI investments fail because of organizational and architectural friction at six recurring boundaries—data, identity and access, security and compliance, governance, change management, and cost—not because models are insufficiently capable. The paper models this as a Deployment Wall with six cumulative stages that mechanically produces the observed low survival rate, and introduces the Seam Index, a reproducible 0–12 score rating how many seams a platform removes natively rather than leaving to the adopter. It argues this score, measured at purchase time, should predict production success better than benchmark rank, and that durable competitive
What carries the argument
The Deployment Wall is a six-stage survival funnel—model selection, integration, governance, workflow redesign, enterprise adoption, realized business value—through which every initiative must pass; moderate attrition at each gate mechanically yields the single-digit survival rates field studies observe. The Seam Index is the measuring instrument: six recurring friction seams, each scored 0 (adopter builds it), 1 (partial support), or 2 (inherited natively), summed to a 0–12 score. Deployment Debt renames unresolved seams as a compounding financial, organizational, and competitive liability. Together these constructs turn a platform choice into an architecture comparison: the buyer counts ho
Load-bearing premise
The framework assumes the six listed seams—data, identity and access, security and compliance, governance, change management, and cost—are the right, complete, and roughly independent set of places where enterprise-AI value leaks.
What would settle it
A longitudinal study that records a platform's Seam Index and the model's benchmark rank at the moment of purchase, then follows initiatives to production: if benchmark rank predicts production success better than the Seam Index, the paper's central claim fails. Even simpler: if any unlisted friction—task selection, sponsorship, business-case quality—accounts for most outcome variance in existing pilot data, the six-seam set is incomplete as the load-bearing decomposition.
If this is right
- A platform's Seam Index at selection should predict production success more strongly than its model's benchmark rank (P1).
- Higher Seam Index scores should come with lower total cost of ownership over the initiative lifecycle (P2).
- Remediating a seam later in the deployment sequence should cost strictly more than remediating it earlier, because deployment debt compounds (P3).
- Workflow redesign, not model quality, should drive measurable value; wrapping workflows around a model should underperform (P4).
- Platforms that leave seams to the adopter should require heavier external consulting, and durable advantage should sit with providers that own identity, data, and governance seams rather than the top model (P5, P6).
Where Pith is reading between the lines
- An immediate extension: apply the Seam Index retrospectively to already-completed pilots; if scores at kickoff separate later production survivors from failures, the instrument's predictive claim gains support without waiting for new deployments.
- The framework implies that 'open' or best-benchmark platforms are systematically over-valued in current procurement, and that procurement teams should weight seam inheritance at least as heavily as capability—a change in practice the paper advocates but does not quantify.
- If deployment debt compounds as described, the cost of delaying deployment-capability building should grow super-linearly; a longitudinal accounting of remediation costs across stalled initiatives would test this compounding function.
- The agentic-AI wave offers a natural experiment: if seam burden amplifies for autonomous agents, organizations that close the six seams for assisted use first should fail less catastrophically than those that grant autonomy early—something the paper predicts but does not test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that the dominant cause of enterprise generative-AI failure is deployment friction rather than model capability. It introduces three linked constructs: the Deployment Wall (a six-stage survival funnel), the Seam Index (a 0–12 platform-scoring instrument across six 'seams'), and Deployment Debt (a compounding liability reframing). It grounds these in a synthesis of field studies and the author's own engagements, explicitly labels figures as illustrative, and proposes six falsifiable propositions. The paper is presented as a design-oriented conceptual contribution, not an empirical validation, and its Limitations section (Sec. 12) acknowledges the untested status of the framework.
Significance. If the framework held up empirically, it would be genuinely useful: it converts platform procurement from benchmark comparison to architecture comparison, and its propositions (P1–P6) are stated in falsifiable form. The paper is unusually honest, with explicit boundary conditions, attributed single-source figures, and a planned research agenda. However, the central instrument—the Seam Index—currently lacks construct-validity evidence, and the DeployWall's 'mechanical reproduction' of the 95% survival rate is a local calibration rather than an independent test. The paper's practical value therefore depends on validation work it does not yet provide.
major comments (4)
- [Sec. 8.1, Table 4] The scoring protocol promises evidence anchors for each seam, but only data, identity, and governance are exemplified; change management—explicitly called 'the seam most correlated with realized value' in Sec. 5—has no anchor defining what a native score of 2 means. Since change management is an adopter-side organizational capability, not a platform-inheritable attribute, the Seam Index conflates platform architecture with organizational readiness. Without an anchor and a rule for who owns this seam, P1's predictor is not well-defined and the instrument is not reproducible for the seam that matters most.
- [Sec. 4, Fig. 2] The 'mechanical reproduction' of the observed ~5% survival rate is achieved by hand-setting per-stage attrition rates so that the funnel lands on the target outcome. The paper acknowledges the rates are illustrative (Sec. 12), but the abstract and Sec. 4 use 'mechanically reproduces' and 'reproduces, mechanically' as if the model were independently confirmed. This is a fit-to-outcome calibration, not a reproduction. Either stage-level survival data are reported, or the wording should be changed to 'consistent with field survival rates'—otherwise the Wall's apparent precision is misleading.
- [Sec. 5, Table 2] The claim that the six seams are 'the' set of decision-relevant friction points is asserted, not demonstrated. No competing-factor analysis rules out omitted variables such as task selection, sponsorship, or business-case quality, and no inter-rater reliability or factorial validity evidence is provided. If an unlisted friction explains most variance in production success, then even a correctly scored Seam Index will not predict P1. This is load-bearing: P1, P2, and P5 all presuppose that the six seams are exhaustive and dominant. The paper should either provide evidence for exhaustiveness or weaken the claim to 'a recurring set' and add an explicit validity agenda.
- [Sec. 8, Table 4] The Seam Index sums six equally weighted 0–2 scores. No justification is given for equal weighting or for the assumption that seams are roughly independent. If, for example, the data seam dominates effort or if governance and security interact, the 0–12 total will systematically misrank platforms. The worked example in Sec. 8.2 depends entirely on this additive form. Before P1 can be tested, the paper should justify the equal-weight assumption, report sensitivity to alternative weights, or present the Index as a profile rather than a single sum.
minor comments (4)
- [Sec. 3, Table 1] The headline 95% figure and the 'roughly twice as often' buy-versus-build result both trace to a single source [28]. The paper does attribute this, which is good, but given the weight these numbers carry, a sentence noting the absence of independent corroboration for each would help calibrate readers.
- [Sec. 6] The section title 'External Validation' overstates what lab hiring demonstrates. The labs' behavior is consistent with a deployment bottleneck, but it is also consistent with standard go-to-market strategy for a commoditizing product. Suggest retitling to 'External Evidence' or 'Consistent Behavior' to match the paper's otherwise careful epistemic hedging.
- [Sec. 10] The first paragraph and Sec. 10.2 both make the same point about agentic systems amplifying seam burden. Consider condensing to avoid repetition.
- [Sec. 8.1] The phrase 'and so on' after listing anchors for data, identity, and governance is insufficient for an instrument that claims reproducibility. At minimum, the supplementary material should specify anchors for all six seams, including change management and cost.
Circularity Check
No significant circularity: external field studies carry the diagnosis, and the funnel model is explicitly illustrative.
full rationale
The paper's central claim — that deployment friction, not model capability, explains the ~95% pilot failure rate — is grounded in independent external sources (MIT Project NANDA, RAND, McKinsey, BCG, S&P Global) and in the author's disclosed participant observation, not in a self-citation chain. The Seam Index is proposed as a new diagnostic instrument with an explicit scoring protocol, and its propositions (P1–P6) are stated as falsifiable hypotheses for future validation rather than as results derived from the instrument; the index is not fitted to production outcomes. The one passage that could resemble a fit is the Deployment Wall's claim to "mechanically reproduce" observed survival rates (Sec. 4, Fig. 2), but the paper explicitly labels the figure illustrative and states that "the precise per-stage rates are illustrative," so the funnel is a structural consistency check rather than a parameterized prediction. The paper also self-reports its limitations (non-random field observations, illustrative figures, propositions offered for testing rather than validated findings). No load-bearing step in the argument reduces by construction to its own inputs, and there are no self-citations invoked as authorities. Therefore no significant circularity is present.
Axiom & Free-Parameter Ledger
free parameters (3)
- Deployment Wall per-stage survival rates =
not stated; "moderate attrition at each stage" chosen to land near ~5% cumulative survival
- Seam Index weighting (equal 0/1/2 per seam) =
equal weights across all six seams
- Deployment-debt compounding curve =
unspecified "steeply rising" shape
axioms (7)
- domain assumption Cited field studies are accurate and representative (MIT NANDA ~95% no-P&L impact; RAND; McKinsey; S&P; Menlo)
- domain assumption The six Deployment Wall stages are cumulative, ordered, and each removes a fraction of initiatives
- ad hoc to paper The six seams are the exhaustive and decision-relevant set of friction points
- domain assumption Frontier-model benchmark convergence implies low marginal competitive value of model choice for enterprise tasks
- domain assumption Frontier labs hiring deployment consultants is revealed evidence that deployment, not intelligence, is the bottleneck
- domain assumption Background theories taken as given: Rogers diffusion, TOE, absorptive capacity, technical debt, complementary assets, AI-factory operating model
- domain assumption Purchased integrated platforms reach production ~2x as often as internally built ones (cited to [28])
invented entities (4)
-
Deployment Wall
independent evidence
-
Seam Index
independent evidence
-
Deployment Debt
independent evidence
-
Seam-layer moat / orchestration layer
no independent evidence
read the original abstract
Enterprise investment in generative artificial intelligence (AI) tripled in a single year to roughly US$37 billion, yet independent field research finds that about 95% of enterprise generative-AI pilots deliver no measurable profit-and-loss impact. We argue that the dominant explanation--that models are not yet capable enough--is mistaken, and that enterprise AI has entered a Deployment Era in which advantage derives not from model intelligence but from the removal of the organizational and architectural friction that prevents a capable model from reaching production. Building on the software-engineering literature on technical debt and machine-learning deployment, and on a structured synthesis of independent field studies, we make the diagnosis operational. We introduce three linked constructs and one measurement instrument: the Deployment Wall, a six-stage value-leak model that mechanically reproduces observed survival rates; the Seam Index, a reproducible 0-12 diagnostic that scores any platform by how many of six recurring friction "seams" it removes natively rather than leaving to the adopter; and Deployment Debt, a construct that reframes unresolved friction as a compounding, quantifiable liability. We specify a scoring protocol with evidence anchors so the instrument can be applied consistently, illustrate it on a worked platform-selection example, and derive six falsifiable propositions with a research agenda for validation. The framework converts an eight-figure platform decision from a benchmark comparison into an architecture comparison.
Figures
Reference graph
Works this paper leans on
-
[1]
Harvard Business Review Press, Boston, MA, USA, 2018
Ajay Agrawal, Joshua Gans, and Avi Goldfarb.Prediction Machines: The Simple Economics of Artificial Intelligence. Harvard Business Review Press, Boston, MA, USA, 2018
2018
-
[2]
Strategic integration of generative AI in organizational settings: Applications, challenges and adoption requirements.IEEE Engineering Management Review, 2025
Mousa Al-Kfairy. Strategic integration of generative AI in organizational settings: Applications, challenges and adoption requirements.IEEE Engineering Management Review, 2025
2025
-
[3]
Software engineering for machine learning: A case study
Saleema Amershi, Andrew Begel, Christian Bird, Robert DeLine, Harald Gall, Ece Kamar, 16 Nachiappan Nagappan, Besmira Nushi, and Thomas Zimmermann. Software engineering for machine learning: A case study. In2019 IEEE/ACM 41st Int. Conf. on Software Engineering: Software Engineering in Practice (ICSE-SEIP), pages 291–300, 2019. doi: 10.1109/ICSE-SEIP. 2019.00042
arXiv 2019
-
[4]
Enterprise alliance to deploy Claude at scale
Anthropic. Enterprise alliance to deploy Claude at scale. Technical report, 2025
2025
-
[5]
Davenport, and Stella Pachidi
Hind Benbya, Thomas H. Davenport, and Stella Pachidi. Artificial intelligence in organizations: Current state and future opportunities.MIS Quarterly Executive, 19(4), 2020
2020
-
[6]
Managing artificial intelligence
Nicholas Berente, Bin Gu, Jan Recker, and Radhika Santhanam. Managing artificial intelligence. MIS Quarterly, 45(3):1433–1450, 2021
2021
-
[7]
Where’s the value in AI? Technical report, 2025
Boston Consulting Group. Where’s the value in AI? Technical report, 2025
2025
-
[8]
Eric Breck, Shanqing Cai, Eric Nielsen, Michael Salib, and D. Sculley. The ML test score: A rubric for ML production readiness and technical debt reduction. In2017 IEEE Int. Conf. on Big Data, pages 1123–1132, 2017
2017
-
[9]
Cohen and Daniel A
Wesley M. Cohen and Daniel A. Levinthal. Absorptive capacity: A new perspective on learning and innovation.Administrative Science Quarterly, 35(1):128–152, 1990
1990
-
[10]
The WyCash portfolio management system
Ward Cunningham. The WyCash portfolio management system. InProc. OOPSLA, pages 29–30, 1992
1992
-
[11]
Davenport and Rajeev Ronanki
Thomas H. Davenport and Rajeev Ronanki. Artificial intelligence for the real world.Harvard Business Review, 96(1):108–116, 2018
2018
-
[12]
State of generative AI in the enterprise
Deloitte. State of generative AI in the enterprise. Technical report, 2025
2025
-
[13]
So what if ChatGPT wrote it?
Yogesh K. Dwivedi et al. Opinion paper: “So what if ChatGPT wrote it?” multidisciplinary perspectives on opportunities, challenges and implications of generative conversational AI. International Journal of Information Management, 71:102642, 2023
2023
-
[14]
Artificial intelligence and business value: A literature review.Information Systems Frontiers, 24:1709– 1734, 2022
Ida Merete Enholm, Emmanouil Papagiannidis, Patrick Mikalef, and John Krogstie. Artificial intelligence and business value: A literature review.Information Systems Frontiers, 24:1709– 1734, 2022
2022
-
[15]
Building the AI-powered organization
Tim Fountaine, Brian McCarthy, and Tamim Saleh. Building the AI-powered organization. Harvard Business Review, 97(4):62–73, 2019
2019
-
[16]
Predicts 2026: Agentic AI in the enterprise
Gartner. Predicts 2026: Agentic AI in the enterprise. Technical report, Stamford, CT, USA, 2025
2026
-
[17]
Partner fund and consultant enablement for enterprise AI
Google Cloud. Partner fund and consultant enablement for enterprise AI. Technical report, 2025
2025
-
[18]
Lakhani.Competing in the Age of AI: Strategy and Leadership When Algorithms and Networks Run the World
Marco Iansiti and Karim R. Lakhani.Competing in the Age of AI: Strategy and Leadership When Algorithms and Networks Run the World. Harvard Business Review Press, Boston, MA, USA, 2020
2020
-
[19]
Global AI adoption and governance index
IBM. Global AI adoption and governance index. Technical report, Armonk, NY, USA, 2024
2024
-
[20]
Enterprise generative-AI data readiness
K2view. Enterprise generative-AI data readiness. Technical report, 2025. 17
2025
-
[21]
Dominik Kreuzberger, Niklas Kühl, and Sebastian Hirschl. Machine learning operations (MLOps): Overview, definition, and architecture.IEEE Access, 11:31866–31879, 2023. doi: 10.1109/ACCESS.2023.3262138
arXiv 2023
-
[22]
Enterprise GenAI security report 2025
LayerX. Enterprise GenAI security report 2025. Technical report, 2025
2025
-
[23]
Large-scale machine learning systems in real-world industrial settings: A review of challenges and solutions.Information and Software Technology, 127:106368, 2020
Lucy Ellen Lwakatare, Aiswarya Raj, Ivica Crnkovic, Jan Bosch, and Helena Holmström Olsson. Large-scale machine learning systems in real-world industrial settings: A review of challenges and solutions.Information and Software Technology, 127:106368, 2020
2020
-
[24]
The state of AI
McKinsey & Company. The state of AI. Technical report, 2025
2025
-
[25]
2025: The state of generative AI in the enterprise
Menlo Ventures. 2025: The state of generative AI in the enterprise. Technical report, San Francisco, CA, USA, 2025
2025
-
[26]
Microsoft 365 Copilot and systems-integrator enablement
Microsoft. Microsoft 365 Copilot and systems-integrator enablement. Technical report, 2025
2025
-
[27]
Patrick Mikalef and Manjul Gupta. Artificial intelligence capability: Conceptualization, measurement calibration, and empirical study on its impact on organizational creativity and firm performance.Information & Management, 58(3):103434, 2021
2021
-
[28]
The GenAI divide: State of AI in business 2025
MIT Project NANDA. The GenAI divide: State of AI in business 2025. Technical report, Massachusetts Institute of Technology, Cambridge, MA, USA, 2025
2025
-
[29]
OpenAI for enterprise and deployment services
OpenAI. OpenAI for enterprise and deployment services. Technical report, 2025
2025
-
[30]
Lawrence
Andrei Paleyes, Raoul-Gabriel Urma, and Neil D. Lawrence. Challenges in deploying machine learning: A survey of case studies.ACM Computing Surveys, 55(6):1–29, 2022. doi: 10.1145/ 3533378
2022
-
[31]
Ramp AI index: Enterprise AI spending trends
Ramp. Ramp AI index: Enterprise AI spending trends. Technical report, 2026
2026
-
[32]
The root causes of failure for artificial intelligence projects and how they can succeed
RAND Corporation. The root causes of failure for artificial intelligence projects and how they can succeed. Technical report, Santa Monica, CA, USA, 2024
2024
-
[33]
Rogers.Diffusion of Innovations
Everett M. Rogers.Diffusion of Innovations. Free Press, New York, NY, USA, 5 edition, 2003
2003
-
[34]
Everyone wants to do the model work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M. Aroyo. “Everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI. InProc. 2021 CHI Conf. on Human Factors in Computing Systems, 2021. doi: 10.1145/3411764.3445518
arXiv 2021
-
[35]
Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison
D. Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-François Crespo, and Dan Dennison. Hidden technical debt in machine learning systems. InAdvances in Neural Information Processing Systems (NeurIPS), volume 28, pages 2503–2511, 2015
2015
-
[36]
Adoption and effects of software engineering best practices in machine learning
Alex Serban, Koen van der Blom, Holger Hoos, and Joost Visser. Adoption and effects of software engineering best practices in machine learning. InProc. 14th ACM/IEEE Int. Symp. on Empirical Software Engineering and Measurement (ESEM), 2020. doi: 10.1145/3382494. 3410681
-
[37]
Generative AI adoption and enterprise outcomes
S&P Global Market Intelligence. Generative AI adoption and enterprise outcomes. Technical report, New York, NY, USA, 2025. 18
2025
-
[38]
David J. Teece. Profiting from technological innovation: Implications for integration, collabora- tion, licensing and public policy.Research Policy, 15(6):285–305, 1986
1986
-
[39]
David J. Teece. Explicating dynamic capabilities: The nature and microfoundations of (sustainable) enterprise performance.Strategic Management Journal, 28(13):1319–1350, 2007
2007
-
[40]
Lexington Books, Lexington, MA, USA, 1990
LouisG.TornatzkyandMitchellFleischer.The Processes of Technological Innovation. Lexington Books, Lexington, MA, USA, 1990
1990
-
[41]
James Wilson and Paul R
H. James Wilson and Paul R. Daugherty. Collaborative intelligence: Humans and AI are joining forces.Harvard Business Review, 96(4):114–123, 2018. 19
2018
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.