REVIEW 2 major objections 5 minor 37 references
The hardest problems shipping ML models into production are organizational, not technical: 17 anti-patterns from 66 hours of practitioner talks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 03:38 UTC pith:AWROVO62
load-bearing objection Solid large-scale RTA catalog of 17 socio-technical anti-patterns from 66h of MLOps talks; useful practitioner framing with honest limits, not a foundational breakthrough. the 2 major comments →
Socio-Technical Anti-Patterns in Building ML-Enabled Software: Insights from Leaders on the Forefront
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Manual reflexive thematic analysis of 73 talks (66 hours) from the MLOps community yields 17 socio-technical anti-patterns whose root causes are predominantly non-technical—leadership vacuum, organizational silos, and communication failures—rather than missing tools or algorithms. Tools and pipelines frequently mitigate only the symptoms; the authors locate the causes inside organization, management, hiring, and process design, and they supply corresponding recommendations that range from technical contracts to structural team redesign.
What carries the argument
The 17 anti-patterns, each framed as a symptom–cause–recommendation triple and grouped under three organizational areas (silos, communication, leadership vacuum). These triples turn unstructured practitioner talk into an actionable diagnostic taxonomy.
Load-bearing premise
That self-selected public talks from one large practitioner community, filtered first by the keyword “team” and dominated by leaders in North America and Europe, give a sufficiently complete and unbiased picture of real industrial failure modes.
What would settle it
A multi-company field study or survey that systematically measures whether the same 17 anti-patterns appear, with comparable frequency and severity, outside the MLOps community and in regions or company sizes underrepresented in the video sample.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a large-scale reflexive thematic analysis (RTA) of 73 public videos (66 hours) from the MLOps community. It induces 17 socio-technical anti-patterns that arise when productionizing ML models, grouped under organizational silos (model-to-product and data-producer-to-consumer handovers), communication failures (redundant development, management–data-science tension), and leadership vacuum (headless-chicken hiring, résumé-driven development, hype-driven product creation). Causes are predominantly non-technical (missing processes, cultural clashes, uneducated hiring, absent strategy). The authors supply practitioner recommendations (feature stores, cross-functional teams, stronger processes, education) and triangulate the catalog against prior interview studies (Nahar et al., Kim et al., Amershi et al., etc.) in Table III, confirming many findings while adding a leadership-centric perspective.
Significance. If the catalog holds, the work supplies the largest qualitative evidence base to date on socio-technical barriers to ML productionization and shifts attention from purely technical MLOps tooling toward organizational root causes. Strengths include the transparent two-phase RTA design, public interim codes, explicit source tracing (Table I), speaker demographics (Table II), and systematic confirmation of independent prior results. The leadership-heavy sample fills a documented gap left by developer-centric studies and yields actionable anti-pattern names and recommendations that practitioners can use immediately. The methodological demonstration that large practitioner-community corpora can complement structured interviews is itself a useful contribution to empirical software engineering.
major comments (2)
- [§II.B–C, §II.E] §II.B–C and Threats (§II.E): The initial keyword filter on “team” followed by a second metadata-driven pass is pragmatic, yet the manuscript never reports how many of the final 17 anti-patterns would have been missed without the second pass, nor whether any candidate themes were discarded solely because they failed the keyword gate. A short sensitivity note (or a count of themes that appeared only in the open-coding phase) would strengthen the claim that the catalog is not an artifact of the first filter.
- [Table III, §VI] Table III and §VI: The triangulation is valuable, but several “confirmed findings” are mapped at a high level of abstraction (e.g., “cultural differences” → C1). For the novel anti-patterns that the authors claim as original (AP12–AP17, Headless-Chicken-Hiring, PoC-hell), the table offers little external corroboration. Either expand the discussion of why these constructs did not surface in the earlier interview studies or explicitly mark them as community-specific hypotheses that still require independent validation.
minor comments (5)
- [Fig. 2] Figure 2 caption and surrounding text: the two-phase diagram is clear, but the numbers (82 → 37, 128 → 38) do not sum exactly to the final 73; a one-sentence reconciliation would remove any reader confusion.
- [§§III–V] Throughout §§III–V the anti-pattern labels (AP1–AP17) are introduced without a compact summary table; adding a one-page overview table that lists each AP, its primary cause(s), and the main recommendation would improve scannability for practitioners.
- [Table II] §II.D demographics: “North America (50.%)” contains a stray period; also the role percentages sum to 100 % only after rounding—state the exact counts or note rounding.
- [References] References [1] and [2] are bare URLs; convert them to proper bibliographic entries with access dates.
- [Abstract, §V.B] Occasional typographic slips remain (e.g., “contextu-alize”, “R´esum´e”, “S¨achsische”); a final proof-reading pass is warranted.
Circularity Check
No circularity: inductive RTA of external practitioner videos yields anti-patterns; triangulation is independent confirmation, not self-definition.
full rationale
The paper's central claim (17 socio-technical anti-patterns rooted in leadership vacuum, silos, and communication) is derived by reflexive thematic analysis of 66 hours of external MLOps-community talks (73 videos), not by fitting parameters to a target quantity or by defining results in terms of themselves. Induction (open coding of 37 videos) produces initial codes/themes; deduction (closed coding of further videos) refines them—standard iterative qualitative procedure, not definitional circularity. Causes and recommendations are extracted from the same corpus of practitioner statements. Table III maps findings onto independent prior interview studies (Nahar et al., Kim et al., Amershi et al., Arpteg et al., Granlund et al.) for external validation; those works have non-overlapping authors and different data sources. No self-citation is load-bearing for the catalog itself, no uniqueness theorem is imported from the authors, and no ansatz is smuggled via citation. The result is therefore self-contained against the external video corpus and prior literature; score 0 is the correct honest finding.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption Reflexive thematic analysis of unstructured practitioner talks can validly surface shared socio-technical anti-patterns without a structured interview protocol.
- domain assumption The MLOps.community video corpus (with >11k Slack members, multi-continent speakers) is representative enough of industry productionization challenges for generalization claims.
- ad hoc to paper Socio-technical challenges of interest arise around and within teams, so keyword 'team' plus later metadata filtering adequately scopes the phenomenon.
- domain assumption Prior interview-based findings (Nahar et al. 2022, Kim et al., Amershi et al., etc.) are reliable enough to serve as external triangulation.
invented entities (1)
-
Named anti-pattern catalog (AP1–AP17), including Headless-Chicken-Hiring, Hype-driven product creation / PoC-hell, and related constructs
independent evidence
read the original abstract
Although machine learning (ML)-enabled software systems seem to be a success story considering their rise in economic power, there are consistent reports from companies and practitioners struggling to bring ML models into production. Many papers have focused on specific, and purely technical aspects, such as testing and pipelines, but only few on socio-technical aspects. Driven by numerous anecdotes and reports from practitioners, our goal is to collect and analyze socio-technical challenges of productionizing ML models centered around and within teams. To this end, we conducted the largest qualitative empirical study in this area, involving the manual analysis of 66 hours of talks that have been recorded by the MLOps community. By analyzing talks from practitioners for practitioners of a community with over 11,000 members in their Slack workspace, we found 17 anti-patterns, often rooted in organizational or management problems. We further list recommendations to overcome these problems, ranging from technical solutions over guidelines to organizational restructuring. Finally, we contextu-alize our findings with previous research, confirming existing results, validating our own, and highlighting new insights.
Figures
Reference graph
Works this paper leans on
-
[1]
Why do 87% of data science projects never make it into production?
“Why do 87% of data science projects never make it into production?” https://venturebeat.com/2019/07/19/why-do-87-of-data-science-projects- never-make-it-into-production/
2019
-
[2]
Gartner identifies the top strategic technology trends for 2022,
“Gartner identifies the top strategic technology trends for 2022,” https://www.gartner.com/en/newsroom/press-releases/2021-10-18- gartner-identifies-the-top-strategic-technology-trends-for-2022
2022
-
[3]
Hidden technical debt in machine learning systems,
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V . Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, “Hidden technical debt in machine learning systems,” inAdvances in Neural Information Processing Systems, vol. 28. Curran Associates, Inc., 2015, pp. 2503–2511,
2015
-
[4]
MLOps: challenges in multi-organization setup: Experiences from two real-world cases,
T. Granlund, A. Kopponen, V . Stirbu, L. Myllyaho, and T. Mikkonen, “MLOps: challenges in multi-organization setup: Experiences from two real-world cases,” inWorkshop on AI Engineering (WAIN). IEEE, 2021, pp. 82–88
2021
-
[5]
Using antipatterns to avoid MLOps mistakes,
N. Muralidhar, S. Muthiah, P. Butler, M. Jain, Y . Yu, K. Burne, W. Li, D. Jones, P. Arunachalam, H. S. McCormick, and N. Ramakrishnan, “Using antipatterns to avoid MLOps mistakes,” 2021, arXiv:2107.00079
Pith/arXiv arXiv 2021
-
[6]
Concolic testing for deep neural networks,
Y . Sun, M. Wu, W. Ruan, X. Huang, M. Kwiatkowska, and D. Kroen- ing, “Concolic testing for deep neural networks,” inProc. Int. Conf. Automated Software Engineering (ASE). ACM, 2018, pp. 109–119
2018
-
[7]
DeepTest: Automated testing of deep-neural-network-driven autonomous cars,
Y . Tian, K. Pei, S. Jana, and B. Ray, “DeepTest: Automated testing of deep-neural-network-driven autonomous cars,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 2018, pp. 303–314
2018
-
[8]
MODE: Automated neural network model debugging via state differential analysis and input selection,
S. Ma, Y . Liu, W.-C. Lee, X. Zhang, and A. Grama, “MODE: Automated neural network model debugging via state differential analysis and input selection,” inProc. Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). ACM, 2018, p. 175–186
2018
-
[9]
Machine learning op- erations (MLOps): Overview, definition, and architecture,
D. Kreuzberger, N. K ¨uhl, and S. Hirschl, “Machine learning op- erations (MLOps): Overview, definition, and architecture,” 2022, arXiv:2205.02302
Pith/arXiv arXiv 2022
-
[10]
Is software engineering research addressing software engineering problems? (keynote),
G. C. Murphy, “Is software engineering research addressing software engineering problems? (keynote),” inProc. Int. Conf. on Software Engineering (ICSE). IEEE, 2020, pp. 4–5
2020
-
[11]
Collaboration challenges in building ml-enabled systems: Communication, documentation, en- gineering, and process,
N. Nahar, S. Zhou, G. Lewis, and C. K ¨astner, “Collaboration challenges in building ml-enabled systems: Communication, documentation, en- gineering, and process,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 2022, pp. 413–425
2022
-
[12]
Garousi, M
V . Garousi, M. Felderer, M. V . M ¨antyl¨a, and A. Rainer,Benefitting from the Grey Literature in Software Engineering Research. Springer International Publishing, 2020, pp. 385–413
2020
-
[13]
Using thematic analysis in psychology,
V . Braun and V . Clarke, “Using thematic analysis in psychology,” Qualitative Research in Psychology, vol. 3, no. 2, pp. 77–101, 2006
2006
-
[14]
Braun, V
V . Braun, V . Clarke, N. Hayfield, and G. Terry,Thematic Analysis. Springer Singapore, 2019, pp. 843–860
2019
-
[15]
Reel life vs. real life: How software developers share their daily life through vlogs,
S. Chattopadhyay, T. Zimmermann, and D. Ford, “Reel life vs. real life: How software developers share their daily life through vlogs,” in Proc. Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). ACM, 2021, pp. 404–415
2021
-
[16]
The pains and gains of microservices: A systematic grey literature review,
J. Soldani, D. A. Tamburri, and W.-J. Van Den Heuvel, “The pains and gains of microservices: A systematic grey literature review,”Journal of Systems and Software, vol. 146, pp. 215–232, 2018
2018
-
[17]
On the practitioners’ under- standing of coupling smells — a grey literature based grounded-theory study,
A. Singjai, G. Simhandl, and U. Zdun, “On the practitioners’ under- standing of coupling smells — a grey literature based grounded-theory study,”Information and Software Technology (IST), vol. 134, p. 106539, 2021
2021
-
[18]
How do committees invent?
M. E. Conway, “How do committees invent?”Datamation, vol. 14, pp. 28–31, 1968
1968
-
[19]
Splitting the organization and integrating the code: Conway’s Law revisited,
J. D. Herbsleb and R. E. Grinter, “Splitting the organization and integrating the code: Conway’s Law revisited,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 1999, pp. 85–95
1999
-
[20]
Reflecting on reflexive thematic analysis,
V . Braun and V . Clarke, “Reflecting on reflexive thematic analysis,” Qualitative research in sport, exercise and health, vol. 11, no. 4, pp. 589–597, 2019
2019
-
[21]
Examining the use of thematic analysis as a tool for informing design of new family communication technologies,
N. Brown and T. Stockman, “Examining the use of thematic analysis as a tool for informing design of new family communication technologies,” inProc. Int. BCS Human Computer Interaction Conference (BCS-HCI). BCS Learning & Development Ltd., 2013, pp. 1–6
2013
-
[22]
Can i use ta? Should i use ta? Should i not use ta? Comparing reflexive thematic analysis and other pattern- based qualitative analytic approaches,
V . Braun and V . Clarke, “Can i use ta? Should i use ta? Should i not use ta? Comparing reflexive thematic analysis and other pattern- based qualitative analytic approaches,”Counselling and Psychotherapy Research, vol. 21, no. 1, pp. 37–47, 2021
2021
-
[23]
F. P. Brooks,The mythical man-month – Essays on Software- Engineering. Addison-Wesley, 1975
1975
-
[24]
R ´esum´e-driven development: A definition and empirical characterization,
J. Fritzsch, M. Wyrich, J. Bogner, and S. Wagner, “R ´esum´e-driven development: A definition and empirical characterization,” inProc. Int. Conf. on Software Engineering (ICSE). IEEE, 2021, pp. 19–28
2021
-
[25]
Data scientists in software teams: State of the art and challenges,
M. Kim, T. Zimmermann, R. DeLine, and A. Begel, “Data scientists in software teams: State of the art and challenges,”Transactions on Software Engineering, vol. 44, no. 11, pp. 1024–1038, 2018
2018
-
[26]
Software engineering challenges of deep learning,
A. Arpteg, B. Brinne, L. Crnkovic-Friis, and J. Bosch, “Software engineering challenges of deep learning,” inProc. Euromicro Conference on Software Engineering and Advanced Applications (SEAA), 2018, pp. 50–59
2018
-
[27]
Software engineering for machine learning: A case study,
S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar, N. Nagap- pan, B. Nushi, and T. Zimmermann, “Software engineering for machine learning: A case study,” inProc. Int. Conf. on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2019, pp. 291– 300
2019
-
[28]
Social debt in software engineering: Insights from industry,
D. A. Tamburri, P. Kruchten, P. Lago, and H. v. Vliet, “Social debt in software engineering: Insights from industry,”Journal of Internet Services and Applications, vol. 6, no. 1, pp. 1–17, 2015
2015
-
[29]
Socio-technical con- gruence: a framework for assessing the impact of technical and work dependencies on software development productivity,
M. Cataldo, J. D. Herbsleb, and K. M. Carley, “Socio-technical con- gruence: a framework for assessing the impact of technical and work dependencies on software development productivity,” inProc. Second ACM-IEEE Int. symposium on Empirical Software Engineering and Measurement (ESEM), 2008, pp. 2–11. MEETUPS [M3] MLOps.community “Hierarchy of MLOps Needs”,...
2008
-
[30]
High Stakes ML: Active Failures, Latent Fac- tors
https://www.youtube.com/watch?v=MRES5IxVnME. [M5] MLOps.community “High Stakes ML: Active Failures, Latent Fac- tors”,YouTube, Apr 16, 2020. https://www.youtube.com/watch?v= 9g4deV1uNZo. [M10] MLOps.community “MLOps - The Blind Men and the Ele- phant”,YouTube, May 11, 2020. https://www.youtube.com/watch?v= RTBq7e3FhEw. [M11] MLOps.community “Machine Learn...
2020
-
[31]
Law of Diminishing Returns for Running AI Proof-of-Concepts
https://www.youtube.com/watch?v=WvwclqkEEpE. [M62] MLOps.community “Law of Diminishing Returns for Running AI Proof-of-Concepts”,YouTube, May 10, 2021. https://www.youtube.com/ watch?v=j09xbtudJgs. [M64] MLOps.community “Your Model is Not An Island: Operationalize Machine Learning at Scale with MLOps”,YouTube, May 25, 2021. https://www.youtube.com/watch?v...
2021
-
[32]
Continuous Integration for ML
https://www.youtube.com/watch?v=rc8zLY15WZU. COFFEESESSIONS [CS6] MLOps.community “Continuous Integration for ML”,YouTube, Aug 10,
-
[33]
How to Choose the Right Machine Learning Tool: A Conversation
https://www.youtube.com/watch?v=L98VxJDHXMM. [CS13] MLOps.community “How to Choose the Right Machine Learning Tool: A Conversation”,YouTube, Oct 15, 2020. https://www.youtube.com/ watch?v=mmTCGkm3ZoQ. [CS18] MLOps.community “Luigi in Production”,YouTube, Premiered Nov 9,
2020
-
[34]
Data Observability: The Next Frontier of Data En- gineering
https://www.youtube.com/watch?v=ShBod1yXUeg. [CS19] MLOps.community “Data Observability: The Next Frontier of Data En- gineering”,YouTube, Nov 23, 2020. https://www.youtube.com/watch?v= IMyI5eKQxMI. [CS20] MLOps.community “Monzo Machine Learning Case Study”,YouTube, Dec 7, 2020. https://www.youtube.com/watch?v=EyLGKmPAZLY. [CS21] MLOps.community “A Conver...
2020
-
[35]
Practical MLOps
https://www.youtube.com/watch?v=-TGp2qKz8tA. [CS27] MLOps.community “Practical MLOps”,YouTube, Jan 26, 2021. https: //www.youtube.com/watch?v=GvAyV8m8ICI. [CS29] MLOps.community “Culture and Architecture in MLOps”,YouTube, Feb 8, 2021. https://www.youtube.com/watch?v=uV676 YLP98. [CS34] MLOps.community “Machine Learning at Atlassian”,YouTube, Apr 12,
2021
-
[36]
War Stories Productionising ML
https://www.youtube.com/watch?v=MI0hqyYSO3c. [CS35] MLOps.community “War Stories Productionising ML”,YouTube, Apr 19, 2021. https://www.youtube.com/watch?v=DmL FncITII. [CS38] MLOps.community “Organisational Challenges of MLOps”,YouTube, May 7, 2021. https://www.youtube.com/watch?v=xe3-ImbPkT0. [CS39] MLOps.community “MLOps: A Leader’s Perspective”,YouTub...
2021
-
[37]
Maturing Machine Learning in Enterprise
https://www.youtube.com/watch?v=iM1tRulj8Xc. [CS43] MLOps.community “Maturing Machine Learning in Enterprise”, YouTube, Jun 15, 2021. https://www.youtube.com/watch?v= kfm3Iozxj8I. [CS44] MLOps.community “Autonomy vs. Alignment: Scaling AI Teams to Deliver Value”,YouTube, Jun 30, 2021. https://www.youtube.com/ watch?v=Gr69acrT8HE. [CS46] MLOps.community “W...
2021
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.