Pith. sign in

REVIEW 2 major objections 5 minor 37 references

The hardest problems shipping ML models into production are organizational, not technical: 17 anti-patterns from 66 hours of practitioner talks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 03:38 UTC pith:AWROVO62

load-bearing objection Solid large-scale RTA catalog of 17 socio-technical anti-patterns from 66h of MLOps talks; useful practitioner framing with honest limits, not a foundational breakthrough. the 2 major comments →

arxiv 2607.03270 v1 pith:AWROVO62 submitted 2026-07-03 cs.SE

Socio-Technical Anti-Patterns in Building ML-Enabled Software: Insights from Leaders on the Forefront

classification cs.SE
keywords MLOpssocio-technical anti-patternsML productionizationorganizational silosqualitative empirical studyreflexive thematic analysisteam collaboration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Companies keep failing to move machine-learning models from notebooks into live products even though the tools and pipelines exist. This paper argues the main reasons are socio-technical, not purely technical. By manually analyzing 66 hours of talks from the large MLOps practitioner community, the authors surface 17 recurring anti-patterns that cluster around three organizational failures: leadership vacuum, silos between teams, and broken communication. They show that tools such as feature stores and model registries often only paper over symptoms, while the real causes sit in hiring practices, missing processes, unclear ownership of data, and management that does not understand how ML work actually progresses. The paper also extracts concrete recommendations ranging from cross-functional teams and pair programming to better product-roadmapping and education of leaders, and it places the new findings against earlier studies to confirm, extend, and re-frame them. A reader cares because the list turns scattered war stories into a usable diagnostic checklist for managers and engineers who keep watching models die on the way to production.

Core claim

Manual reflexive thematic analysis of 73 talks (66 hours) from the MLOps community yields 17 socio-technical anti-patterns whose root causes are predominantly non-technical—leadership vacuum, organizational silos, and communication failures—rather than missing tools or algorithms. Tools and pipelines frequently mitigate only the symptoms; the authors locate the causes inside organization, management, hiring, and process design, and they supply corresponding recommendations that range from technical contracts to structural team redesign.

What carries the argument

The 17 anti-patterns, each framed as a symptom–cause–recommendation triple and grouped under three organizational areas (silos, communication, leadership vacuum). These triples turn unstructured practitioner talk into an actionable diagnostic taxonomy.

Load-bearing premise

That self-selected public talks from one large practitioner community, filtered first by the keyword “team” and dominated by leaders in North America and Europe, give a sufficiently complete and unbiased picture of real industrial failure modes.

What would settle it

A multi-company field study or survey that systematically measures whether the same 17 anti-patterns appear, with comparable frequency and severity, outside the MLOps community and in regions or company sizes underrepresented in the video sample.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports a large-scale reflexive thematic analysis (RTA) of 73 public videos (66 hours) from the MLOps community. It induces 17 socio-technical anti-patterns that arise when productionizing ML models, grouped under organizational silos (model-to-product and data-producer-to-consumer handovers), communication failures (redundant development, management–data-science tension), and leadership vacuum (headless-chicken hiring, résumé-driven development, hype-driven product creation). Causes are predominantly non-technical (missing processes, cultural clashes, uneducated hiring, absent strategy). The authors supply practitioner recommendations (feature stores, cross-functional teams, stronger processes, education) and triangulate the catalog against prior interview studies (Nahar et al., Kim et al., Amershi et al., etc.) in Table III, confirming many findings while adding a leadership-centric perspective.

Significance. If the catalog holds, the work supplies the largest qualitative evidence base to date on socio-technical barriers to ML productionization and shifts attention from purely technical MLOps tooling toward organizational root causes. Strengths include the transparent two-phase RTA design, public interim codes, explicit source tracing (Table I), speaker demographics (Table II), and systematic confirmation of independent prior results. The leadership-heavy sample fills a documented gap left by developer-centric studies and yields actionable anti-pattern names and recommendations that practitioners can use immediately. The methodological demonstration that large practitioner-community corpora can complement structured interviews is itself a useful contribution to empirical software engineering.

major comments (2)
  1. [§II.B–C, §II.E] §II.B–C and Threats (§II.E): The initial keyword filter on “team” followed by a second metadata-driven pass is pragmatic, yet the manuscript never reports how many of the final 17 anti-patterns would have been missed without the second pass, nor whether any candidate themes were discarded solely because they failed the keyword gate. A short sensitivity note (or a count of themes that appeared only in the open-coding phase) would strengthen the claim that the catalog is not an artifact of the first filter.
  2. [Table III, §VI] Table III and §VI: The triangulation is valuable, but several “confirmed findings” are mapped at a high level of abstraction (e.g., “cultural differences” → C1). For the novel anti-patterns that the authors claim as original (AP12–AP17, Headless-Chicken-Hiring, PoC-hell), the table offers little external corroboration. Either expand the discussion of why these constructs did not surface in the earlier interview studies or explicitly mark them as community-specific hypotheses that still require independent validation.
minor comments (5)
  1. [Fig. 2] Figure 2 caption and surrounding text: the two-phase diagram is clear, but the numbers (82 → 37, 128 → 38) do not sum exactly to the final 73; a one-sentence reconciliation would remove any reader confusion.
  2. [§§III–V] Throughout §§III–V the anti-pattern labels (AP1–AP17) are introduced without a compact summary table; adding a one-page overview table that lists each AP, its primary cause(s), and the main recommendation would improve scannability for practitioners.
  3. [Table II] §II.D demographics: “North America (50.%)” contains a stray period; also the role percentages sum to 100 % only after rounding—state the exact counts or note rounding.
  4. [References] References [1] and [2] are bare URLs; convert them to proper bibliographic entries with access dates.
  5. [Abstract, §V.B] Occasional typographic slips remain (e.g., “contextu-alize”, “R´esum´e”, “S¨achsische”); a final proof-reading pass is warranted.

Circularity Check

0 steps flagged

No circularity: inductive RTA of external practitioner videos yields anti-patterns; triangulation is independent confirmation, not self-definition.

full rationale

The paper's central claim (17 socio-technical anti-patterns rooted in leadership vacuum, silos, and communication) is derived by reflexive thematic analysis of 66 hours of external MLOps-community talks (73 videos), not by fitting parameters to a target quantity or by defining results in terms of themselves. Induction (open coding of 37 videos) produces initial codes/themes; deduction (closed coding of further videos) refines them—standard iterative qualitative procedure, not definitional circularity. Causes and recommendations are extracted from the same corpus of practitioner statements. Table III maps findings onto independent prior interview studies (Nahar et al., Kim et al., Amershi et al., Arpteg et al., Granlund et al.) for external validation; those works have non-overlapping authors and different data sources. No self-citation is load-bearing for the catalog itself, no uniqueness theorem is imported from the authors, and no ansatz is smuggled via citation. The result is therefore self-contained against the external video corpus and prior literature; score 0 is the correct honest finding.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 1 invented entities

This is qualitative empirical SE work, not a formal derivation. Load-bearing commitments are methodological and sampling assumptions, not free parameters or new physical entities. The anti-pattern names are analytical constructs induced from data; a few (e.g., résumé-driven development) import prior definitions.

axioms (4)
  • domain assumption Reflexive thematic analysis of unstructured practitioner talks can validly surface shared socio-technical anti-patterns without a structured interview protocol.
    Core method choice (§II); RTA is justified by flexibility for conversation data, but validity of late theme development is assumed rather than externally validated.
  • domain assumption The MLOps.community video corpus (with >11k Slack members, multi-continent speakers) is representative enough of industry productionization challenges for generalization claims.
    Stated in §II.A and defended under threats (§II.E); single-community and public-performance bias remain.
  • ad hoc to paper Socio-technical challenges of interest arise around and within teams, so keyword 'team' plus later metadata filtering adequately scopes the phenomenon.
    Explicit filtering rationale in induction phase (§II.B); ethics/legal topics deliberately out of scope.
  • domain assumption Prior interview-based findings (Nahar et al. 2022, Kim et al., Amershi et al., etc.) are reliable enough to serve as external triangulation.
    Table III and §VI treat overlap as confirmation of generality.
invented entities (1)
  • Named anti-pattern catalog (AP1–AP17), including Headless-Chicken-Hiring, Hype-driven product creation / PoC-hell, and related constructs independent evidence
    purpose: Organize recurring practitioner-reported failures into actionable diagnostic units with causes and recommendations
    Analytical constructs induced from the corpus; some names are paper-coined, others adapt prior terms (e.g., résumé-driven development). Independent evidence is the public video corpus and partial overlap with prior studies, not a separate measurement instrument.

pith-pipeline@v1.1.0-grok45 · 26208 in / 2879 out tokens · 35458 ms · 2026-07-12T03:38:10.289639+00:00 · methodology

0 comments
read the original abstract

Although machine learning (ML)-enabled software systems seem to be a success story considering their rise in economic power, there are consistent reports from companies and practitioners struggling to bring ML models into production. Many papers have focused on specific, and purely technical aspects, such as testing and pipelines, but only few on socio-technical aspects. Driven by numerous anecdotes and reports from practitioners, our goal is to collect and analyze socio-technical challenges of productionizing ML models centered around and within teams. To this end, we conducted the largest qualitative empirical study in this area, involving the manual analysis of 66 hours of talks that have been recorded by the MLOps community. By analyzing talks from practitioners for practitioners of a community with over 11,000 members in their Slack workspace, we found 17 anti-patterns, often rooted in organizational or management problems. We further list recommendations to overcome these problems, ranging from technical solutions over guidelines to organizational restructuring. Finally, we contextu-alize our findings with previous research, confirming existing results, validating our own, and highlighting new insights.

Figures

Figures reproduced from arXiv: 2607.03270 by Alina Mailach, Norbert Siegmund.

Figure 1
Figure 1. Figure 1: Anecdotal evidence of non-technical issues of failed ML projects. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Two-phased qualitative research design. We apply RTA by having an inductive (bottom-up) phase and a deductive (top-down) phase as displayed in [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

37 extracted references · 2 linked inside Pith

  1. [1]

    Why do 87% of data science projects never make it into production?

    “Why do 87% of data science projects never make it into production?” https://venturebeat.com/2019/07/19/why-do-87-of-data-science-projects- never-make-it-into-production/

  2. [2]

    Gartner identifies the top strategic technology trends for 2022,

    “Gartner identifies the top strategic technology trends for 2022,” https://www.gartner.com/en/newsroom/press-releases/2021-10-18- gartner-identifies-the-top-strategic-technology-trends-for-2022

  3. [3]

    Hidden technical debt in machine learning systems,

    D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner, V . Chaudhary, M. Young, J.-F. Crespo, and D. Dennison, “Hidden technical debt in machine learning systems,” inAdvances in Neural Information Processing Systems, vol. 28. Curran Associates, Inc., 2015, pp. 2503–2511,

  4. [4]

    MLOps: challenges in multi-organization setup: Experiences from two real-world cases,

    T. Granlund, A. Kopponen, V . Stirbu, L. Myllyaho, and T. Mikkonen, “MLOps: challenges in multi-organization setup: Experiences from two real-world cases,” inWorkshop on AI Engineering (WAIN). IEEE, 2021, pp. 82–88

  5. [5]

    Using antipatterns to avoid MLOps mistakes,

    N. Muralidhar, S. Muthiah, P. Butler, M. Jain, Y . Yu, K. Burne, W. Li, D. Jones, P. Arunachalam, H. S. McCormick, and N. Ramakrishnan, “Using antipatterns to avoid MLOps mistakes,” 2021, arXiv:2107.00079

  6. [6]

    Concolic testing for deep neural networks,

    Y . Sun, M. Wu, W. Ruan, X. Huang, M. Kwiatkowska, and D. Kroen- ing, “Concolic testing for deep neural networks,” inProc. Int. Conf. Automated Software Engineering (ASE). ACM, 2018, pp. 109–119

  7. [7]

    DeepTest: Automated testing of deep-neural-network-driven autonomous cars,

    Y . Tian, K. Pei, S. Jana, and B. Ray, “DeepTest: Automated testing of deep-neural-network-driven autonomous cars,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 2018, pp. 303–314

  8. [8]

    MODE: Automated neural network model debugging via state differential analysis and input selection,

    S. Ma, Y . Liu, W.-C. Lee, X. Zhang, and A. Grama, “MODE: Automated neural network model debugging via state differential analysis and input selection,” inProc. Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). ACM, 2018, p. 175–186

  9. [9]

    Machine learning op- erations (MLOps): Overview, definition, and architecture,

    D. Kreuzberger, N. K ¨uhl, and S. Hirschl, “Machine learning op- erations (MLOps): Overview, definition, and architecture,” 2022, arXiv:2205.02302

  10. [10]

    Is software engineering research addressing software engineering problems? (keynote),

    G. C. Murphy, “Is software engineering research addressing software engineering problems? (keynote),” inProc. Int. Conf. on Software Engineering (ICSE). IEEE, 2020, pp. 4–5

  11. [11]

    Collaboration challenges in building ml-enabled systems: Communication, documentation, en- gineering, and process,

    N. Nahar, S. Zhou, G. Lewis, and C. K ¨astner, “Collaboration challenges in building ml-enabled systems: Communication, documentation, en- gineering, and process,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 2022, pp. 413–425

  12. [12]

    Garousi, M

    V . Garousi, M. Felderer, M. V . M ¨antyl¨a, and A. Rainer,Benefitting from the Grey Literature in Software Engineering Research. Springer International Publishing, 2020, pp. 385–413

  13. [13]

    Using thematic analysis in psychology,

    V . Braun and V . Clarke, “Using thematic analysis in psychology,” Qualitative Research in Psychology, vol. 3, no. 2, pp. 77–101, 2006

  14. [14]

    Braun, V

    V . Braun, V . Clarke, N. Hayfield, and G. Terry,Thematic Analysis. Springer Singapore, 2019, pp. 843–860

  15. [15]

    Reel life vs. real life: How software developers share their daily life through vlogs,

    S. Chattopadhyay, T. Zimmermann, and D. Ford, “Reel life vs. real life: How software developers share their daily life through vlogs,” in Proc. Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). ACM, 2021, pp. 404–415

  16. [16]

    The pains and gains of microservices: A systematic grey literature review,

    J. Soldani, D. A. Tamburri, and W.-J. Van Den Heuvel, “The pains and gains of microservices: A systematic grey literature review,”Journal of Systems and Software, vol. 146, pp. 215–232, 2018

  17. [17]

    On the practitioners’ under- standing of coupling smells — a grey literature based grounded-theory study,

    A. Singjai, G. Simhandl, and U. Zdun, “On the practitioners’ under- standing of coupling smells — a grey literature based grounded-theory study,”Information and Software Technology (IST), vol. 134, p. 106539, 2021

  18. [18]

    How do committees invent?

    M. E. Conway, “How do committees invent?”Datamation, vol. 14, pp. 28–31, 1968

  19. [19]

    Splitting the organization and integrating the code: Conway’s Law revisited,

    J. D. Herbsleb and R. E. Grinter, “Splitting the organization and integrating the code: Conway’s Law revisited,” inProc. Int. Conf. on Software Engineering (ICSE). ACM, 1999, pp. 85–95

  20. [20]

    Reflecting on reflexive thematic analysis,

    V . Braun and V . Clarke, “Reflecting on reflexive thematic analysis,” Qualitative research in sport, exercise and health, vol. 11, no. 4, pp. 589–597, 2019

  21. [21]

    Examining the use of thematic analysis as a tool for informing design of new family communication technologies,

    N. Brown and T. Stockman, “Examining the use of thematic analysis as a tool for informing design of new family communication technologies,” inProc. Int. BCS Human Computer Interaction Conference (BCS-HCI). BCS Learning & Development Ltd., 2013, pp. 1–6

  22. [22]

    Can i use ta? Should i use ta? Should i not use ta? Comparing reflexive thematic analysis and other pattern- based qualitative analytic approaches,

    V . Braun and V . Clarke, “Can i use ta? Should i use ta? Should i not use ta? Comparing reflexive thematic analysis and other pattern- based qualitative analytic approaches,”Counselling and Psychotherapy Research, vol. 21, no. 1, pp. 37–47, 2021

  23. [23]

    F. P. Brooks,The mythical man-month – Essays on Software- Engineering. Addison-Wesley, 1975

  24. [24]

    R ´esum´e-driven development: A definition and empirical characterization,

    J. Fritzsch, M. Wyrich, J. Bogner, and S. Wagner, “R ´esum´e-driven development: A definition and empirical characterization,” inProc. Int. Conf. on Software Engineering (ICSE). IEEE, 2021, pp. 19–28

  25. [25]

    Data scientists in software teams: State of the art and challenges,

    M. Kim, T. Zimmermann, R. DeLine, and A. Begel, “Data scientists in software teams: State of the art and challenges,”Transactions on Software Engineering, vol. 44, no. 11, pp. 1024–1038, 2018

  26. [26]

    Software engineering challenges of deep learning,

    A. Arpteg, B. Brinne, L. Crnkovic-Friis, and J. Bosch, “Software engineering challenges of deep learning,” inProc. Euromicro Conference on Software Engineering and Advanced Applications (SEAA), 2018, pp. 50–59

  27. [27]

    Software engineering for machine learning: A case study,

    S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar, N. Nagap- pan, B. Nushi, and T. Zimmermann, “Software engineering for machine learning: A case study,” inProc. Int. Conf. on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2019, pp. 291– 300

  28. [28]

    Social debt in software engineering: Insights from industry,

    D. A. Tamburri, P. Kruchten, P. Lago, and H. v. Vliet, “Social debt in software engineering: Insights from industry,”Journal of Internet Services and Applications, vol. 6, no. 1, pp. 1–17, 2015

  29. [29]

    Socio-technical con- gruence: a framework for assessing the impact of technical and work dependencies on software development productivity,

    M. Cataldo, J. D. Herbsleb, and K. M. Carley, “Socio-technical con- gruence: a framework for assessing the impact of technical and work dependencies on software development productivity,” inProc. Second ACM-IEEE Int. symposium on Empirical Software Engineering and Measurement (ESEM), 2008, pp. 2–11. MEETUPS [M3] MLOps.community “Hierarchy of MLOps Needs”,...

  30. [30]

    High Stakes ML: Active Failures, Latent Fac- tors

    https://www.youtube.com/watch?v=MRES5IxVnME. [M5] MLOps.community “High Stakes ML: Active Failures, Latent Fac- tors”,YouTube, Apr 16, 2020. https://www.youtube.com/watch?v= 9g4deV1uNZo. [M10] MLOps.community “MLOps - The Blind Men and the Ele- phant”,YouTube, May 11, 2020. https://www.youtube.com/watch?v= RTBq7e3FhEw. [M11] MLOps.community “Machine Learn...

  31. [31]

    Law of Diminishing Returns for Running AI Proof-of-Concepts

    https://www.youtube.com/watch?v=WvwclqkEEpE. [M62] MLOps.community “Law of Diminishing Returns for Running AI Proof-of-Concepts”,YouTube, May 10, 2021. https://www.youtube.com/ watch?v=j09xbtudJgs. [M64] MLOps.community “Your Model is Not An Island: Operationalize Machine Learning at Scale with MLOps”,YouTube, May 25, 2021. https://www.youtube.com/watch?v...

  32. [32]

    Continuous Integration for ML

    https://www.youtube.com/watch?v=rc8zLY15WZU. COFFEESESSIONS [CS6] MLOps.community “Continuous Integration for ML”,YouTube, Aug 10,

  33. [33]

    How to Choose the Right Machine Learning Tool: A Conversation

    https://www.youtube.com/watch?v=L98VxJDHXMM. [CS13] MLOps.community “How to Choose the Right Machine Learning Tool: A Conversation”,YouTube, Oct 15, 2020. https://www.youtube.com/ watch?v=mmTCGkm3ZoQ. [CS18] MLOps.community “Luigi in Production”,YouTube, Premiered Nov 9,

  34. [34]

    Data Observability: The Next Frontier of Data En- gineering

    https://www.youtube.com/watch?v=ShBod1yXUeg. [CS19] MLOps.community “Data Observability: The Next Frontier of Data En- gineering”,YouTube, Nov 23, 2020. https://www.youtube.com/watch?v= IMyI5eKQxMI. [CS20] MLOps.community “Monzo Machine Learning Case Study”,YouTube, Dec 7, 2020. https://www.youtube.com/watch?v=EyLGKmPAZLY. [CS21] MLOps.community “A Conver...

  35. [35]

    Practical MLOps

    https://www.youtube.com/watch?v=-TGp2qKz8tA. [CS27] MLOps.community “Practical MLOps”,YouTube, Jan 26, 2021. https: //www.youtube.com/watch?v=GvAyV8m8ICI. [CS29] MLOps.community “Culture and Architecture in MLOps”,YouTube, Feb 8, 2021. https://www.youtube.com/watch?v=uV676 YLP98. [CS34] MLOps.community “Machine Learning at Atlassian”,YouTube, Apr 12,

  36. [36]

    War Stories Productionising ML

    https://www.youtube.com/watch?v=MI0hqyYSO3c. [CS35] MLOps.community “War Stories Productionising ML”,YouTube, Apr 19, 2021. https://www.youtube.com/watch?v=DmL FncITII. [CS38] MLOps.community “Organisational Challenges of MLOps”,YouTube, May 7, 2021. https://www.youtube.com/watch?v=xe3-ImbPkT0. [CS39] MLOps.community “MLOps: A Leader’s Perspective”,YouTub...

  37. [37]

    Maturing Machine Learning in Enterprise

    https://www.youtube.com/watch?v=iM1tRulj8Xc. [CS43] MLOps.community “Maturing Machine Learning in Enterprise”, YouTube, Jun 15, 2021. https://www.youtube.com/watch?v= kfm3Iozxj8I. [CS44] MLOps.community “Autonomy vs. Alignment: Scaling AI Teams to Deliver Value”,YouTube, Jun 30, 2021. https://www.youtube.com/ watch?v=Gr69acrT8HE. [CS46] MLOps.community “W...