REVIEW 3 major objections 5 minor 42 references
Agile Management for Machine Learning: A Systematic Mapping Study
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A systematic mapping of 27 studies organizes agile management for ML-enabled systems into eight approaches and eight themes, with effort estimation as the dominant open challenge.
desk verdict A solid, honestly presented mapping study whose usefulness is real but whose 'comprehensive' claim overreaches its single-database, narrow-string retrieval. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism that carries the argument is the systematic mapping protocol. It combines a bibliographic database search built on the string ('machine learning' OR 'artificial intelligence') AND (('management' OR 'practices') AND ('agile' OR 'scrum')) with iterative backward and forward snowballing that screened over 2,400 records, followed by thematic synthesis with open coding to cluster practices and recommendations into eight themes and challenges into three themes. This protocol is what turns a list of 27 papers into an ordered taxonomy, and a failure in the search or coding steps would make the mapping's categories and frequencies unreliable.
What would settle it
A replication using the same inclusion criteria across several major digital libraries with expanded terms such as 'MLOps,' 'data science project management,' and 'AI project management' that recovers additional qualifying primary studies with new approaches or challenge themes would falsify the paper's claim of a comprehensive state-of-the-art mapping.
Extended reading notes
Core claim
The study's central claim is that the fragmented literature on agile management for ML-enabled systems can be organized into a coherent landscape. From 27 primary studies collected through a bibliographic database search plus iterative backward and forward snowballing, the authors identify eight distinct management approaches—Agile4MLS, STAMP 4 NLP, SKI, Scrum-DS, Data Driven Scrum, Agile-facilitated Knowledge Discovery, ADS, and ASD-DM—and synthesize their recommendations into eight themes: iteration flexibility, ML-specific artifacts, decoupled ceremonies, hybrid agile/data-mining approaches, minimal viable models or demo APIs, Kanban adoption, business alignment, and ethical considerations. The reported challenges consolidate into three themes: sprint planning and effort estimation, methodological/training/strategic alignment deficiencies, and ethical considerations. The authors further claim that effort estimation is the most persistent reported challenge and that the field is dominated by solution proposals and case studies, with only one controlled experiment among the 27 papers.
Load-bearing premise
The mapping assumes that a single bibliographic database search plus backward and forward snowballing retrieves essentially all relevant studies, so if important agile-for-ML work uses different vocabulary or is indexed only in other databases, the taxonomy and theme frequencies will be incomplete.
Editorial extensions
If this is right
- Researchers gain a common vocabulary for positioning new work: a new framework can be described by which of the eight approaches it extends and which of the eight practice themes it addresses.
- Effort estimation is identified as the priority open problem, so techniques that improve estimation for experimental ML tasks would directly attack the most frequently reported pain point.
- The near absence of controlled experiments (one among 27 studies) implies that the next wave of research should test existing frameworks such as SKI and Data Driven Scrum under controlled or quasi-experimental conditions.
- The dominance of hybrid agile/data-mining integrations suggests that reconciling the sequential dependencies of ML workflows with iterative delivery is the central design problem future methods must solve.
- Practitioners get a shortlist of concrete adaptations—flexible iterations, decoupled ceremonies, model and data stories, demo APIs, Kanban adoption—instead of an unstructured collection of suggestions.
Reading between the lines
- Editorial extension: a multi-database replication with broader vocabulary (MLOps, data science project management, AI project management) would likely surface additional frameworks, so the eight-approach count is best treated as a lower bound rather than the true population.
- Editorial extension: the mapping's theme of capability-based iterations suggests a testable hypothesis—teams using decoupled ceremonies or capability-based iterations should show fewer sprint overruns than teams on fixed time-boxes—which the single controlled experiment in the corpus is far too small to assess.
- Editorial extension: because effort estimation is the dominant challenge, a concrete design target is an estimation method tied to experimental uncertainty, such as sizing work by data-versioning or experiment count rather than by feature points, which the mapped literature does not yet provide.
- Editorial extension: the eight themes could be turned into a lightweight assessment checklist for practitioners, converting the mapping from a literature summary into a diagnostic tool for agile-for-ML adoption, though the paper itself stops at description.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a systematic mapping study (SMS) of agile management for ML-enabled systems. Following the guidelines of Petersen et al. and the hybrid search strategy of Wohlin et al., the authors searched Scopus with a PICO-based string and supplemented the results with backward and forward snowballing, finally selecting 27 primary studies published between 2008 and 2024. From this corpus, they identify eight agile management approaches, synthesize recommendations and practices into eight themes and challenges into three themes, classify research types and empirical evaluation methods, and conclude that effort estimation is the dominant open challenge. The paper also provides a Zenodo repository with supplementary material. The main contribution is an organized map of the state of the art and a set of research gaps, but the paper's 'comprehensive' claim is not fully supported by the retrieval evidence.
Significance. If the mapping is accepted, the paper provides a useful first structured synthesis of agile management practices for ML-enabled systems, with the practical finding that effort estimation is the most frequently reported obstacle. The study's transparency—explicit protocol, documented inclusion/exclusion criteria, and a public Zenodo supplement—is a genuine strength and supports replicability. The paper also gives a fair assessment of the weak empirical evidence base in the primary studies, which is valuable for steering future research. However, the significance is moderated by the fact that the corpus is built on a single-database, narrow-vocabulary search, and the central claim of comprehensiveness is therefore stronger than the evidence. The identified themes and frequencies are still informative as a map of the included 27 studies, but they are not yet a definitive map of the field.
major comments (3)
- [Section III-B and Section VI] The retrieval strategy is the load-bearing risk for the paper's main claim. The only database searched is Scopus, and the search string is ("machine learning" OR "artificial intelligence") AND (("management" OR "practices") AND ("agile" OR "scrum")). This string will not retrieve studies that describe their subject as "data science," "data analytics," "MLOps," "iterative ML development," "kanban for data science," or "AI project management" unless those papers are connected by citation to the 10 Scopus seeds. Backward and forward snowballing from only 10 seed papers cannot recover a literature network that is disconnected from those seeds. Section VI argues that repeating the hybrid search multiple times in 2024 increases confidence in comprehensiveness, but repetition does not fix vocabulary recall. Because the abstract and Section VII claim a "comprehensive mapping," and because the theme frequencies (e.g., Hybrid Approaches 10/27, effort estimation 10 papers) are counts over this corpus, the comprehensiveness claim is not supported by the reported retrieval evidence. I recommend either broadening the search string to include additional terms (e.g., "data science," "data analytics," "MLOps," "AI project management") or softening the claim to describe a map of the identified corpus, with a sensitivity analysis or a recall check against a set of known relevant studies.
- [Section IV-A, Table III] The text states that eight agile management approaches were identified, but Table III lists nine rows, including a row labeled "None [15]" that contains adaptations for ethical user stories. If the "None" entry is not an approach, it should not be listed among the approaches in Table III; if it is intended as a ninth approach, the count and the abstract should be corrected. This inconsistency affects the summary of RQ1 and should be resolved before publication.
- [Section IV-C and Section VI] The thematic synthesis described in Section IV-C was performed by the first author and reviewed by the last author, while the reliability section claims that coding was "independently peer reviewed" and discrepancies resolved through consensus. No inter-rater reliability measure (e.g., Cohen's kappa) or a detailed coding audit trail is reported. Given that the paper's principal findings are frequency counts over thematic categories, the absence of any agreement metric makes it difficult to assess the robustness of the themes for RQ3 and RQ4. I recommend reporting coding agreement statistics or at least a more detailed description of how the independent review was conducted and resolved.
minor comments (5)
- [Section III-B] In the search string, the second "or" is lowercase ("agile" or "scrum") while the first is uppercase; use consistent casing for readability.
- [Section VI] The phrase "there's the possibility we have missed studies" is informal; consider "there is the possibility" and revise the surrounding sentence for formal style.
- [Section III-C and Section VI] Section III-C says the initial selection process "was conducted by the main author and subsequently reviewed by the other authors," while Section VI claims "all steps were conducted by more than one researcher." These statements are in tension and should be reconciled.
- [Section IV-F and Table II] Table II lists the empirical evaluation types as "experiment, case study, survey," but the text in Section IV-F additionally mentions a Proof-of-Concept study and a Design Science Research study. Align the classification scheme with the reported categories.
- [General] The paper uses "frameworks" in the abstract and "approaches" in Section IV-A for the same set of items; choose one term consistently, and fix the spacing in the Table III header "AGILEMANAGEMENTAPPROACHES ANDTHEIRADAPTATIONS."
Circularity Check
No material circularity: the mapping synthesizes 27 primary studies, and its central challenge findings rest on external primary sources.
full rationale
The paper's derivation chain is retrieval, selection, coding, and thematic synthesis, with no fitted parameters or equations that reduce a predicted quantity to an input. The eight management themes and three challenge themes are open-coded categories derived from the 27 selected primary studies; frequencies such as Hybrid Approaches (10 of 27 papers) and effort estimation are descriptive counts over that corpus, not forced outputs of the method. The authors' self-citations are not load-bearing: reference [7] supports a background claim about data scientist and software engineer interaction, and reference [41] supports the standard methodological choice of combining database search with snowballing, while the central finding that effort estimation is the dominant challenge is attributed to a list of external primary studies, e.g., [4, 9, 17, 19, 27, 29, 31, 32, 35, 36]. The paper's acknowledged risk of missing studies, stated in Section VI as 'there's the possibility we have missed studies,' is an external-validity and completeness threat, not a circularity, because the map's conclusions are not equivalent to the search string or selection criteria by construction. No imported uniqueness theorem, renamed result, or ansatz smuggled via citation is present. The minor self-citations are visible but independent of the substantive mapping claims, so the appropriate score is 1 rather than 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The Scopus search string combined with backward and forward snowballing retrieves a representative and sufficiently complete set of primary studies.
- domain assumption The inclusion and exclusion criteria (EC1-EC6), including the exclusion of grey literature, posters, and short papers, do not systematically bias the synthesized themes and challenges.
- domain assumption The thematic synthesis with open coding by two of the authors yields reliable category labels.
Cite this review
Pith. "Pith review of Agile Management for Machine Learning: A Systematic Mapping Study." pith.science (2026). https://pith.science/paper/JHVQ4CZH
@misc{pith2026250620759,
author = {Pith},
title = {Pith review of: Agile Management for Machine Learning: A Systematic Mapping Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/JHVQ4CZH}},
note = {Machine review of arXiv:2506.20759}
}
read the original abstract
[Context] Machine learning (ML)-enabled systems are present in our society, driving significant digital transformations. The dynamic nature of ML development, characterized by experimental cycles and rapid changes in data, poses challenges to traditional project management. Agile methods, with their flexibility and incremental delivery, seem well-suited to address this dynamism. However, it is unclear how to effectively apply these methods in the context of ML-enabled systems, where challenges require tailored approaches. [Goal] Our goal is to outline the state of the art in agile management for ML-enabled systems. [Method] We conducted a systematic mapping study using a hybrid search strategy that combines database searches with backward and forward snowballing iterations. [Results] Our study identified 27 papers published between 2008 and 2024. From these, we identified eight frameworks and categorized recommendations and practices into eight key themes, such as Iteration Flexibility, Innovative ML-specific Artifacts, and the Minimal Viable Model. The main challenge identified across studies was accurate effort estimation for ML-related tasks. [Conclusion] This study contributes by mapping the state of the art and identifying open gaps in the field. While relevant work exists, more robust empirical evaluation is still needed to validate these contributions.
Figures
Reference graph
Works this paper leans on
-
[15]
Utilizing user stories to bring ai ethics into practice in soft- ware engineering,
K.-K. Kemell, V . Vakkuri, and E. Halme, “Utilizing user stories to bring ai ethics into practice in soft- ware engineering,” inInternational Conference on Product-Focused Software Process Improvement, 2022, pp. 553–558
work page 2022
-
[1]
Kanban in software development: A systematic literature review,
M. O. Ahmad, J. Markkula, and M. Oivo, “Kanban in software development: A systematic literature review,” in2Euromicro Conference on Software Engineering and Advanced Applications, 2013, pp. 9–16
work page 2013
-
[2]
M. Alnoukari, Z. Alzoabi, and S. Hanna, “Applying adaptive software development (asd) agile model- ing on predictive data mining applications: Asd- dm methodology,” inInternational Symposium on Information Technology, vol. 2, 2008, pp. 1–6
work page 2008
-
[3]
Practices for managing machine learning products: A multivocal literature review,
I. Alves, L. A. Leite, P. Meirelles, F. Kon, and C. S. R. Aguiar, “Practices for managing machine learning products: A multivocal literature review,” IEEE Transactions on Engineering Management, vol. 71, pp. 7425–7455, 2023
work page 2023
-
[4]
Applying scrum in data science projects,
J. Baijens, R. Helms, and D. Iren, “Applying scrum in data science projects,” inIEEE Conference on Business Informatics (CBI), vol. 1, 2020, pp. 30– 38
work page 2020
-
[5]
Data analyt- ics project methodologies: Which one to choose?
J. Baijens, R. Helms, and R. Kusters, “Data analyt- ics project methodologies: Which one to choose?” inInternational Conference on Big Data in Man- agement, 2020, pp. 41–47
work page 2020
-
[6]
Manifesto for agile software development,
K. Beck, M. Beedle, A. Van Bennekum, A. Cock- burn, W. Cunningham, M. Fowler, J. Grenning, J. Highsmith, A. Hunt, R. Jeffrieset al., “Manifesto for agile software development,” 2001
work page 2001
-
[7]
G. Busquim, H. Villamizar, M. J. Lima, and M. Kalinowski, “On the interaction between soft- ware engineers and data scientists when build- ing machine learning-enabled systems,” inInterna- tional Conference on Software Quality, 2024, pp. 55–75
work page 2024
Show all 42 references
-
[8]
Cohn,User stories applied: For agile software development
M. Cohn,User stories applied: For agile software development. Addison-Wesley Professional, 2004
2004
-
[9]
Being agile in a data science project,
R. Cordeiro, I. Alves, S. Alves, and A. Goldman, “Being agile in a data science project,” inAgile Processes in Software Engineering and Extreme Programming – Workshops, P. Kruchten and P. Gre- gory, Eds., 2024
2024
-
[10]
Recommended steps for thematic synthesis in software engineering,
D. S. Cruzes and T. Dyba, “Recommended steps for thematic synthesis in software engineering,” inInternational Symposium on Empirical Software Engineering and Measurement, 2011, pp. 275–284
2011
-
[11]
On the ap- propriate methodologies for data science projects,
A. K. Dastgerdi and T. J. Gandomani, “On the ap- propriate methodologies for data science projects,” inInternational Conference on Information Tech- nology (ICIT), 2021, pp. 667–673
2021
-
[12]
Ai lifecycle models need to be revised: An exploratory study in fintech,
M. Haakman, L. Cruz, H. Huijgens, and A. Van Deursen, “Ai lifecycle models need to be revised: An exploratory study in fintech,” Empirical Software Engineering, vol. 26, no. 5, p. 95, 2021
2021
-
[13]
How to write ethical user stories? impacts of the eccola method,
E. Halme, V . Vakkuri, J. Kultanen, M. Jantunen, K.- K. Kemell, R. Rousi, and P. Abrahamsson, “How to write ethical user stories? impacts of the eccola method,” inInternational Conference on Agile Soft- ware Development, 2021, pp. 36–52
2021
-
[14]
The agile deployment of machine learning models in health- care,
S. Jackson, M. Yaqub, and C.-X. Li, “The agile deployment of machine learning models in health- care,”Frontiers in Big Data, vol. 1, p. 7, 2019
2019
-
[16]
Stamp 4 nlp–an agile framework for rapid quality-driven nlp applications develop- ment,
P. Kohl, O. Schmidts, L. Kl ¨oser, H. Werth, B. Kraft, and A. Z ¨undorf, “Stamp 4 nlp–an agile framework for rapid quality-driven nlp applications develop- ment,” inInternational Conference on the Quality of Information and Communications Technology, 2021, pp. 156–166
2021
-
[17]
On the application of scrum in data science projects,
N. Kraut and F. Transchel, “On the application of scrum in data science projects,” inInternational Conference on Big Data Analytics (ICBDA), 2022, pp. 1–9
2022
-
[18]
What makes agile software development agile?
M. Kuhrmann, P. Tell, R. Hebig, andet al., “What makes agile software development agile?”IEEE Transactions on Software Engineering, vol. 48, no. 9, pp. 3523–3539, 2022
2022
-
[19]
Evaluating data science project agility by exploring process frameworks used by data science teams,
S. Lahiri and J. Saltz, “Evaluating data science project agility by exploring process frameworks used by data science teams,” 2023
2023
-
[20]
Agile clinical research: A data science approach to scrumban in clinical medicine,
H. Lei, R. O’Connell, L. Ehwerhemuepha, S. Tara- man, W. Feaster, and A. Chang, “Agile clinical research: A data science approach to scrumban in clinical medicine,”Intelligence-based medicine, vol. 3, p. 100009, 2020
2020
-
[21]
An agile framework for trustworthy ai
S. Leijnen, H. Aldewereld, R. van Belkom, R. Bij- vank, and R. Ossewaarde, “An agile framework for trustworthy ai.” inNeHuAI@ ECAI, 2020, pp. 75– 78
2020
-
[22]
Pico: model for clinical questions,
R. Leonardo, “Pico: model for clinical questions,” Evid Based Med Pract, vol. 3, no. 115, p. 2, 2018
2018
-
[23]
Mitchell,Machine Learning, ser
T. Mitchell,Machine Learning, ser. McGraw- Hill International Editions. McGraw-Hill, 1997. [Online]. Available: https://books.google.com.br/ books?id=EoYBngEACAAJ
1997
-
[24]
A meta-summary of challenges in building products with ml components–collecting experiences from 4758+ practitioners,
N. Nahar, H. Zhang, G. Lewis, S. Zhou, and C. K ¨astner, “A meta-summary of challenges in building products with ml components–collecting experiences from 4758+ practitioners,” inInter- national Conference on AI Engineering–Software Engineering for AI (CAIN), 2023, pp. 171–183
2023
-
[25]
Guidelines for conducting systematic mapping studies in software engineering: An update,
K. Petersen, S. Vakkalanka, and L. Kuzniarz, “Guidelines for conducting systematic mapping studies in software engineering: An update,”Infor- mation and Software Technology, vol. 64, pp. 1–18, 2015
2015
-
[26]
An improved agile framework for implementing data science initia- tives in the government,
W. Qadadeh and S. Abdallah, “An improved agile framework for implementing data science initia- tives in the government,” inInternational Confer- ence on Information and Computer Technologies (ICICT), 2020, pp. 24–30
2020
-
[27]
Comparing data science project management methodologies via a controlled experiment,
J. Saltz, K. Crowstonet al., “Comparing data science project management methodologies via a controlled experiment,” 2017
2017
-
[28]
Achieving lean data science agility via data driven scrum,
J. Saltz, A. Sutherland, and N. Hotz, “Achieving lean data science agility via data driven scrum,” 2022
2022
-
[29]
Ski: An agile framework for data science,
J. Saltz and A. Suthrland, “Ski: An agile framework for data science,” inIEEE International Conference on Big Data (Big Data), 2019, pp. 3468–3476
2019
-
[30]
Identifying the most common frameworks data science teams use to structure and coordinate their projects,
J. S. Saltz and N. Hotz, “Identifying the most common frameworks data science teams use to structure and coordinate their projects,” inInterna- tional Conference on Big Data (Big Data), 2020, pp. 2038–2042
2020
-
[31]
Achieving agile big data science: the evolution of a team’s agile process methodology,
J. S. Saltz and I. Shamshurin, “Achieving agile big data science: the evolution of a team’s agile process methodology,” inIEEE International Conference on Big Data (Big Data), 2019, pp. 3477–3485
2019
-
[32]
Iden- tifying and addressing 6 key questions when using data driven scrum,
J. S. Saltz, A. Sutherland, and T. Jombart, “Iden- tifying and addressing 6 key questions when using data driven scrum,” inInternational Conference on Big Data (Big Data), 2021, pp. 2345–2352
2021
-
[33]
Synthesizing agile and knowledge discovery: case study results,
C. Schmidt and W. N. Sun, “Synthesizing agile and knowledge discovery: case study results,”Journal of Computer Information Systems, vol. 58, no. 2, pp. 142–150, 2018
2018
-
[34]
Scrum development process,
K. Schwaber, “Scrum development process,” in Business Object Design and Implementation: OOP- SLA’95 Workshop Proceedings, 1997, pp. 117–134
1997
-
[35]
Analysis of software engineering for agile machine learning projects,
K. Singla, J. Bose, and C. Naik, “Analysis of software engineering for agile machine learning projects,” inIEEE India Council International Con- ference (INDICON), 2018, pp. 1–5
2018
-
[36]
Machine learning and data science project management from an agile perspective: Methods and challenges,
M. P. Uysal, “Machine learning and data science project management from an agile perspective: Methods and challenges,” inContemporary chal- lenges for agile project management. IGI Global, 2022, pp. 73–88
2022
-
[37]
Toward a method engineering framework for project management and machine learning,
——, “Toward a method engineering framework for project management and machine learning,” inAnnual Computers, Software, and Applications Conference (COMPSAC), 2023, pp. 1186–1190
2023
-
[38]
Agile4mls—leveraging agile practices for developing machine learning-enabled systems: An industrial experience,
K. Vaidhyanathan, A. Chandran, H. Muccini, and R. Roy, “Agile4mls—leveraging agile practices for developing machine learning-enabled systems: An industrial experience,”IEEE Software, vol. 39, no. 6, pp. 43–50, 2022
2022
-
[39]
Managing artificial intelligence projects: Key in- sights from an ai consulting firm,
G. Vial, A.-F. Cameron, T. Giannelia, and J. Jiang, “Managing artificial intelligence projects: Key in- sights from an ai consulting firm,”Information Systems Journal, vol. 33, no. 3, pp. 669–691, 2023
2023
-
[40]
Requirements engineering paper classification and evaluation criteria: a proposal and a discussion,
R. Wieringa, N. Maiden, N. Mead, and C. Rolland, “Requirements engineering paper classification and evaluation criteria: a proposal and a discussion,” Requirements Engineering, vol. 11, pp. 102–107, 2006
2006
-
[41]
Successful combination of database search and snowballing for identification of primary studies in systematic literature studies,
C. Wohlin, M. Kalinowski, K. R. Felizardo, and E. Mendes, “Successful combination of database search and snowballing for identification of primary studies in systematic literature studies,”Information and Software Technology, vol. 147, p. 106908, 2022
2022
-
[42]
Wohlin, P
C. Wohlin, P. Runeson, M. H ¨ost, M. C. Ohlsson, B. Regnell, and A. Wessl ´en,Experimentation in Software Engineering. Springer, 2024
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.