{"id":"d454f1c7-33c4-4bec-a03f-10b534bf7817","arxiv_id":"2506.20759","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic map of 27 studies finds eight agile-for-ML frameworks, eight practice themes, and effort estimation as the top reported challenge.","lead":"This paper systematically maps 27 published studies on how agile project management has been adapted for machine learning systems. It identifies eight management approaches, eight practice themes, and names effort estimation as the most common challenge.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-database search with narrow vocabulary is the load-bearing risk: snowballing from only 10 Scopus seed papers cannot fully compensate, so 'comprehensive mapping' is asserted beyond the retrieval evidence.","rationale":"The reader's weakest assumption—that the single Scopus query plus snowballing may not retrieve a comprehensive set—is precisely the most load-bearing concern, because the paper's main contribution is a comprehensive map and the frequency-based findings (dominant challenge, theme prominence) are computed from that retrieval set. I agree with the conditional verdict: the mapping is useful and methodologically reasonable, but the completeness claim is stated more strongly than a single-database, fixed-vocabulary search warrants. The concrete test I propose directly measures recall by comparing against expanded vocabulary and other databases, as well as against related reviews that cover similar ground. If the test shows minimal additional papers, the concern is resolved; if it shows substantial additional papers, the paper's conclusions about the absence of certain themes or the dominance of others would need revision. I have not found a separate internal-inconsistency concern that would move the verdict to reject; the paper's own limitations section and supplementary repository are consistent with a careful but incomplete search.","tokens_in":9803,"tokens_out":4045,"duration_ms":41110,"concrete_test":"Re-run the search with an expanded vocabulary across multiple databases: replace ('management' OR 'practices') with ('management' OR 'practices' OR 'process*' OR 'methodolog*' OR 'lifecycle') and ('agile' OR 'scrum') with ('agile' OR 'scrum' OR 'iterative' OR 'kanban' OR 'sprint*'), searching Scopus, IEEE Xplore, ACM Digital Library, and Web of Science. Also check whether primary studies from Alves et al. (2023) and Nahar et al. (2023) on agile management of ML/AI projects appear in the 27. If the expanded search adds more than roughly 3–4 additional relevant primary studies, the comprehensiveness claim in Sections VI–VII and the dominant-challenge finding in RQ4 should be tempered accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a comprehensive mapping of agile management for ML-enabled systems (Abstract, Section VII). The retrieval rests on a single Scopus string—('machine learning' OR 'artificial intelligence') AND (('management' OR 'practices') AND ('agile' OR 'scrum'))—plus backward/forward snowballing from the 10 Scopus hits (Section III-B/C). Relevant work phrased as 'iterative ML development', 'kanban for data science', 'sprint-based MLOps', or 'AI project management' without those exact words will not be retrieved; snowballing only recovers such papers if they are adjacent to the 10 seeds. The paper's own Section VI admits 'the possibility we have missed studies' yet asserts confidence because the hybrid search was repeated; repetition does not fix vocabulary recall. Because the 'eight themes' and 'dominant challenge' findings are frequencies over this 27-paper corpus (e.g., Hybrid Approaches 10/27, effort estimation 10 citations), a structurally biased retrieval set directly biases the map. This is load-bearing, not fatal: the 27 studies may be correctly analyzed, but 'comprehensive' is stronger than the retrieval evidence supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a systematic mapping study (SMS) of agile management for ML-enabled systems. Following the guidelines of Petersen et al. and the hybrid search strategy of Wohlin et al., the authors searched Scopus with a PICO-based string and supplemented the results with backward and forward snowballing, finally selecting 27 primary studies published between 2008 and 2024. From this corpus, they identify eight agile management approaches, synthesize recommendations and practices into eight themes and challenges into three themes, classify research types and empirical evaluation methods, and conclude that effort estimation is the dominant open challenge. The paper also provides a Zenodo repository with supplementary material. The main contribution is an organized map of the state of the art and a set of research gaps, but the paper's 'comprehensive' claim is not fully supported by the retrieval evidence.","tokens_in":10117,"tokens_out":5618,"duration_ms":60554,"significance":"If the mapping is accepted, the paper provides a useful first structured synthesis of agile management practices for ML-enabled systems, with the practical finding that effort estimation is the most frequently reported obstacle. The study's transparency—explicit protocol, documented inclusion/exclusion criteria, and a public Zenodo supplement—is a genuine strength and supports replicability. The paper also gives a fair assessment of the weak empirical evidence base in the primary studies, which is valuable for steering future research. However, the significance is moderated by the fact that the corpus is built on a single-database, narrow-vocabulary search, and the central claim of comprehensiveness is therefore stronger than the evidence. The identified themes and frequencies are still informative as a map of the included 27 studies, but they are not yet a definitive map of the field.","major_comments":[{"comment":"The retrieval strategy is the load-bearing risk for the paper's main claim. The only database searched is Scopus, and the search string is (\"machine learning\" OR \"artificial intelligence\") AND ((\"management\" OR \"practices\") AND (\"agile\" OR \"scrum\")). This string will not retrieve studies that describe their subject as \"data science,\" \"data analytics,\" \"MLOps,\" \"iterative ML development,\" \"kanban for data science,\" or \"AI project management\" unless those papers are connected by citation to the 10 Scopus seeds. Backward and forward snowballing from only 10 seed papers cannot recover a literature network that is disconnected from those seeds. Section VI argues that repeating the hybrid search multiple times in 2024 increases confidence in comprehensiveness, but repetition does not fix vocabulary recall. Because the abstract and Section VII claim a \"comprehensive mapping,\" and because the theme frequencies (e.g., Hybrid Approaches 10/27, effort estimation 10 papers) are counts over this corpus, the comprehensiveness claim is not supported by the reported retrieval evidence. I recommend either broadening the search string to include additional terms (e.g., \"data science,\" \"data analytics,\" \"MLOps,\" \"AI project management\") or softening the claim to describe a map of the identified corpus, with a sensitivity analysis or a recall check against a set of known relevant studies.","section":"Section III-B and Section VI"},{"comment":"The text states that eight agile management approaches were identified, but Table III lists nine rows, including a row labeled \"None [15]\" that contains adaptations for ethical user stories. If the \"None\" entry is not an approach, it should not be listed among the approaches in Table III; if it is intended as a ninth approach, the count and the abstract should be corrected. This inconsistency affects the summary of RQ1 and should be resolved before publication.","section":"Section IV-A, Table III"},{"comment":"The thematic synthesis described in Section IV-C was performed by the first author and reviewed by the last author, while the reliability section claims that coding was \"independently peer reviewed\" and discrepancies resolved through consensus. No inter-rater reliability measure (e.g., Cohen's kappa) or a detailed coding audit trail is reported. Given that the paper's principal findings are frequency counts over thematic categories, the absence of any agreement metric makes it difficult to assess the robustness of the themes for RQ3 and RQ4. I recommend reporting coding agreement statistics or at least a more detailed description of how the independent review was conducted and resolved.","section":"Section IV-C and Section VI"}],"minor_comments":[{"comment":"In the search string, the second \"or\" is lowercase (\"agile\" or \"scrum\") while the first is uppercase; use consistent casing for readability.","section":"Section III-B"},{"comment":"The phrase \"there's the possibility we have missed studies\" is informal; consider \"there is the possibility\" and revise the surrounding sentence for formal style.","section":"Section VI"},{"comment":"Section III-C says the initial selection process \"was conducted by the main author and subsequently reviewed by the other authors,\" while Section VI claims \"all steps were conducted by more than one researcher.\" These statements are in tension and should be reconciled.","section":"Section III-C and Section VI"},{"comment":"Table II lists the empirical evaluation types as \"experiment, case study, survey,\" but the text in Section IV-F additionally mentions a Proof-of-Concept study and a Design Science Research study. Align the classification scheme with the reported categories.","section":"Section IV-F and Table II"},{"comment":"The paper uses \"frameworks\" in the abstract and \"approaches\" in Section IV-A for the same set of items; choose one term consistently, and fix the spacing in the Table III header \"AGILEMANAGEMENTAPPROACHES ANDTHEIRADAPTATIONS.\"","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The central concern is the overclaimed comprehensiveness of the search, which is a load-bearing issue for the paper's main contribution. The paper is otherwise in scope and the protocol is transparent. If the authors either broaden the search and re-run the analysis or explicitly reframe the contribution as a map of the identified corpus, the paper would be worth reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this one. First, it is a genuinely useful piece of secondary research: it rounds up 27 papers on agile management for ML-enabled systems, distills eight management approaches and eight practice/recommendation themes, and names effort estimation as the dominant open problem. The thematic synthesis (iteration flexibility, decoupled ceremonies, hybrid agile/CRISP-DM cycles, ML-specific artifacts) is clearly organized and will save anyone new to this area a lot of reading. The authors also did the right reproducibility things: a Petersen/Wohlin-style protocol, inclusion/exclusion criteria in a table, a Zenodo supplement, and a transparent account of the selection flow. For a subfield that is mostly solution proposals and single-case experience reports, this consolidation is a legitimate contribution, even though none of the individual frameworks are new.\n\nSecond, the soft spot is exactly where the stress-test note puts it: the retrieval. The initial database search is a single Scopus string using 'management' or 'practices' crossed with 'agile' or 'scrum'. Work phrased as 'kanban for data science', 'sprint-based MLOps', or 'AI project management' without those exact terms will slip through, and snowballing from ten Scopus seeds only recovers vocabulary-neighbors of those seeds. The paper's own threats section admits missed studies are possible but then asserts confidence because the hybrid search was repeated. Repetition does not fix recall; that is a logical gap in the mitigation. The frequencies that drive the main findings (e.g., hybrid approaches in 10/27 papers, effort estimation cited by ten studies) are frequencies over this retrieved corpus, so a biased retrieval set tilts the map.\n\nThat said, I do not think this is a load-bearing flaw in the sense that the conclusions would collapse. The 27 studies that were found are relevant, the coding looks careful (two authors, consensus, open data), and the gap analysis—weak empirical validation, dominance of case studies, need for generalizable guidelines—matches what I know of the area. The paper is a fair map of what these papers say; it is just not proof that nothing important was missed. A tempering of the word 'comprehensive' and ideally a second database (IEEE Xplore or ACM DL) or an independent screening pass would fix most of it.\n\nI would send this to serious peer review. The method is sound, the synthesis is valuable, and the limitations are stated openly, even if the confidence in completeness is overstated. A careful referee can push for the retrieval fix without rejecting the substance. I would cite it as the current best entry point into this literature, and I would bring it to a reading group for the discussion of hybrid search strategies alone.","headline":"A solid, honestly presented mapping study whose usefulness is real but whose 'comprehensive' claim overreaches its single-database, narrow-string retrieval.","tokens_in":10518,"tokens_out":692,"would_cite":true,"duration_ms":10460,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic mapping of 27 studies organizes agile management for ML-enabled systems into eight approaches and eight themes, with effort estimation as the dominant open challenge.","keywords":["agile management","machine learning","artificial intelligence","systematic mapping study","ML-enabled systems","Scrum","effort estimation","agile practices"],"falsifier":"A replication using the same inclusion criteria across several major digital libraries with expanded terms such as 'MLOps,' 'data science project management,' and 'AI project management' that recovers additional qualifying primary studies with new approaches or challenge themes would falsify the paper's claim of a comprehensive state-of-the-art mapping.","tokens_in":9588,"feed_emoji":"🤖","tokens_out":10886,"duration_ms":99803,"temperature":0.7,"pith_summary":"This paper is a systematic mapping study, a literature-review protocol that classifies studies rather than aggregating evidence, applied here to agile management of machine-learning-enabled systems. The authors set out to establish a consolidated picture of how agile methods are being adapted for ML projects and identify 27 peer-reviewed studies published between 2008 and 2024. From those studies they construct eight management approaches, eight themes of practices and recommendations, and three challenge themes, and they report that effort estimation for experimental ML tasks is the most frequently cited challenge. A reader should care because practitioners currently lack a consolidated map of these adaptations, and the taxonomy gives both researchers and teams a structured list of existing practices and remaining gaps, the largest of which is the near absence of rigorous empirical validation.","feed_headline":"27 studies map agile ML management into eight approaches","feed_subtitle":"Effort estimation for experimental ML tasks is the most reported hurdle across the mapped studies.","key_machinery":"The mechanism that carries the argument is the systematic mapping protocol. It combines a bibliographic database search built on the string ('machine learning' OR 'artificial intelligence') AND (('management' OR 'practices') AND ('agile' OR 'scrum')) with iterative backward and forward snowballing that screened over 2,400 records, followed by thematic synthesis with open coding to cluster practices and recommendations into eight themes and challenges into three themes. This protocol is what turns a list of 27 papers into an ordered taxonomy, and a failure in the search or coding steps would make the mapping's categories and frequencies unreliable.","core_discovery":"The study's central claim is that the fragmented literature on agile management for ML-enabled systems can be organized into a coherent landscape. From 27 primary studies collected through a bibliographic database search plus iterative backward and forward snowballing, the authors identify eight distinct management approaches—Agile4MLS, STAMP 4 NLP, SKI, Scrum-DS, Data Driven Scrum, Agile-facilitated Knowledge Discovery, ADS, and ASD-DM—and synthesize their recommendations into eight themes: iteration flexibility, ML-specific artifacts, decoupled ceremonies, hybrid agile/data-mining approaches, minimal viable models or demo APIs, Kanban adoption, business alignment, and ethical considerations. The reported challenges consolidate into three themes: sprint planning and effort estimation, methodological/training/strategic alignment deficiencies, and ethical considerations. The authors further claim that effort estimation is the most persistent reported challenge and that the field is dominated by solution proposals and case studies, with only one controlled experiment among the 27 papers.","pith_inferences":["Editorial extension: a multi-database replication with broader vocabulary (MLOps, data science project management, AI project management) would likely surface additional frameworks, so the eight-approach count is best treated as a lower bound rather than the true population.","Editorial extension: the mapping's theme of capability-based iterations suggests a testable hypothesis—teams using decoupled ceremonies or capability-based iterations should show fewer sprint overruns than teams on fixed time-boxes—which the single controlled experiment in the corpus is far too small to assess.","Editorial extension: because effort estimation is the dominant challenge, a concrete design target is an estimation method tied to experimental uncertainty, such as sizing work by data-versioning or experiment count rather than by feature points, which the mapped literature does not yet provide.","Editorial extension: the eight themes could be turned into a lightweight assessment checklist for practitioners, converting the mapping from a literature summary into a diagnostic tool for agile-for-ML adoption, though the paper itself stops at description."],"forward_implications":["Researchers gain a common vocabulary for positioning new work: a new framework can be described by which of the eight approaches it extends and which of the eight practice themes it addresses.","Effort estimation is identified as the priority open problem, so techniques that improve estimation for experimental ML tasks would directly attack the most frequently reported pain point.","The near absence of controlled experiments (one among 27 studies) implies that the next wave of research should test existing frameworks such as SKI and Data Driven Scrum under controlled or quasi-experimental conditions.","The dominance of hybrid agile/data-mining integrations suggests that reconciling the sequential dependencies of ML workflows with iterative delivery is the central design problem future methods must solve.","Practitioners get a shortlist of concrete adaptations—flexible iterations, decoupled ceremonies, model and data stories, demo APIs, Kanban adoption—instead of an unstructured collection of suggestions."],"supporting_citations":[{"why":"Supplies the hybrid database-search plus backward and forward snowballing strategy on which the comprehensiveness claim rests.","marker":"[41]"},{"why":"Provides the systematic mapping study guidelines governing the protocol.","marker":"[25]"},{"why":"Supplies the thematic synthesis and open-coding steps from which the eight themes and three challenge themes are derived.","marker":"[10]"},{"why":"Provides the research-type classification scheme used to categorize the 27 studies.","marker":"[40]"},{"why":"Provides the empirical-strategy classification (experiment, case study, survey) used to assess validation methods.","marker":"[42]"},{"why":"Gives the question-formulation strategy used to construct the search string.","marker":"[22]"},{"why":"Primary study contributing the Agile4MLS approach, including Model/Data Stories and Demo ML-API.","marker":"[38]"},{"why":"Primary study contributing the SKI framework with capability-based iterations and decoupled Scrum ceremonies.","marker":"[29]"},{"why":"Primary study contributing Data Driven Scrum, a capability-based iteration approach central to the mapping.","marker":"[28]"},{"why":"Primary study contributing Scrum-DS, the hybrid CRISP-DM-per-sprint approach.","marker":"[4]"}],"fun_headline_variants":["Agile ML management mapped: 27 studies, 8 approaches","Effort estimation is the top agile ML challenge","What 27 studies reveal about agile ML management","Eight agile approaches for ML development found","Agile ML: 8 frameworks, one key hurdle"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The mapping assumes that a single bibliographic database search plus backward and forward snowballing retrieves essentially all relevant studies, so if important agile-for-ML work uses different vocabulary or is indexed only in other databases, the taxonomy and theme frequencies will be incomplete.","fun_headline_variants_meta":{"raw":{"variants":["Agile ML management mapped: 27 studies, 8 approaches","Effort estimation is the top agile ML challenge","What 27 studies reveal about agile ML management","Eight agile approaches for ML development found","Agile ML: 8 frameworks, one key hurdle"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000589,"raw_usage":{"total_tokens":2765,"prompt_tokens":948,"completion_tokens":1817,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":1742}},"tokens_in":564,"tokens_out":1817,"duration_ms":15890,"temperature":1.0,"reasoning_tokens":1742,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:41:59.568006+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication using the same inclusion criteria across several major digital libraries with expanded terms such as 'MLOps,' 'data science project management,' and 'AI project management' that recovers additional qualifying primary studies with new approaches or challenge themes would falsify the paper's claim of a comprehensive state-of-the-art mapping.","supporting_citations":[{"cited_title":"Successful combination of database search and snowballing for identification of primary studies in systematic literature studies,","cited_arxiv_id":null,"evidence_quote":"Supplies the hybrid database-search plus backward and forward snowballing strategy on which the comprehensiveness claim rests."},{"cited_title":"Guidelines for conducting systematic mapping studies in software engineering: An update,","cited_arxiv_id":null,"evidence_quote":"Provides the systematic mapping study guidelines governing the protocol."},{"cited_title":"Recommended steps for thematic synthesis in software engineering,","cited_arxiv_id":null,"evidence_quote":"Supplies the thematic synthesis and open-coding steps from which the eight themes and three challenge themes are derived."},{"cited_title":"Requirements engineering paper classification and evaluation criteria: a proposal and a discussion,","cited_arxiv_id":null,"evidence_quote":"Provides the research-type classification scheme used to categorize the 27 studies."},{"cited_title":"Wohlin, P","cited_arxiv_id":null,"evidence_quote":"Provides the empirical-strategy classification (experiment, case study, survey) used to assess validation methods."},{"cited_title":"Pico: model for clinical questions,","cited_arxiv_id":null,"evidence_quote":"Gives the question-formulation strategy used to construct the search string."},{"cited_title":"Agile4mls—leveraging agile practices for developing machine learning-enabled systems: An industrial experience,","cited_arxiv_id":null,"evidence_quote":"Primary study contributing the Agile4MLS approach, including Model/Data Stories and Demo ML-API."},{"cited_title":"Ski: An agile framework for data science,","cited_arxiv_id":null,"evidence_quote":"Primary study contributing the SKI framework with capability-based iterations and decoupled Scrum ceremonies."},{"cited_title":"Achieving lean data science agility via data driven scrum,","cited_arxiv_id":null,"evidence_quote":"Primary study contributing Data Driven Scrum, a capability-based iteration approach central to the mapping."},{"cited_title":"Applying scrum in data science projects,","cited_arxiv_id":null,"evidence_quote":"Primary study contributing Scrum-DS, the hybrid CRISP-DM-per-sprint approach."}],"review_version":1}