{"id":"02884278-4dbb-40c5-8ccd-b149b442c477","arxiv_id":"1908.04674","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Based on four interviews, the paper argues that requirements engineering for ML needs new requirement types, including explainability, freedom from discrimination, data requirements, and an understanding of ML performance measures.","lead":"The authors interviewed four data scientists about how they handle requirements for machine learning systems, and found that requirements engineering must adapt to a world where programs are trained, not coded. It is a compact, early map of the open problems that will matter to anyone building ML systems or the teams and processes around them.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Conclusion overreaches its evidence: four data-scientist interviews do not establish what requirements engineers are 'demanded' to do.","rationale":"The reader's weakest assumption was the generalizability of the four interviews. My concern is related but distinct: even if the four data scientists were representative of data scientists generally, the study would still not directly support the conclusion about what requirements engineers should do, because the participants were not requirements engineers and the prescriptive step is the authors' interpretation (Section VI, Section I). This is a real soft spot, but it does not overturn the reader's ACCEPT verdict. The paper transparently positions itself as a first step, acknowledges validity threats in Section III-D, provides an interview guide and direct quotes, and its high-level message aligns with the cited literature (e.g., Ishikawa and Yoshioka; Horkoff). As an exploratory research agenda contribution, the paper stands; the authors should simply temper the 'demands' wording or explicitly label the three conclusions as hypotheses. Thus the verdict remains UNCHANGED, but with the recommendation to soften the abstract's prescriptive language until the RE-practitioner perspective is studied.","tokens_in":11093,"tokens_out":6729,"duration_ms":74042,"concrete_test":"Have two independent coders re-code the four interview transcripts and map each of the three abstract conclusions to explicit participant statements, requiring at least two distinct supporting statements per conclusion. Then run a small confirmatory interview or survey with requirements engineers working on ML systems, asking whether they experience these as demands. If the mapping fails or the requirements engineers disagree, the conclusions should be reframed as hypotheses to be tested rather than as demands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that ML development 'demands' requirements engineers to perform three specific activities—rests on an inferential leap. The study interviewed four data scientists (Section III-B), not requirements engineers, and the authors explicitly defer the RE perspective to future work (Section VI: 'augment it with the view of requirements engineers... Do they agree that RE for ML is different?'). The interviews show that these data scientists perceive challenges (unclear metrics, explainability, legal constraints, data quality), but the prescriptive step—that these challenges must be addressed by requirements engineers—is the authors' synthesis, stated as early as Section I ('From our perspective, this falls into the profession of a requirements engineer'). With n=4 and no reported saturation check or independent confirmation, the three abstract conclusions are plausible hypotheses for a research agenda, not empirically established demands. This is the weakest load-bearing point: if the conclusions are read as strong prescriptions rather than exploratory proposals, the evidence does not support them. The paper remains useful, but the strength of the wording is not matched by the empirical basis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an exploratory interview study with four data scientists, aiming to characterize what is unique about Requirements Engineering (RE) for machine learning (ML) systems. The authors describe their study design, present interview findings organized around quantitative targets, explainability, freedom from discrimination, legal requirements, and data requirements, and then map these onto RE activities (elicitation, analysis, specification, verification and validation). The paper concludes that ML development \"demands\" requirements engineers to understand ML performance measures, to be aware of new quality requirements, and to integrate ML specifics into the RE process. The paper positions this as a first step toward an RE methodology for ML systems and explicitly defers the perspective of requirements engineers themselves to future work.","tokens_in":11272,"tokens_out":2904,"duration_ms":32146,"significance":"If the conclusions are read as well-founded hypotheses, this is a useful early contribution to an important and rapidly growing area. The paper is transparent in its methodology: the interview guide is published, quotes are provided, and the authors discuss several threats to validity. It also offers a compact, actionable taxonomy of challenges (Table I) that can inform both practitioners and follow-up research. The strength of the contribution is, however, limited by the evidence base: four interviewees, no requirements engineers among them, no saturation check, and coding validated only by mutual cross-checking between the two authors. The prescriptive framing in the abstract and conclusions goes beyond what the data can establish. As exploratory qualitative research, the paper is credible; as an empirical demonstration that RE for ML must change in the claimed ways, it is not.","major_comments":[{"comment":"The central conclusion—\"development of ML systems demands requirements engineers to...\"—is stated as an established finding, but the evidence consists of perceptions from four data scientists. The step from \"data scientists report challenges with metrics, explainability, and legal constraints\" to \"requirements engineers must perform these three activities\" is the authors' analytical synthesis, not an empirical result of the interviews. The authors themselves acknowledge in Section VI that the requirements-engineering perspective is future work. This mismatch between evidence and prescriptive wording is load-bearing because the abstract and introduction invite readers to treat these as validated demands. I recommend reframing the conclusions as hypotheses or propositions, e.g., \"our analysis suggests that RE for ML will require...\", and adding an explicit external-validity limitation to Section III-D.","section":"Abstract, Section I, and Section IV/V"},{"comment":"The subject selection and validation method limit the strength of any generalizable claim. Only four data scientists were interviewed (two in research, two in industry, with one domain each for the industry participants), and no requirements engineers were included. The authors do not report a saturation check or any formal inter-rater agreement measure; the validation consisted of each author reviewing the other's coding and meeting to synthesize themes. This is acceptable for an exploratory study, but it cannot support strong prescriptive claims. The threats-to-validity section discusses descriptive validity, interpretation validity, researcher bias, and reactivity, but does not address generalizability or transferability. Please add a discussion of how the findings might transfer to other settings and what would be needed to confirm the conclusions.","section":"Section III-B and III-D"},{"comment":"The paper equates quantitative ML performance targets with functional requirements, using \"Quantitative Targets a.k.a. Functional Requirements\" as a heading and quoting P3 as saying \"I consider predictive power as functional requirement.\" In conventional RE terminology, predictive power is typically a quality or performance characteristic, not a functional requirement. This conflation could confuse the RE audience that the paper addresses. The later discussion in Section V-C correctly differentiates expected versus desired performance, which suggests the authors are aware of the distinction. Please either justify the terminology or replace it with language such as \"functional requirements expressed quantitatively\" and clarify how this relates to the RE literature.","section":"Section IV-A and Table I"}],"minor_comments":[{"comment":"Grammar: \"ML engineering constitute a paradigm shift\" should be \"ML engineering constitutes a paradigm shift.\"","section":"Section I"},{"comment":"The sentence \"changes in the development paradigm... also demands changes in RE\" has a subject-verb agreement error; it should be \"also demand changes.\"","section":"Abstract"},{"comment":"\"We then combined data from all transcripts in a meeting to ensure that we do cover the full data\" mixes tenses; \"ensure that we covered\" would be clearer.","section":"Section III-C"},{"comment":"The claim \"training data needs specified and validated requirements like code\" is an interesting synthesis, but it is presented without direct interviewee support; the preceding P3 quote states the opposite for data requirements (\"You could try, but it won't help\"). Please clarify how the authors derive this normative conclusion from the data.","section":"Section IV-E"}],"recommendation":"major_revision","confidential_remarks":"The paper reads like a workshop contribution that is being submitted to a journal. The core issue is calibration between evidence and claims. The authors are honest about the exploratory nature in the body of the paper, but the abstract and conclusion overstate the findings. With a careful rewording and an explicit external-validity limitation, the paper would be a solid research-in-progress or vision paper. For a full archival journal paper, I would expect a larger interview base or a mixed-method design including the requirements-engineer perspective that the authors themselves identify as missing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a small qualitative study worth knowing about, and it probably should be published as an exploratory contribution—but read the abstract as a proposal, not a finding.\n\nWhat's new: the authors interviewed four data scientists about how they handle requirements-related activities for ML systems, and they synthesize three concrete implications for RE: understanding performance measures, new quality requirements (explainability, nondiscrimination, legal), and process integration. The individual items are not new—Horkoff and Ishikawa & Yoshioka already say similar things, and the authors cite them. The new bit is the empirical grounding: direct quotes from practitioners, a described coding process, an interview guide on figshare, and explicit threats to validity. That is genuine, if modest, evidence.\n\nWhat it does well: it is honest. The authors say this is a first step, and they explicitly defer the requirements-engineer perspective to future work. They acknowledge the small sample in the validity section. No quantitative fitting, no hidden parameters, no invented entities. The coding procedure is collaborative and described in enough detail to follow. The citation pattern is also fine: the self-citations are on topic and not self-promotional.\n\nSoft spots: the main one is the distance between what the interviews support and what the abstract claims. Four data scientists saying 'we struggle with metrics, explainability, and legal issues' supports a claim that these are relevant challenges. It does not establish that requirements engineers are 'demanded' to do three specific things. That prescriptive step is the authors' synthesis, and it is plausible—but it is a hypothesis, not a measurement. The stress-test note is right about this overreach; it is not fatal, but the paper should soften the wording in the abstract and conclusion. A second, lighter issue is that there is no saturation check and no independent auditor; for an exploratory study with n=4, that is more a limitation than a defect, and the authors mostly frame it that way.\n\nWho it's for: researchers in requirements engineering and software engineering for ML who want a compact, citable baseline for the 'RE for ML is different' argument. Practitioners may find the table of RE activities useful as a checklist.\n\nVerdict: this deserves a serious referee. I would accept it with minor revisions asking them to temper 'demands' to something like 'suggests that requirements engineers may need to' and to label the conclusions as research hypotheses. That does not diminish the paper; it makes it more accurate.\n\nRecommendation: send it out. If I were the editor, I would treat it as a lightweight empirical/position paper and not hold it to a standard that a n=4 qualitative study cannot meet.","headline":"A transparent n=4 interview study that is worth reading as an exploratory research agenda, but the abstract's 'demands' wording outruns the evidence.","tokens_in":11768,"tokens_out":1986,"would_cite":true,"duration_ms":21114,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that building software by training models rather than writing code changes what requirements engineering must do, and it maps the needed changes from interviews with four data scientists.","keywords":["machine learning","requirements engineering","data scientists","interview study","explainability","freedom from discrimination","data requirements","ML performance measures"],"falsifier":"A broad survey or interview series covering, say, fifty ML projects that finds most data scientists see conventional RE practices as sufficient, or that teams using only classical requirements methods deliver ML systems just as successfully, would directly contradict the paper's central claim. A less severe check: if data scientists in other domains do not recognize the new quality requirements (explainability, non-discrimination, data requirements) as part of their work, the proposed RE methodology loses its empirical basis.","tokens_in":10901,"feed_emoji":"🤖","tokens_out":4031,"duration_ms":39721,"temperature":0.7,"pith_summary":"The paper argues that building software by training models rather than writing code changes what requirements engineering must do. Based on four interviews with data scientists, it concludes that requirements engineers need to understand ML performance measures such as accuracy, precision, recall, and lift in order to state good functional requirements, and that they must handle new quality requirements such as explainability, freedom from discrimination, and legal constraints. It also claims that training data itself needs explicit requirements, and that verification and validation of ML systems continue during operation through monitoring and retraining. The contribution is an early map of how RE activities—elicitation, analysis, specification, and verification and validation—need to adjust for ML-based systems.","feed_headline":"ML systems need a new kind of requirements engineering","feed_subtitle":"Four data scientists say specs must cover accuracy, recall, explainability, and training-data constraints.","key_machinery":"The carrying mechanism is the Software 2.0 paradigm shift: instead of manually coding rules, developers generate behavior by fitting a model to training data with a fitness function. The paper uses this shift to explain why ML systems introduce requirements types that conventional RE lacks. The second mechanism is the standard RE activity framework—elicitation, analysis, specification, and verification and validation—used as the grid for organizing interview findings; this lets the authors turn scattered practitioner statements into concrete changes to each RE activity. Named techniques such as performance measures and data lineage carry the detailed argument.","core_discovery":"On the paper's own terms, the central discovery is that the shift from coding to training forces requirements engineering to change, not just adapt its notation. Functional requirements for ML systems become quantitative performance targets: the interviewees treat predictive power measured by accuracy, precision, recall, or lift as a functional requirement, and emphasize that customers often do not understand these measures. New quality requirements appear that are absent from standard quality models, including explainability (both of the model and of single predictions) and freedom from discrimination, meaning only societally and legally accepted logics of discrimination may be used. Training data becomes a first-class object of specification, covering data quantity (diversity rather than raw count), data quality (completeness, consistency, correctness), provenance, and legal constraints such as GDPR consent. Finally, the paper claims that verification and validation do not end at deployment: the requirements engineer should specify when retraining happens, how runtime data is monitored, and what counts as a data anomaly.","pith_inferences":["If the paper is right, RE training and hiring criteria should include data-science literacy; a testable extension would be comparing project outcomes between teams with and without ML-aware RE support.","The framework suggests a concrete checklist for ML requirements—single-number evaluation metric, protected attributes, explanation situations, data source list, retraining triggers—that could be validated on industrial ML projects.","The paper's distinction between trained and continuously learning systems implies that learning systems may also need requirements on the update mechanism itself, not just on model behavior."],"forward_implications":["Requirements engineers will need enough statistical literacy to translate stakeholder goals into the right ML performance measure and to explain what precision, recall, or lift mean in a given domain.","Requirements specifications for ML systems will include new quality requirements—explainability, freedom from discrimination, and legal and regulatory constraints—alongside functional and data requirements.","Data requirements become a distinct class, specifying data quantity as diversity, data quality dimensions, collection conditions, and provenance.","Verification and validation extends into operations: monitoring runtime data, detecting anomalies, and scheduling retraining become requirements-level concerns.","Elicitation must bring data scientists and legal experts into requirements activities from the start, and identify protected characteristics before model training."],"supporting_citations":[{"why":"Supplies the survey evidence that RE is perceived as the most difficult activity for ML-based systems, motivating the study.","marker":"[2]"},{"why":"Identifies the lack of a unified treatment of non-functional requirements for ML and suggests asking ML experts, which the paper does.","marker":"[13]"},{"why":"Provides CRISP-DM, a reference process whose business-understanding step is left underspecified, which the paper extends with RE-specific detail.","marker":"[12]"},{"why":"Contributes the analogy that ML training is like compilation and data needs testing like code, supporting the paper's data-requirements argument.","marker":"[29]"},{"why":"Supplies the standard quality model that lacks explainability and freedom from discrimination, which the paper proposes to extend.","marker":"[26]"},{"why":"Recommends a single-number evaluation metric for ML projects, supporting the paper's treatment of performance measures as functional requirements.","marker":"[23]"}],"fun_headline_variants":["Data scientists reveal new rules for requirements engineering in ML","ML shifts from coding to training, so requirements must change","Requirements specs for ML must cover accuracy, explainability, and data","Experts: ML demands new quality requirements like explainability and fairness","For ML systems, requirements engineering is no longer about code alone"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study's conclusions rest on the assumption that four interviewed data scientists' experiences and opinions represent the wider population of ML practitioners; if those four are atypical, the proposed changes to RE may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["Data scientists reveal new rules for requirements engineering in ML","ML shifts from coding to training, so requirements must change","Requirements specs for ML must cover accuracy, explainability, and data","Experts: ML demands new quality requirements like explainability and fairness","For ML systems, requirements engineering is no longer about code alone"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1276,"prompt_tokens":867,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":325}},"tokens_in":483,"tokens_out":409,"duration_ms":4460,"temperature":1.0,"reasoning_tokens":325,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:34:30.563993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A broad survey or interview series covering, say, fifty ML projects that finds most data scientists see conventional RE practices as sufficient, or that teams using only classical requirements methods deliver ML systems just as successfully, would directly contradict the paper's central claim. A less severe check: if data scientists in other domains do not recognize the new quality requirements (explainability, non-discrimination, data requirements) as part of their work, the proposed RE methodology loses its empirical basis.","supporting_citations":[{"cited_title":"How do engineers perceive d ifﬁculties in engineering of machine-learning systems? - Questionnai re survey,","cited_arxiv_id":null,"evidence_quote":"Supplies the survey evidence that RE is perceived as the most difficult activity for ML-based systems, motivating the study."},{"cited_title":"Non-functional requirements for machine learning: Chal- lenges and new directions,","cited_arxiv_id":null,"evidence_quote":"Identifies the lack of a unified treatment of non-functional requirements for ML and suggests asking ML experts, which the paper does."},{"cited_title":"The CRISP-DM model: the new blueprint for d ata mining,","cited_arxiv_id":null,"evidence_quote":"Provides CRISP-DM, a reference process whose business-understanding step is left underspecified, which the paper extends with RE-specific detail."},{"cited_title":"T he ML test score: A rubric for ML production readiness and technic al debt reduction,","cited_arxiv_id":null,"evidence_quote":"Contributes the analogy that ML training is like compilation and data needs testing like code, supporting the paper's data-requirements argument."},{"cited_title":"Systems and software engineering – Systems a nd software quality requirements and evaluation (SQuaRE) – System and s oftware quality models,","cited_arxiv_id":null,"evidence_quote":"Supplies the standard quality model that lacks explainability and freedom from discrimination, which the paper proposes to extend."},{"cited_title":"Ng, Machine Learning Yearning","cited_arxiv_id":null,"evidence_quote":"Recommends a single-number evaluation metric for ML projects, supporting the paper's treatment of performance measures as functional requirements."}],"review_version":1}