{"id":"fa072ee8-d950-4e93-bc7c-2b270db07b90","arxiv_id":"2506.20621","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Define-ML extends Lean Inception with Data Source Mapping, Feature-to-Data Source Mapping, and ML Mapping activities, and practitioners in two studies perceived it as useful and expressed intention to adopt it.","lead":"Define-ML adds three structured mapping activities to the Lean Inception workshop to help teams ideate machine learning products with data constraints in mind. In two industry validation studies, practitioners reported that the approach improved clarity, alignment, and their intention to adopt it.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'validated approach' claim rests on post-workshop self-reported perceptions with no baseline or objective outcome measure; the inferential gap from participant satisfaction to actual ideation effectiveness is the load-bearing weakness.","rationale":"The reader's weakest assumption identifies the same load-bearing issue: self-reported perceptions collected via TAM after facilitated workshops are treated as evidence of actual effectiveness. My stress-test sharpens this by emphasizing that no baseline, no objective outcome measure, and no longitudinal follow-up exist anywhere in the paper, so the inferential gap from perception to validation is the entire empirical foundation of the central claim. This is not an internal inconsistency; the paper is internally coherent and honestly reports its limitations. The concern is about the strength of the evidence relative to the claim. The authors do disclose the relevant threats in §VII.B, which is a credit to their transparency, but the abstract and conclusion still use the word 'validated' without the qualifiers that the methods section would warrant. Given that the paper presents a novel, clearly described framework with reproducible artifacts and honest reporting, a conditional acceptance is appropriate: the authors should either add a comparative baseline study or temper 'validated approach' to 'perceived as useful in initial evaluations.' My analysis does not change the reader's conditional verdict, so I recommend UNCHANGED.","tokens_in":9943,"tokens_out":2954,"duration_ms":37229,"concrete_test":"Conduct a comparative evaluation with a counterbalanced design: two teams ideate the same ML product brief, one using standard Lean Inception and one using Define-ML (or cross over after a washout). Have independent expert raters, blind to condition, evaluate the resulting product definitions on pre-registered criteria: data-feasibility of the ML components, alignment with business goals, clarity of ML requirements, and scope realism. Additionally, track whether teams actually adopt and implement the ideated features over a 6-month follow-up. If Define-ML does not significantly outperform Lean Inception on these objective criteria, the 'validated approach' claim fails and should be relabeled as preliminary perceptual evidence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that Define-ML is a 'validated approach' for ML product ideation (Abstract, §VIII)—depends on the assumption that participants' TAM-based self-reports after researcher-facilitated workshops are a reliable proxy for the framework actually improving ideation outcomes. This assumption is load-bearing because it is the only evidence offered for the central claim. In §VI.C the data collection is described as a TAM questionnaire, and §VII.B explicitly concedes that 'TAM and qualitative feedback basically capture perceptions' and that 'all workshops were facilitated by the researchers, which could introduce facilitator bias.' There is no baseline comparison: all participants used only Define-ML, so high ratings could reflect the general benefit of any structured workshop rather than the specific ML-focused activities. There is no objective measure of ideation quality, such as feasibility of the resulting product definition, alignment with available data, or downstream adoption. The samples are small (11 static, 9 dynamic), and in the dynamic case only 9 of 11 business participants answered, introducing possible self-selection. The unanimous 'intention to adopt' is a behavioral intention collected at the end of a workshop, not actual adoption; no longitudinal tracking is reported. Therefore, the evidence supports 'participants perceived Define-ML as useful in these two settings,' not 'Define-ML is a validated approach for ideating ML-enabled systems.' The central claim overstates the empirical support, and the overstatement is not merely cosmetic—it is the basis on which practitioners and researchers would decide to adopt the method.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Define-ML, an extension of the Lean Inception ideation workshop with three new activities (Data Source Mapping, Feature-to-Data Source Mapping, and ML Mapping) intended to address ML-specific concerns in early product ideation. The authors followed the Technology Transfer Model, first presenting the approach in a lab session, then running a static validation with 11 practitioners on a toy problem, and a dynamic validation in a three-day workshop with a multinational energy drink company, where 9 of 11 business participants completed a TAM-based questionnaire. The reported results show high perceived usefulness for the new activities and unanimous (static) or near-unanimous (dynamic) intention to adopt the framework. The paper concludes that Define-ML is a 'validated approach' for ML product ideation. The manuscript includes an open science repository and a Miro template for the boards.","tokens_in":10309,"tokens_out":2925,"duration_ms":36333,"significance":"If the results are taken at face value, Define-ML addresses a real and current gap: most ideation methods for software products do not explicitly handle data dependencies, technical feasibility, and the alignment of ML capabilities with business goals. The paper's strengths are its concrete, reproducible artifacts (open Miro template, Zenodo repository), its adherence to a systematic technology-transfer methodology with defined research questions, and its candid threats-to-validity section, which explicitly acknowledges that TAM captures perceptions and that facilitator bias is possible. The evaluation, however, rests entirely on self-reported perceptions and behavioral intentions, with no baseline comparison, no objective outcome measures, small samples, and researcher-facilitated sessions. The paper would be a useful contribution as a design study reporting practitioner perceptions of a new ideation framework; the current wording overstates the strength of the evidence.","major_comments":[{"comment":"The conclusion that Define-ML is a 'validated approach' for ML product ideation is not supported by the evidence presented. The evaluation in Sections V and VI uses TAM-based questionnaires that measure perceived usefulness, ease of use, and intention to use (Section VI.C), and the authors themselves concede in Section VII.B that 'TAM and qualitative feedback basically capture perceptions.' There is no objective measure of ideation outcome quality, such as feasibility of the produced product definition, alignment with available data, or downstream project success. I recommend either tempering the claim to 'an approach that practitioners perceived as useful' or adding evidence that links the workshop outputs to actual ideation effectiveness.","section":"Abstract; Section VIII"},{"comment":"Neither validation includes a comparison baseline. All participants used only Define-ML, so the high agreement ratings could reflect the general value of any structured, facilitated workshop rather than the specific contribution of the three ML-focused activities. To support the central claim that these activities add value over and above Lean Inception, a comparison with an unmodified Lean Inception session, or at least an explicit within-workshop evaluation of what each activity contributed beyond the baseline method, would be needed. Without such a comparison, the paper's motivation that traditional ideation methods 'lack explicit support for ML-specific considerations' (Section II.A) is plausible but not empirically backed by the reported data.","section":"Sections V and VI"},{"comment":"The characterization 'strong perceived usefulness' in the conclusion is too strong for the ML Mapping activity in the dynamic validation, where only 6 of 9 practitioners agreed, 2 partially agreed, and 1 was neutral, and where one participant reported that the artifact 'confused me most.' The static validation also had 2 of 11 participants only partially agreeing. The paper should either report this activity as only moderately supported or explain why these results still justify the overall claim of strong usefulness.","section":"Section VI.D, RQ3; Section V.D, RQ3"},{"comment":"The 'intention to adopt' measure was collected immediately after the workshop and is a behavioral intention, not actual adoption. The authors correctly note in Section VII.B that longitudinal tracking is future work, but the conclusion in Section VIII states a 'clear practitioners' intent to adopt' as if it were evidence of adoption readiness. I recommend distinguishing clearly between reported intention and evidenced adoption in the abstract and conclusion, and noting the self-selection issue: only 9 of the 11 business participants answered the questionnaire, and the non-respondents may differ systematically from respondents.","section":"Section VI.D, RQ4; Section VII.B"}],"minor_comments":[{"comment":"The phrase 'intent of adoption' should be 'intention to adopt' to match the terminology used in Section VI.C and the TAM literature.","section":"Abstract"},{"comment":"Typo: 'aproach' should be 'approach'.","section":"Section VII.A"},{"comment":"The manuscript frequently has stray spaces before periods in headings and section references (e.g., 'V . Yildirim' in Section II.B, 'F . Step 6' in Section III.F). Please fix these formatting issues in the final version.","section":"Sections V and VI"},{"comment":"The term 'corporate governance' in the Data Source Mapping board may be unclear to readers outside the specific organizational context; consider defining it explicitly or using a more self-explanatory label such as 'IT-managed data' vs. 'locally managed data'.","section":"Section IV.A"},{"comment":"Please ensure that the frequency bars are accompanied by the exact counts or percentages, since the numbers are central to the reader's ability to assess the strength of the agreement findings.","section":"Figures 5 and 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a software engineering venue and the artifacts are commendable. The main issue is the gap between the evidence (perception-based, no baseline) and the 'validated' claim in the abstract and conclusion. The authors have already written a thoughtful threats section, so the fix should be tractable: reframe the contributions as an approach with promising practitioner-reported perceptions, add an explicit comparison or a pilot study as future work, and soften the concluding overstatements. I would not go to reject, as the core design is plausible and the artifacts are reusable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: Define-ML is a new combination of Lean Inception and the Mix & Match ML toolkit, with two additional mapping activities that force teams to think about data sources and data-to-feature dependencies early. That integration is not in the cited literature, and the paper gives you enough detail to run the workshop yourself. The Miro template and Zenodo data are a nice touch; this is reproducible enough for a methods paper.\n\nWhat the paper does well: it is honest in the places that matter. The threats-to-validity section openly concedes that TAM captures perceptions and that facilitator bias is possible. The static and dynamic validations are described concretely, with participant quotes that give a real sense of what worked and what confused people. The learning curve for ML Mapping, the request for more time on Feature-to-Data Source Mapping, the one participant who found the artifact confusing—none of that is hidden. For an early-stage method proposal, this is solid qualitative groundwork.\n\nThe soft spot is the one the reader and stress-test both flagged: the jump from \"participants said they liked it and intend to use it\" to \"validated approach\" in the abstract and conclusion. There is no baseline, no objective measure of ideation quality, tiny samples (11 and 9), and the dynamic case lost two of eleven respondents. Unanimous intention to adopt is a behavioral intention measured at the end of a facilitated workshop, not adoption. The authors themselves write that the evidence is perceptual, so the overstatement is mainly in the abstract and conclusion, but it is still an overstatement. This is not a fatal flaw for a software engineering methods paper—TAM-based initial evaluation is common and acceptable when clearly labeled as preliminary—but the claims should be tempered.\n\nWho this is for: researchers working on ML requirements and ideation, and practitioners who want a structured workshop format. It is a modest but useful step beyond existing toolkits because it embeds ML awareness into a full product-ideation flow. I would not cite it as evidence that Define-ML works; I would cite it as an example of a structured ideation method with initial acceptance data.\n\nMy recommendation: send it to peer review. The method is clearly described, the evaluation is appropriately scoped if you read the caveats, and the limitations are acknowledged rather than buried. A serious referee should ask the authors to soften \"validated\" to something like \"perceived as useful in initial evaluations,\" and to add a sentence in the conclusion that longitudinal and comparative studies are still needed. That is a revision, not a rejection.","headline":"A clearly written method paper with a genuinely new integration of existing ideation tools, but the evidence is perceptual and the 'validated' wording oversells it; still worth a serious peer review.","tokens_in":10780,"tokens_out":1283,"would_cite":true,"duration_ms":16432,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Define-ML, a Lean Inception extension, overcomes traditional ideation methods' blind spots for ML systems by making data sources and ML feasibility explicit, and reports unanimous intention to adopt it among…","keywords":["Define-ML","ML-enabled systems","ideation","Lean Inception","data mapping","machine learning","product discovery","technology acceptance"],"falsifier":"A concrete test would be a controlled longitudinal comparison where teams ideate ML-enabled products using either Define-ML or the unmodified Lean Inception, with facilitation kept neutral, and success measured by objective outcomes such as the share of ideated features that reach a deployed product using available data, the rate of feature pivots due to data unavailability, or stakeholder alignment scores over time. If Define-ML teams show no fewer data-related dead ends and no better alignment than the Lean Inception baseline, the central claim would be refuted.","tokens_in":9728,"feed_emoji":"🧠","tokens_out":7830,"duration_ms":78861,"temperature":0.7,"pith_summary":"Define-ML is a workshop framework that extends the Lean Inception ideation method with three activities built around data and machine-learning feasibility: Data Source Mapping, Feature-to-Data Source Mapping, and ML Mapping. The authors' central claim is that these activities close a gap in traditional ideation methods, which lack explicit support for data dependencies and probabilistic behavior when conceiving ML-enabled systems. They validate the framework with a toy-problem session among industry practitioners and a real-world workshop on retail demand forecasting for an energy drink company. Participants reported that the activities clarified data constraints, aligned ML capabilities with business goals, and improved cross-functional collaboration, and every questionnaire respondent indicated intention to adopt the approach. The paper positions Define-ML as an openly released, validated approach for grounding ML product ideas in what data and models can actually support.","feed_headline":"Three mapping activities steer ML product ideas toward feasible builds","feed_subtitle":"Workshop activities map data sources to features and ML models, aligning business goals with what data can support.","key_machinery":"The central mechanism is the three-activity sequence, each supported by a visual board. Data Source Mapping positions data sources on axes of public versus private and corporate governance, with quality circles for high, medium, and low quality. Feature-to-Data Source Mapping links each prioritized feature to the data sources needed to build it, exposing the MVP's data dependencies. ML Mapping separates ML-intensive from non-ML features, attaches each ML-intensive feature to a business objective, then uses tokens for data types and ML capabilities to match data sources to plausible model types. These activities are carried out during a collaborative workshop that includes ML experts, who provide real-time feasibility judgments that keep ideas grounded in what is technically achievable.","core_discovery":"On the paper's own terms, the discovery is that a structured extension of Lean Inception—adding Data Source Mapping, Feature-to-Data Source Mapping, and ML Mapping—can address the distinct challenges of ideating ML-enabled systems, namely aligning business objectives with probabilistic behavior and making data dependencies explicit early. The three activities ask teams to map data sources by access and governance, connect each proposed feature to the data it needs, and match ML-intensive features to feasible model types while tying each to a business objective. Static validation with a simulated loan-approval problem and dynamic validation in a retail demand-forecasting case both yielded high perceived usefulness on the new activities and unanimous expressed intention to adopt the framework. The authors conclude that Define-ML provides a structured framework for addressing the unique challenges of ideating ML-enabled systems.","pith_inferences":["Editorial inference: the same data-to-feature mapping logic could extend beyond ML to any data-intensive product ideation, since data dependency is not unique to predictive features.","Editorial inference: because all workshops were researcher-facilitated and evaluation was self-reported, the unanimous adoption intent may overstate what would happen when teams run Define-ML on their own; a replication with neutral facilitation would test the persistence of the effect.","Editorial inference: the ML Mapping step, built on data-type and capability tokens, could be extended to generative AI and agent-based features (as the authors suggest as future work) by adding tokens for those capability classes.","Editorial inference: the artifacts Define-ML produces, such as the feature-to-data-source matrix, could feed directly into requirements engineering for ML, potentially reducing the specification gap documented in prior work."],"forward_implications":["If Define-ML works as reported, teams can surface data gaps and feasibility constraints before committing to an ML product vision, reducing later rework.","The feature-to-data matrix gives the MVP definition an explicit data-dependency input, making scope decisions more realistic.","The structured activities give business and technical stakeholders a shared vocabulary, fostering cross-functional alignment on data-driven products.","The open template release means practitioners can adopt the framework without proprietary tools, and the documented facilitation role suggests organizations should budget for ML-expert support.","Acceptance across two industry contexts (energy sector and retail/beverage) offers initial evidence of applicability beyond a single domain."],"supporting_citations":[{"why":"Documents practitioner challenges in building ML products—expectation management and data alignment—that Define-ML targets.","marker":"[1]"},{"why":"Provides the Lean Inception workshop method that Define-ML extends.","marker":"[2]"},{"why":"Surveys requirements engineering problems for ML systems, motivating the data-feasibility alignment.","marker":"[3]"},{"why":"Supplies the staged technology-transfer process used to develop and validate the approach.","marker":"[5]"},{"why":"Provides the data-type and ML-capability token inspiration for the ML Mapping activity.","marker":"[9]"},{"why":"Gives case-study reporting guidelines that structure the dynamic validation.","marker":"[17]"},{"why":"Supplies the perceived-usefulness and ease-of-use questionnaire constructs used in the acceptance evaluation.","marker":"[19]"}],"fun_headline_variants":["Three mapping activities align ML features with feasible data","Define-ML extends Lean Inception with data and ML mapping","Mapping data sources to features steers ML product ideation","Industrial case validates framework for ML product ideation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that participants' questionnaire answers about perceived usefulness, ease of use, and intention to adopt—collected right after researcher-facilitated workshops—are a reliable stand-in for the framework's true effectiveness in real product ideation.","fun_headline_variants_meta":{"raw":{"variants":["Three mapping activities align ML features with feasible data","Define-ML extends Lean Inception with data and ML mapping","Mapping data sources to features steers ML product ideation","Industrial case validates framework for ML product ideation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000184,"raw_usage":{"total_tokens":1326,"prompt_tokens":961,"completion_tokens":365,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":577,"completion_tokens_details":{"reasoning_tokens":301}},"tokens_in":577,"tokens_out":365,"duration_ms":4485,"temperature":1.0,"reasoning_tokens":301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:43:21.440979+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be a controlled longitudinal comparison where teams ideate ML-enabled products using either Define-ML or the unmodified Lean Inception, with facilitation kept neutral, and success measured by objective outcomes such as the share of ideated features that reach a deployed product using available data, the rate of feature pivots due to data unavailability, or stakeholder alignment scores over time. If Define-ML teams show no fewer data-related dead ends and no better alignment than the Lean Inception baseline, the central claim would be refuted.","supporting_citations":[{"cited_title":"A meta- summary of challenges in building products with ml components– collecting experiences from 4758+ practitioners,","cited_arxiv_id":null,"evidence_quote":"Documents practitioner challenges in building ML products—expectation management and data alignment—that Define-ML targets."},{"cited_title":"Lean inception: how to align people and build the right product,","cited_arxiv_id":null,"evidence_quote":"Provides the Lean Inception workshop method that Define-ML extends."},{"cited_title":"Status quo and problems of requirements engineering for machine learning: Results from an international survey,","cited_arxiv_id":null,"evidence_quote":"Surveys requirements engineering problems for ML systems, motivating the data-feasibility alignment."},{"cited_title":"A model for technology transfer in practice,","cited_arxiv_id":null,"evidence_quote":"Supplies the staged technology-transfer process used to develop and validate the approach."},{"cited_title":"Mix & match machine learning: An ideation toolkit to design machine learning-enabled solutions,","cited_arxiv_id":null,"evidence_quote":"Provides the data-type and ML-capability token inspiration for the ML Mapping activity."},{"cited_title":"Runeson, M","cited_arxiv_id":null,"evidence_quote":"Gives case-study reporting guidelines that structure the dynamic validation."},{"cited_title":"Perceived usefulness, perceived ease of use, and user acceptance of information technology,","cited_arxiv_id":null,"evidence_quote":"Supplies the perceived-usefulness and ease-of-use questionnaire constructs used in the acceptance evaluation."}],"review_version":1}