{"id":"585f2deb-fc2d-47e9-844c-beb626306e18","arxiv_id":"2507.02819","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Data scientists construct prediction targets through bricolage, applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining) to balance five criteria: validity, simplicity, predictability, portability, and resource requirements.","lead":"Researchers interviewed fifteen data scientists in education and healthcare about how they pick what to predict. They found that data scientists cobble together target variables from available data, balancing validity, simplicity, predictive performance, portability, and resource costs, rather than following a top-down measurement plan. The findings give HCI and machine learning researchers a map of where to build tools for better measurement choices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim rests solely on retrospective self-reports and a hypothetical vignette; without observational or artifact-based evidence, the five criteria and five strategies may describe narrative accounts rather than actual practice.","rationale":"I agree with the reader's verdict of CONDITIONAL. The most load-bearing assumption is that retrospective and hypothetical self-reports are reliable evidence of real practice. This is the hinge on which the entire empirical contribution turns: if data scientists' accounts are systematically rationalized or incomplete, the five-criteria/five-strategies framework may not describe actual target variable construction. The paper's methods are otherwise sound for a qualitative study: open coding with reconciliation, reflexive thematic analysis, and saturation claimed after 15 interviews. It also credits prior ethnographic work (Passi and Barocas, 2019) that independently observed iterative reformulation, which provides some convergent support. However, that prior work observed only one team and did not produce the specific five-by-five taxonomy. The missing limitation statement in Section 5.5 is notable: the authors acknowledge a single stakeholder perspective and limited domains, but never the risk that retrospective accounts diverge from practice. A concrete test that would settle this is an observational or artifact-based trace study; short of that, making transcripts and codebook available for independent audit would at least test the reliability of the coding. Because the reader already assigned CONDITIONAL, my stress-test does not change the verdict.","tokens_in":33094,"tokens_out":5205,"duration_ms":64852,"concrete_test":"Run an observational or artifact-based study of 3-5 real data science projects: collect git histories, notebook versions, and Slack/email threads around target variable decisions, and have independent coders map documented changes to the five strategies (piggybacking, composing, swapping, bridging, refining) and the five criteria. If a substantive fraction of target variable decisions (e.g., more than 30%) cannot be mapped, or if coders' mapping reliability is low (e.g., Cohen's kappa below 0.5), the self-report basis is suspect. Alternatively, re-interview a subset of participants with their own project artifacts as prompts; if their recounted sequence of formulation changes contradicts the artifacts, the retrospective accounts are unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that data scientists 'construct target variables through a bricolage process' (Abstract; Section 4). All supporting evidence comes from directed storytelling about past projects (Section 3.1.1) and a vignette-based hypothetical task (Section 3.1.2). Neither source is direct evidence of practice. Retrospective accounts are subject to hindsight bias, post-hoc rationalization, and narrative smoothing; a vignette elicits reasoning under hypothetical conditions, not observed behavior. The paper does not triangulate with project artifacts, logs, or direct observation, and Section 5.5 (Study Limitations) does not acknowledge this self-report reliability threat. The internal consistency of the analysis is not the issue: the coding process and quoted examples are plausible. The issue is that the empirical foundation cannot distinguish 'what data scientists do' from 'what data scientists say they do.' Because the framework's credibility depends on the assumption that verbal accounts faithfully reflect practice, an unverified divergence would invalidate the central claim about actual target variable construction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a semi-structured interview study with fifteen data scientists from education and healthcare. Using directed storytelling about past projects and a hypothetical vignette-based formulation task, the authors develop a process model in which target variable construction is a form of bricolage: data scientists iterate among candidate outcome definitions, balancing validity against predictive performance, simplicity, portability, and resource requirements, and applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining). The paper also describes theory- and data-driven validity evaluation practices and proposes design implications for tool support and pedagogy.","tokens_in":33192,"tokens_out":3342,"duration_ms":41102,"significance":"The topic is important and understudied: target variable construction is a pivotal but often invisible step in predictive modeling, and the paper gives it a rich, grounded vocabulary. The five-criteria/five-strategies framework is plausible and supported by concrete participant quotes, and the authors connect their findings to measurement theory and to prior work on problem formulation in a useful way. The study is also transparent about its qualitative methods, including positionality and limitations. If the findings hold, they shift the design conversation from enforcing top-down measurement to scaffolding resource-constrained, iterative practice. The main weakness is evidential: the central claim about how data scientists actually construct target variables rests on retrospective self-reports and a single hypothetical task, with no triangulation from artifacts or observation, and the paper does not qualify this claim sufficiently.","major_comments":[{"comment":"The central claim that data scientists 'construct target variables through a bricolage process' (Abstract; §4) rests entirely on directed storytelling about past projects and a hypothetical vignette. Section 5.5 lists several limitations but does not acknowledge the self-report reliability threat: retrospective accounts are subject to hindsight bias and narrative smoothing, and vignette reasoning may differ from behavior under real resource constraints. Because the paper offers no observational or artifact-based triangulation, it cannot distinguish what data scientists do from what they say they do. This is a load-bearing issue for the empirical contribution; it can be fixed by reframing the findings as participants' reported and hypothetical reasoning and by adding an explicit limitation.","section":"§3.1.1, §3.1.2, §5.5"},{"comment":"The vignette stimulus was refined across successive interviews ('we refined the specificity of the scenario description and accompanying evaluation plots'), and one education-domain participant (P1) was shown the healthcare vignette. The paper does not address whether the changing stimulus affected the comparability of the reasoning elicited, nor how the mismatched-domain data were handled beyond 'taken into consideration during analysis.' This is relevant because the vignette is presented as a way to 'observe their reasoning in a new modeling context.' The authors should treat the vignette data as a supplementary source, analyze and report any version-related differences, and justify the inclusion of the mismatched-domain case.","section":"§3.1.2, Appendix A"},{"comment":"The paper generalizes from fifteen network-recruited participants in two domains to 'data scientists' in the abstract and throughout §4, and its process model (Fig. 2) is presented as the target variable construction process. Saturation is asserted rather than demonstrated, and the sample is small and domain-specific. The findings should be hedged as a provisional framework based on the studied population, with explicit statements that the criteria and strategies are likely to be extended or refined in other domains, organizational contexts, and experience levels.","section":"§3.2, §3.3, Fig. 2"}],"minor_comments":[{"comment":"The abstract calls one criterion 'predictability,' whereas the findings and Figure 3 use 'predictive performance'; please unify the terminology.","section":"Abstract, §4.1, Fig. 3"},{"comment":"The paper reports that two authors independently performed open coding, but it does not report inter-coder agreement or provide a codebook excerpt; adding a table of themes with definitions and representative quotes would increase transparency and allow readers to assess the taxonomy's grounding.","section":"§3.3"},{"comment":"The sentence beginning 'Mussgnug [97] argue that...' has a subject-verb agreement error; the reference is to a single author and should read 'argues.'","section":"§2.5"},{"comment":"Several screenshots in the arXiv version are barely legible or not described in detail; please ensure that all evaluation plots are readable and, if possible, accompanied by a description of what participants saw.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"This is a solid qualitative study that fits the CSCW scope and addresses a real gap. My main concern is epistemic: the manuscript makes categorical claims about actual practice from self-report data alone. That is fixable through careful reframing and a sharper limitations section, so I recommend major revision rather than rejection. The self-citation pattern is acceptable for this focused line of work, though the authors should ensure the related-work discussion gives due weight to adjacent empirical studies of data science practice beyond their own prior papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper: it is the most explicit process model of target-variable construction I have seen, and the five-criteria/five-strategy taxonomy is genuinely usable. It builds on Passi and Barocas and Muller et al. rather than discovering a new continent, but it turns 'problem formulation is messy' into a named structure: criteria (validity, simplicity, predictive performance, portability, resource requirements) and strategies (piggybacking, composing, swapping, bridging, refining). That synthesis is the contribution, and it is a solid one. The design implications in Section 5 are concrete enough to guide tool builders without overpromising.\n\nThe method is appropriate for the question. Semi-structured interviews with directed storytelling and a vignette probe are standard for mapping practitioner reasoning. The authors describe their open coding, saturation, and positionality honestly. The quotes are illustrative and well-chosen. They also report a project that was discontinued when reformulation failed, which suggests they were not just cherry-picking success stories. The limitations section is explicit about domain scope and the collaborative nature of the work.\n\nNow the soft spots, in proportion. The central process model rests entirely on retrospective self-reports plus one hypothetical vignette. There is no observation, no artifact analysis, no logs. That means we are learning how data scientists narrate their target-variable construction, not necessarily how they do it. I do not think this sinks the paper; it is a first qualitative map, and the claim could be re-read as 'how data scientists describe their practice,' which the evidence supports. But Section 5.5 does not directly acknowledge this self-report threat, and the stress-test note is right that it should. The vignette also changed across interviews, and one participant saw the wrong-domain vignette; the authors disclose both, but they do weaken comparability. The sample is fifteen network-recruited people in two high-stakes domains, so generalizability beyond education and healthcare is an open question. Finally, the team adopted bricolage as the central lens before analysis, so there is some risk of the frame shaping interpretation; the richness of the taxonomy makes me think this was not forced, but it is worth flagging in review.\n\nNet: a well-executed qualitative study that earns its place. This deserves peer review, and I would engage with it. If I were handling it, I would ask the authors to add one paragraph on the self-report limitation and to be slightly more careful in the abstract about generalizing beyond the two domains. But none of this changes the verdict.","headline":"A careful interview study that gives the field a usable taxonomy of target-variable bricolage; the self-report evidence is the main soft spot, but not disqualifying.","tokens_in":33785,"tokens_out":1541,"would_cite":true,"duration_ms":21606,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Data scientists construct prediction targets through bricolage, making do with the data at hand rather than following a measurement plan.","keywords":["target variable construction","bricolage","problem formulation","measurement validity","predictive modeling","data science practice","interview study","model evaluation"],"falsifier":"A longitudinal trace of real data science projects, from version-controlled labels or screen recordings, that showed target variables are fixed before any data exploration and never swapped, composed, or refined in response to resource or predictability constraints would contradict the bricolage process model. Direct observation that predictive performance never outranks validity in teams' deliberations would similarly contradict the claim that criteria are balanced rather than ordered.","tokens_in":32823,"feed_emoji":"🎯","tokens_out":5585,"duration_ms":62063,"temperature":0.7,"pith_summary":"This paper tries to establish that target variable construction in data science is a bricolage process rather than a top-down measurement process. When data scientists translate fuzzy concepts such as \"healthcare need\" or \"student writing authenticity\" into a concrete prediction target, they make do with whatever data are available, and they iteratively rework the target until it is fit for purpose. Based on interviews with fifteen data scientists in education and healthcare, the paper synthesizes five evaluation criteria — validity, simplicity, predictive performance, portability, and resource requirements — and five reformulation strategies: piggybacking, composing, swapping, bridging, and refining. If this account is right, many AI failures tied to target choice are not mainly errors in measurement theory but products of a constrained, improvisational work process, and interventions should scaffold that process rather than replace it with pre-planned protocols.","feed_headline":"Bricolage, not blueprints, shapes how data scientists pick targets","feed_subtitle":"Interviews with 15 data scientists reveal the five criteria and five strategies that turn fuzzy concepts into model targets.","key_machinery":"The central object is the concept of bricolage, borrowed from Lévi-Strauss, applied as a three-part structure: repertoire (the available data, tools, and domain knowledge), dialogue (iterative interaction with the data), and outcome (the final formulation). The paper's concrete machinery is the process model in Figure 2, together with the five evaluation criteria and five reformulation strategies that drive each iteration. The model does the explanatory work by showing how an apparently messy, improvised practice is orderly: each strategy is a response to a defect on a specific criterion, and stopping or discontinuing is a reasoned judgement that criteria are met or cannot be met.","core_discovery":"The central claim is that data scientists are bricoleurs of measurement: they assemble a prediction target from the materials already in hand, assess the result along several criteria, and apply reformulation strategies when a criterion is not met. The process model has data scientists starting from an initial formulation, then cycling through strategies — adopting an established precedent (piggybacking), combining outcomes (composing), changing the outcome (swapping), substituting a low-cost proxy for a gold standard (bridging), or adjusting the definition (refining) — until the target satisfies the criteria or the project is abandoned. Validity is one criterion among several, and participants treated its standards as elastic depending on stakes and resource constraints; predictive performance, by contrast, was often held to hard thresholds. Data-driven checks, such as poking holes in label definitions, and theory-driven reasoning, such as identifying spurious causes of outcomes, are both used to evaluate validity, though without formal measurement-theory vocabulary.","pith_inferences":["The bricolage account plausibly extends beyond classical predictive modeling to benchmark and evaluation design for generative AI, where \"toxicity\" or \"helpfulness\" targets are also built under resource constraints; one testable extension is to check whether evaluation designers use the same five strategies.","A practical consequence the paper leaves implicit is that recording the reformulation history of a target variable, which proxies were considered and why they were swapped, could serve as accountability documentation as valuable as model cards for understanding what a model really measures.","The reliance on experienced practitioners in two high-stakes domains raises the question of whether novices or data-rich domains such as social media analytics show a narrower or wider strategy repertoire; a comparative interview or trace study could test this.","The elasticity of validity standards invites a testable prediction: teams that document explicit per-criterion minimum standards before starting are less likely to erode validity under performance pressure than teams that do not."],"forward_implications":["If target variables are built by bricolage, then requiring pre-registered, fixed outcome definitions is likely fighting the process; more effective interventions would set minimum standards per criterion while preserving iterative reformulation.","Tooling for data scientists should support weighing trade-offs among validity, simplicity, predictive performance, portability, and resource requirements, rather than centering only predictive metrics such as AU-ROC.","Data scientists already evaluate validity through concrete practices such as testing predictions in deployment, probing mislabeled cases, and checking base-rate heuristics; these are natural hooks for validity scaffolding.","When no formulation satisfies the criteria within resource limits, discontinuing the project is a legitimate outcome of the bricolage process, and support tools should make that judgement easier rather than pushing teams to proceed."],"supporting_citations":[{"why":"Supplies the concept of bricolage and the \"science of the concrete\" that the paper applies to target variable construction.","marker":"[85]"},{"why":"Prior account of problem formulation as a negotiated translation between objectives and data, which this paper generalizes into a process model.","marker":"[106]"},{"why":"Defines the top-down measurement and validity framework that the paper argues is in tension with data science practice.","marker":"[62]"},{"why":"Provides the programmer-as-bricoleur analogy and the notion of dialogue with materials used to frame data scientists' work.","marker":"[136]"},{"why":"Prior interview study of label design in data science teams; this study shifts the focus to target variable construction.","marker":"[96]"},{"why":"Argues machine learning reframes measurement as prediction; the paper responds by showing validity concerns do surface in practice.","marker":"[97]"},{"why":"Concrete example of a proxy target variable (cost as a measure of healthcare need) that motivates the study.","marker":"[102]"},{"why":"Supplies validity sub-criteria used to classify participants' theory-driven evaluation practices.","marker":"[28]"}],"fun_headline_variants":["Data scientists bricolage target variables from fuzzy concepts","Bricolage process explains how data scientists pick model targets","Five criteria and five strategies shape data scientists' target choices","Data scientists assemble prediction targets using bricolage","Bricolage tactics turn fuzzy ideas into concrete model targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings depend on participants' retrospective stories and hypothetical-vignette reasoning being reliable evidence of how they actually build target variables in real projects.","fun_headline_variants_meta":{"raw":{"variants":["Data scientists bricolage target variables from fuzzy concepts","Bricolage process explains how data scientists pick model targets","Five criteria and five strategies shape data scientists' target choices","Data scientists assemble prediction targets using bricolage","Bricolage tactics turn fuzzy ideas into concrete model targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000561,"raw_usage":{"total_tokens":2673,"prompt_tokens":962,"completion_tokens":1711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":1629}},"tokens_in":578,"tokens_out":1711,"duration_ms":13787,"temperature":1.0,"reasoning_tokens":1629,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:19:55.672945+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A longitudinal trace of real data science projects, from version-controlled labels or screen recordings, that showed target variables are fixed before any data exploration and never swapped, composed, or refined in response to resource or predictability constraints would contradict the bricolage process model. Direct observation that predictive performance never outranks validity in teams' deliberations would similarly contradict the claim that criteria are balanced rather than ordered.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the concept of bricolage and the \"science of the concrete\" that the paper applies to target variable construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the programmer-as-bricoleur analogy and the notion of dialogue with materials used to frame data scientists' work."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior interview study of label design in data science teams; this study shifts the focus to target variable construction."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues machine learning reframes measurement as prediction; the paper responds by showing validity concerns do surface in practice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Concrete example of a proxy target variable (cost as a measure of healthcare need) that motivates the study."}],"review_version":1}