Pith. sign in

REVIEW 3 major objections 4 minor 152 references

Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Data scientists construct prediction targets through bricolage, making do with the data at hand rather than following a measurement plan.

desk verdict A careful interview study that gives the field a usable taxonomy of target-variable bricolage; the self-report evidence is the main soft spot, but not disqualifying. read the letter →

arxiv 2507.02819 v3 pith:3YL6O3M7 submitted 2025-07-03 cs.HC cs.CYcs.LG

classification cs.HCcs.CYcs.LG
keywords targetvariableconstructionbricolageproblemformulationmeasurementvaliditypredictivemodelingdatasciencepracticeinterviewstudymodelevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that target variable construction in data science is a bricolage process rather than a top-down measurement process. When data scientists translate fuzzy concepts such as "healthcare need" or "student writing authenticity" into a concrete prediction target, they make do with whatever data are available, and they iteratively rework the target until it is fit for purpose. Based on interviews with fifteen data scientists in education and healthcare, the paper synthesizes five evaluation criteria — validity, simplicity, predictive performance, portability, and resource requirements — and five reformulation strategies: piggybacking, composing, swapping, bridging, and refining. If this account is right, many AI failures tied to target choice are not mainly errors in measurement theory but products of a constrained, improvisational work process, and interventions should scaffold that process rather than replace it with pre-planned protocols.

What carries the argument

The central object is the concept of bricolage, borrowed from Lévi-Strauss, applied as a three-part structure: repertoire (the available data, tools, and domain knowledge), dialogue (iterative interaction with the data), and outcome (the final formulation). The paper's concrete machinery is the process model in Figure 2, together with the five evaluation criteria and five reformulation strategies that drive each iteration. The model does the explanatory work by showing how an apparently messy, improvised practice is orderly: each strategy is a response to a defect on a specific criterion, and stopping or discontinuing is a reasoned judgement that criteria are met or cannot be met.

What would settle it

A longitudinal trace of real data science projects, from version-controlled labels or screen recordings, that showed target variables are fixed before any data exploration and never swapped, composed, or refined in response to resource or predictability constraints would contradict the bricolage process model. Direct observation that predictive performance never outranks validity in teams' deliberations would similarly contradict the claim that criteria are balanced rather than ordered.

Watch

Extended reading notes

Core claim

The central claim is that data scientists are bricoleurs of measurement: they assemble a prediction target from the materials already in hand, assess the result along several criteria, and apply reformulation strategies when a criterion is not met. The process model has data scientists starting from an initial formulation, then cycling through strategies — adopting an established precedent (piggybacking), combining outcomes (composing), changing the outcome (swapping), substituting a low-cost proxy for a gold standard (bridging), or adjusting the definition (refining) — until the target satisfies the criteria or the project is abandoned. Validity is one criterion among several, and participants treated its standards as elastic depending on stakes and resource constraints; predictive performance, by contrast, was often held to hard thresholds. Data-driven checks, such as poking holes in label definitions, and theory-driven reasoning, such as identifying spurious causes of outcomes, are both used to evaluate validity, though without formal measurement-theory vocabulary.

Load-bearing premise

The findings depend on participants' retrospective stories and hypothetical-vignette reasoning being reliable evidence of how they actually build target variables in real projects.

Editorial extensions

If this is right

  • If target variables are built by bricolage, then requiring pre-registered, fixed outcome definitions is likely fighting the process; more effective interventions would set minimum standards per criterion while preserving iterative reformulation.
  • Tooling for data scientists should support weighing trade-offs among validity, simplicity, predictive performance, portability, and resource requirements, rather than centering only predictive metrics such as AU-ROC.
  • Data scientists already evaluate validity through concrete practices such as testing predictions in deployment, probing mislabeled cases, and checking base-rate heuristics; these are natural hooks for validity scaffolding.
  • When no formulation satisfies the criteria within resource limits, discontinuing the project is a legitimate outcome of the bricolage process, and support tools should make that judgement easier rather than pushing teams to proceed.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The bricolage account plausibly extends beyond classical predictive modeling to benchmark and evaluation design for generative AI, where "toxicity" or "helpfulness" targets are also built under resource constraints; one testable extension is to check whether evaluation designers use the same five strategies.
  • A practical consequence the paper leaves implicit is that recording the reformulation history of a target variable, which proxies were considered and why they were swapped, could serve as accountability documentation as valuable as model cards for understanding what a model really measures.
  • The reliance on experienced practitioners in two high-stakes domains raises the question of whether novices or data-rich domains such as social media analytics show a narrower or wider strategy repertoire; a comparative interview or trace study could test this.
  • The elasticity of validity standards invites a testable prediction: teams that document explicit per-criterion minimum standards before starting are less likely to erode validity under performance pressure than teams that do not.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper reports a semi-structured interview study with fifteen data scientists from education and healthcare. Using directed storytelling about past projects and a hypothetical vignette-based formulation task, the authors develop a process model in which target variable construction is a form of bricolage: data scientists iterate among candidate outcome definitions, balancing validity against predictive performance, simplicity, portability, and resource requirements, and applying five reformulation strategies (piggybacking, composing, swapping, bridging, refining). The paper also describes theory- and data-driven validity evaluation practices and proposes design implications for tool support and pedagogy.

Significance. The topic is important and understudied: target variable construction is a pivotal but often invisible step in predictive modeling, and the paper gives it a rich, grounded vocabulary. The five-criteria/five-strategies framework is plausible and supported by concrete participant quotes, and the authors connect their findings to measurement theory and to prior work on problem formulation in a useful way. The study is also transparent about its qualitative methods, including positionality and limitations. If the findings hold, they shift the design conversation from enforcing top-down measurement to scaffolding resource-constrained, iterative practice. The main weakness is evidential: the central claim about how data scientists actually construct target variables rests on retrospective self-reports and a single hypothetical task, with no triangulation from artifacts or observation, and the paper does not qualify this claim sufficiently.

major comments (3)
  1. [§3.1.1, §3.1.2, §5.5] The central claim that data scientists 'construct target variables through a bricolage process' (Abstract; §4) rests entirely on directed storytelling about past projects and a hypothetical vignette. Section 5.5 lists several limitations but does not acknowledge the self-report reliability threat: retrospective accounts are subject to hindsight bias and narrative smoothing, and vignette reasoning may differ from behavior under real resource constraints. Because the paper offers no observational or artifact-based triangulation, it cannot distinguish what data scientists do from what they say they do. This is a load-bearing issue for the empirical contribution; it can be fixed by reframing the findings as participants' reported and hypothetical reasoning and by adding an explicit limitation.
  2. [§3.1.2, Appendix A] The vignette stimulus was refined across successive interviews ('we refined the specificity of the scenario description and accompanying evaluation plots'), and one education-domain participant (P1) was shown the healthcare vignette. The paper does not address whether the changing stimulus affected the comparability of the reasoning elicited, nor how the mismatched-domain data were handled beyond 'taken into consideration during analysis.' This is relevant because the vignette is presented as a way to 'observe their reasoning in a new modeling context.' The authors should treat the vignette data as a supplementary source, analyze and report any version-related differences, and justify the inclusion of the mismatched-domain case.
  3. [§3.2, §3.3, Fig. 2] The paper generalizes from fifteen network-recruited participants in two domains to 'data scientists' in the abstract and throughout §4, and its process model (Fig. 2) is presented as the target variable construction process. Saturation is asserted rather than demonstrated, and the sample is small and domain-specific. The findings should be hedged as a provisional framework based on the studied population, with explicit statements that the criteria and strategies are likely to be extended or refined in other domains, organizational contexts, and experience levels.
minor comments (4)
  1. [Abstract, §4.1, Fig. 3] The abstract calls one criterion 'predictability,' whereas the findings and Figure 3 use 'predictive performance'; please unify the terminology.
  2. [§3.3] The paper reports that two authors independently performed open coding, but it does not report inter-coder agreement or provide a codebook excerpt; adding a table of themes with definitions and representative quotes would increase transparency and allow readers to assess the taxonomy's grounding.
  3. [§2.5] The sentence beginning 'Mussgnug [97] argue that...' has a subject-verb agreement error; the reference is to a single author and should read 'argues.'
  4. [Appendix A] Several screenshots in the arXiv version are barely legible or not described in detail; please ensure that all evaluation plots are readable and, if possible, accompanied by a description of what participants saw.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical taxonomy is grounded in interview data rather than derived from its own premises.

full rationale

The paper's central claim is a qualitative process model synthesized from fifteen interviews, not a mathematical derivation. There are no fitted parameters, no predictive equations, and no target quantity that is defined in terms of the inputs. The bricolage framing is acknowledged as a pre-analytic lens (Section 3.4: 'Team members with design experience had previously encountered Claude Lévi-Strauss’ concept of bricolage, which became our central theoretical framework'), but an interpretive framework is not a circular reduction: the authors report variation and cases that push against the frame (e.g., P12's project discontinuation in Section 4.2.6, P2's rejection of a latent-variable model, and participants' explicit trade-offs of validity against other criteria), which shows the data could have disconfirmed the bricolage characterization. The five criteria and five strategies are coding products illustrated with participant quotes rather than restatements of the authors' definitions. The few self-citations (e.g., [46], [47], [71]) appear only in background sections and design-implication discussions, not as load-bearing evidence for the empirical findings; the central claims rest on transcribed interviews, dual coding with reconciliation, and saturation. Methodological limitations (vignette refinement across interviews, P1 receiving the healthcare vignette, self-report reliance) are evidence-quality concerns, not circularity. No step in the paper reduces to its own inputs by construction, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or physical entities are introduced. The central claim rests on qualitative assumptions: self-reports proxy practice, saturation after 15 interviews, and the a priori bricolage lens. These are disclosed in the paper but not independently verified.

assumptions (3)
  • domain assumption Participants' retrospective interview accounts and hypothetical vignette reasoning reflect actual target variable construction practices.
    The entire empirical base is self-report (Sections 3.1 and 3.3); no observational or logged data on real formulation processes is used.
  • domain assumption Thematic saturation was reached after fifteen interviews in two domains, so the five criteria and five strategies are reasonably complete.
    Saturation is asserted in Section 3.3 based on recurring themes; the small, network-recruited sample and evolving vignette limit confidence in completeness.
  • ad hoc to paper Levi-Strauss's bricolage is an appropriate interpretive lens for target variable construction.
    The authors adopted bricolage as the central framework before and during analysis (Section 3.4); this frames what is noticed and reported, though the data are allowed to refine it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks." pith.science (2026). https://pith.science/paper/3YL6O3M7

@misc{pith2026250702819,
  author       = {Pith},
  title        = {Pith review of: Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling Tasks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3YL6O3M7}},
  note         = {Machine review of arXiv:2507.02819}
}
read the original abstract

Data scientists often formulate predictive modeling tasks involving fuzzy, hard-to-define concepts, such as the "authenticity" of student writing or the "healthcare need" of a patient. Yet the process by which data scientists translate fuzzy concepts into a concrete, proxy target variable remains poorly understood. We interview fifteen data scientists in education (N=8) and healthcare (N=7) to understand how they construct target variables for predictive modeling tasks. Our findings suggest that data scientists construct target variables through a bricolage process, in which they use creative and pragmatic approaches to make do with the limited data at hand. Data scientists attempt to satisfy five major criteria for a target variable through bricolage: validity, simplicity, predictability, portability, and resource requirements. To achieve this, data scientists adaptively apply problem (re)formulation strategies, such as swapping out one candidate target variable for another when the first fails to meet certain criteria (e.g., predictability), or composing multiple outcomes into a single target variable to capture a more holistic set of modeling objectives. Based on our findings, we present opportunities for future HCI, CSCW, and ML research to better support the art and science of target variable construction.

Figures

Figures reproduced from arXiv: 2507.02819 by the authors.

Figure 1
Figure 1. An illustration of the relationship between a target variable, prediction task, and problem formulation. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. An illustration of the target variable construction process presented in our findings. During target [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. The five criteria data scientists evaluated during target variable construction. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

152 extracted references · 63 canonical work pages

  1. [1]

    https://www2.ed.gov/rschstat/eval/high-school/early-warning-systems-brief.pdf

    2016. https://www2.ed.gov/rschstat/eval/high-school/early-warning-systems-brief.pdf

  2. [2]

    Amina A Abdu, Irene V Pasquetto, and Abigail Z Jacobs. 2023. An empirical analysis of racial categories in the algorithmic fairness literature. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 1324–1333. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW447. Publication date: November 2025. Measurement as ...

  3. [3]

    Robert Adcock and David Collier. 2001. Measurement validity: A shared standard for qualitative and quantitative research. American political science review 95, 3, 529–546

  4. [4]

    Yongsu Ahn and Yu-Ru Lin. 2019. Fairsight: Visual analytics for fairness in decision making. IEEE transactions on visualization and computer graphics 26, 1, 1086–1095

  5. [5]

    Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini, Shiladitya Dutta, Frances Dean, Inioluwa Deborah Raji, and Travis Zack. 2025. Medical Large Language Model Benchmarks Should Prioritize Construct Validity. arXiv preprint arXiv:2503.10694 (2025)

  6. [6]

    Sara Alspaugh, Nava Zokaei, Andrea Liu, Cindy Jin, and Marti A Hearst. 2018. Futzing and moseying: Interviews with professional data analysts on exploration practices. IEEE transactions on visualization and computer graphics 25, 1, 22–31

  7. [7]

    Audrey Amrein-Beardsley. 2014. Rethinking value-added models in education: Critical perspectives on tests and assessment-based accountability. Routledge

  8. [8]

    Zahra Ashktorab, Michael Desmond, Qian Pan, James M Johnson, Martin Santillan Cooper, Elizabeth M Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintalwar, and Werner Geyer. 2024. Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences.arXiv preprint arXiv:2410.00873 (2024)

Show all 152 references
  1. [9]

    Ted Baker and Reed E Nelson. 2005. Creating something from nothing: Resource construction through entrepreneurial bricolage. Administrative science quarterly 50, 3, 329–366

  2. [10]

    David J Bartholomew, Martin Knott, and Irini Moustaki. 2011. Latent variable models and factor analysis: A unified approach. John Wiley & Sons

  3. [11]

    Anton Barua, Stephen W Thomas, and Ahmed E Hassan. 2014. What are developers talking about? an analysis of topics and trends in stack overflow. Empirical software engineering 19, 619–654

  4. [12]

    Rachel KE Bellamy, Kuntal Dey, Michael Hind, Samuel C Hoffman, Stephanie Houde, Kalapriya Kannan, Pranay Lohia, Jacquelyn Martino, Sameep Mehta, Aleksandra Mojsilović, et al. 2019. AI Fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journa...

  5. [13]

    Aditya Bhattacharya, Simone Stumpf, Lucija Gosak, Gregor Stiglic, and Katrien Verbert. 2024. EXMOS: Explanatory Model Steering Through Multifaceted Explanations and Data Configurations. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–27

  6. [14]

    Christopher M Bishop. 1998. Latent variable models. In Learning in graphical models . Springer, 371–403

  7. [15]

    Alex Bogatu, Alvaro AA Fernandes, Norman W Paton, and Nikolaos Konstantinou. 2020. Dataset discovery in data lakes. In 2020 ieee 36th international conference on data engineering (icde) . IEEE, 709–720

  8. [16]

    Geoffrey C Bowker. 2000. Sorting Things Out: Classification and Its Consequences . MIT press

  9. [17]

    Jonathan Bragg, Mausam, and Daniel S Weld. 2018. Sprout: Crowd-powered task design for crowdsourcing. In Proceedings of the 31st annual acm symposium on user interface software and technology . 165–176

  10. [18]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2, 77–101

  11. [19]

    Monika Büscher, Satinder Gill, Preben Mogensen, and Dan Shapiro. 2001. Landscapes of practice: Bricolage as a method for situated design. Computer Supported Cooperative Work (CSCW) 10, 1–28

  12. [20]

    Ángel Alexander Cabrera, Will Epperson, Fred Hohman, Minsuk Kahng, Jamie Morgenstern, and Duen Horng Chau

  13. [21]

    Ángel Alexander Cabrera, Erica Fu, Donald Bertucci, Kenneth Holstein, Ameet Talwalkar, Jason I Hong, and Adam Perer. 2023. Zeno: An interactive framework for behavioral evaluation of machine learning. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Syst...

  14. [22]

    Ángel Alexander Cabrera, Marco Tulio Ribeiro, Bongshin Lee, Robert Deline, Adam Perer, and Steven M Drucker. 2023. What did my AI learn? How data scientists make sense of model behavior. ACM Transactions on Computer-Human Interaction 30, 1, 1–27

  15. [23]

    Raul Castro Fernandez. 2023. Data-sharing markets: model, protocol, and algorithms to incentivize the formation of data-sharing consortia. Proceedings of the ACM on Management of Data 1, 2, 1–25

  16. [24]

    Joseph Chee Chang, Saleema Amershi, and Ece Kamar. 2017. Revolt: Collaborative crowdsourcing for labeling machine learning datasets. In Proceedings of the 2017 CHI conference on human factors in computing systems . 2334–2346

  17. [25]

    Quan Ze Chen and Amy X Zhang. 2023. Judgment sieve: Reducing uncertainty in group judgments through interventions targeting ambiguity versus disagreement. Proceedings of the ACM on Human-Computer Interaction 7, CSCW2, 1–26

  18. [26]

    Hao-Fei Cheng, Logan Stapleton, Anna Kawakami, Venkatesh Sivaraman, Yanghuidi Cheng, Diana Qing, Adam Perer, Kenneth Holstein, Zhiwei Steven Wu, and Haiyi Zhu. 2022. How child welfare workers reduce racial disparities in algorithmic decisions. In Proceedings of the 2022 CHI Co...

  19. [27]

    Victoria Clarke and Virginia Braun. 2017. Thematic analysis. The journal of positive psychology 12, 3, 297–298

  20. [28]

    Amanda Coston, Anna Kawakami, Haiyi Zhu, Ken Holstein, and Hoda Heidari. 2023. A validity perspective on evaluating the justified use of data-driven decision-making algorithms. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, 690–704

  21. [29]

    Alfred W Crosby. 1997. The measure of reality: Quantification in Western Europe, 1250-1600 . Cambridge University Press

  22. [30]

    Tobias Daudert. 2020. A web-based collaborative annotation and consolidation tool. In Proceedings of the Twelfth Language Resources and Evaluation Conference . 7053–7059

  23. [31]

    Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 2022. Exploring how machine learning practitioners (try to) use fairness toolkits. In Proceedings of the 2022 ACM Conference on Fairness, Accountabili...

  24. [32]

    Michael Desmond, Michael Muller, Zahra Ashktorab, Casey Dugan, Evelyn Duesterwald, Kristina Brimijoin, Catherine Finegan-Dollak, Michelle Brachman, Aabhas Sharma, Narendra Nath Joshi, et al . 2021. Increasing the speed and accuracy of data labeling through an ai assisted inter...

  25. [33]

    Dennis Dingen, Marcel van’t Veer, Patrick Houthuizen, Eveline HJ Mestrom, Erik HHM Korsten, Arthur RA Bouwman, and Jarke Van Wijk. 2018. RegressionExplorer: Interactive exploration of logistic regression models with subgroup analysis. IEEE transactions on visualization and com...

  26. [34]

    Ellen A Drost. 2011. Validity and reliability in social science research. Education Research and perspectives 38, 1, 105–123

  27. [35]

    Pablo Duboue. 2020. The art of feature engineering: essentials for machine learning . Cambridge University Press

  28. [36]

    Raffi Duymedjian and Charles-Clemens Rüling. 2010. Towards a foundation of bricolage in organization and management theory. Organization studies 31, 2, 133–151

  29. [37]

    Shelley Evenson. 2016. Driving Service Design By Directed Storytelling. In Design for services. Routledge, 66–72

  30. [38]

    B Everett. 2013. An introduction to latent variable models . Springer Science & Business Media

  31. [39]

    Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker

  32. [40]

    Martin Fuller and Ryan Moore. 2017. An Analysis of Jane Jacobs’s The Death and Life of Great American Cities . Macat Library

  33. [41]

    Dalia Gala, Milo Phillips-Brown, Naman Goel, Carinal Prunkl, Laura Alvarez Jubete, Ray Eitel-Porter, et al . 2024. FairTargetSim: An Interactive Simulator for Understanding and Explaining the Fairness Effects of Target Variable Definition. arXiv preprint arXiv:2403.06031

  34. [42]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12, 86–92

  35. [43]

    Darren Gergle and Desney S Tan. 2014. Experimental research in HCI. In Ways of Knowing in HCI. Springer, 191–227

  36. [44]

    Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley. 2000. PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiolo...

  37. [45]

    Nitesh Goyal, Ian D Kivlichan, Rachel Rosen, and Lucy Vasserman. 2022. Is your toxicity my toxicity? exploring the impact of rater identity on toxicity annotation. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2, 1–28

  38. [46]

    Luke Guerdan, Amanda Coston, Kenneth Holstein, and Zhiwei Steven Wu. 2023. Counterfactual prediction under outcome measurement error. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . 1584–1598

  39. [47]

    Luke Guerdan, Amanda Coston, Zhiwei Steven Wu, and Kenneth Holstein. 2023. Ground (less) truth: A causal framework for proxy labels in human-algorithm decision-making. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency. 688–704

  40. [48]

    Philip J Guo, Sean Kandel, Joseph M Hellerstein, and Jeffrey Heer. 2011. Proactive wrangling: Mixed-initiative end-user programming of data transformation scripts. In Proceedings of the 24th annual ACM symposium on User interface software and technology. 65–74

  41. [49]

    Ian Hacking. 1999. The social construction of what? Harvard university press

  42. [50]

    MD Romael Haque, Devansh Saxena, Katy Weathington, Joseph Chudzik, and Shion Guha. 2024. Are we asking the right questions?: Designing for community stakeholders’ interactions with ai in policing. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . 1–20

  43. [51]

    Bill Harding. 2021. Software effort estimates vs popular developer productivity metrics . Technical Report. GitClear. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW447. Publication date: November 2025. Measurement as Bricolage CSCW447:29

  44. [52]

    Emma Harvey, Hauke Sandhaus, Abigail Z Jacobs, Emanuel Moss, and Mona Sloane. 2024. The Cadaver in the Machine: The Social Practices of Measurement and Validation in Motion Capture Technology. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–23

  45. [53]

    Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, and Hanna Wallach. 2025. Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems. arXiv preprint arXiv:2506.04482 (2025)

  46. [54]

    Orit Hazzan and Jim Tomayko. 2004. Human aspects of software engineering: The case of extreme programming. In International Conference on Extreme Programming and Agile Processes in software Engineering . Springer, 303–311

  47. [55]

    Zeyu He, Saniya Naphade, and Ting-Hao Kenneth Huang. 2025. Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–33

  48. [56]

    Douglas D Heckathorn. 2011. Comment: Snowball versus respondent-driven sampling. Sociological methodology 41, 1, 355–366

  49. [57]

    Nicole Hoess, Carlos Paradis, Rick Kazman, and Wolfgang Mauerer. 2025. Does the Tool Matter? Exploring Some Causes of Threats to Validity in Mining Software Repositories. arXiv preprint arXiv:2501.15114 (2025)

  50. [58]

    Jake M Hofman, Angelos Chatzimparmpas, Amit Sharma, Duncan J Watts, and Jessica Hullman. 2023. Pre-registration for predictive modeling. arXiv preprint arXiv:2311.18807

  51. [59]

    Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learning systems: What do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–16

  52. [60]

    Youyang Hou and Dakuo Wang. 2017. Hacking with NPOs: collaborative analytics and broker roles in civic data hackathons. Proceedings of the ACM on Human-Computer Interaction 1, CSCW, 1–16

  53. [61]

    Steven J Ingels, Daniel J Pratt, James E Rogers, Peter H Siegel, and Ellen S Stutts. 2004. Education Longitudinal Study of 2002: Base Year Data File User’s Manual. NCES 2004-405. National Center for Education Statistics

  54. [62]

    Abigail Z Jacobs and Hanna Wallach. 2021. Measurement and fairness. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. 375–385

  55. [63]

    Alistair Johnson, Lucas Bulgarelli, Tom Pollard, Steven Horng, Leo Anthony Celi, and Roger Mark. 2020. Mimic-iv. PhysioNet. A vailable online at: https://physionet. org/content/mimiciv/1.0/(accessed August 23, 2021), 49–55

  56. [64]

    Eunice Jun, Melissa Birchfield, Nicole De Moura, Jeffrey Heer, and Rene Just. 2022. Hypothesis formalization: Empirical findings, software limitations, and design implications. ACM Transactions on Computer-Human Interaction (TOCHI) 29, 1 (2022), 1–28

  57. [65]

    Eunice Jun, Audrey Seo, Jeffrey Heer, and René Just. 2022. Tisane: Authoring statistical models via formal reasoning from conceptual and data relationships. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–16

  58. [66]

    Ji-Youn Jung, Devansh Saxena, Minjung Park, Jini Kim, Jodi Forlizzi, Kenneth Holstein, and John Zimmerman. 2025. Making the Right Thing: Bridging HCI and Responsible AI in Early-Stage AI Concept Selection. arXiv preprint arXiv:2506.17494 (2025)

  59. [67]

    Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J

    V K. Chaithanya Manam, Dwarakanath Jampani, Mariam Zaim, Meng-Han Wu, and Alexander J. Quinn. 2019. Taskmate: A mechanism to improve the quality of instructions in crowdsourcing. In Companion Proceedings of The 2019 World Wide Web Conference. 1121–1130

  60. [68]

    Sean Kandel, Andreas Paepcke, Joseph Hellerstein, and Jeffrey Heer. 2011. Wrangler: Interactive visual specification of data transformation scripts. In Proceedings of the sigchi conference on human factors in computing systems . 3363–3372

  61. [69]

    Sean Kandel, Andreas Paepcke, Joseph M Hellerstein, and Jeffrey Heer. 2012. Enterprise data analysis and visualization: An interview study. IEEE transactions on visualization and computer graphics 18, 12, 2917–2926

  62. [70]

    Sayash Kapoor and Arvind Narayanan. 2023. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4, 9 (2023)

  63. [71]

    Anna Kawakami, Amanda Coston, Haiyi Zhu, Hoda Heidari, and Kenneth Holstein. 2024. The Situate AI Guidebook: Co-Designing a Toolkit to Support Multi-Stakeholder, Early-stage Deliberations Around Public Sector AI Proposals. In Proceedings of the CHI Conference on Human Factors ...

  64. [72]

    Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, and Kenneth Holstein. 2022. Improving human-AI partnerships in child welfare: understanding worker practices, challenges, and desires for a...

  65. [73]

    Mary Beth Kery and Brad A Myers. 2017. Exploring exploratory programming. In 2017 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC) . IEEE, 25–29

  66. [74]

    Mary Beth Kery, Marissa Radensky, Mahima Arya, Bonnie E John, and Brad A Myers. 2018. The story in the notebook: Exploratory data science using a literate programming tool. InProceedings of the 2018 CHI conference on human factors Proc. ACM Hum.-Comput. Interact., Vol. 9, No. ...

  67. [75]

    Miryung Kim, Thomas Zimmermann, Robert DeLine, and Andrew Begel. 2016. The emerging role of data scientists on software development teams. In Proceedings of the 38th International Conference on Software Engineering . 96–107

  68. [76]

    Amy J Ko, Robin Abraham, Laura Beckwith, Alan Blackwell, Margaret Burnett, Martin Erwig, Chris Scaffidi, Joseph Lawrance, Henry Lieberman, Brad Myers, et al. 2011. The state of the art in end-user software engineering. ACM Computing Surveys (CSUR) 43, 3, 1–44

  69. [77]

    Sean Kross and Philip Guo. 2021. Orienting, framing, bridging, magic, and counseling: How data scientists navigate the outer loop of client collaborations in industry and academia. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2, 1–28

  70. [78]

    Tzu-Sheng Kuo, Hong Shen, Jisoo Geum, Nev Jones, Jason I Hong, Haiyi Zhu, and Kenneth Holstein. 2023. Under- standing Frontline Workers’ and Unhoused Individuals’ Perspectives on AI Used in Homeless Services. In Proceedings of the 2023 CHI Conference on Human Factors in Comput...

  71. [79]

    Bruno Latour and Steve Woolgar. 2013. Laboratory life: The construction of scientific facts . Princeton university press

  72. [80]

    Jason Lefever, Yuanfang Cai, Humberto Cervantes, Rick Kazman, and Hongzhou Fang. 2021. On the lack of consensus among technical debt detection tools. In 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP) . IEEE, 121–130

  73. [81]

    Jessica Liu, Huaming Chen, Jun Shen, and Kim-Kwang Raymond Choo. 2024. FairCompass: Operationalising fairness in machine learning. IEEE Transactions on Artificial Intelligence

  74. [82]

    Yu Lu Liu, Su Lin Blodgett, Jackie Chi Kit Cheung, Q Vera Liao, Alexandra Olteanu, and Ziang Xiao. 2024. ECBD: Evidence-centered benchmark design for NLP. arXiv preprint arXiv:2406.08723 (2024)

  75. [83]

    Panagiotis Louridas. 1999. Design as bricolage: Anthropology meets design thinking. Design Studies 20, 6 (1999), 517–535

  76. [84]

    Scott Lundberg. 2017. A unified approach to interpreting model predictions. arXiv preprint arXiv:1705.07874

  77. [85]

    Claude Lévi-Strauss. 1966. The savage mind. University of Chicago Press

  78. [86]

    VK Manam and Alexander Quinn. 2018. Wingit: Efficient refinement of unclear task instructions. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing , Vol. 6. 108–116

  79. [87]

    Yaoli Mao, Dakuo Wang, Michael Muller, Kush R Varshney, Ioana Baldini, Casey Dugan, and Aleksandra Mojsilović

  80. [88]

    Sara Mateus and Soumodip Sarkar. 2024. Bricolage–a systematic review, conceptualization, and research agenda. Entrepreneurship & Regional Development, 1–22

  81. [89]

    Jennifer Mickel. 2024. Racial/Ethnic Categories in AI and Algorithmic Fairness: Why They Matter and What They Represent. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 2484–2494

  82. [90]

    How data scientistswork together with domain experts in scientific collaborations: To find the right answer or to ask the right question? Proceedings of the ACM on Human-Computer Interaction 3, GROUP, 1–23

  83. [91]

    Smitha Milli, Luca Belli, and Moritz Hardt. 2021. From optimizing engagement to measuring value. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency . 714–722

  84. [92]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model cards for model reporting. In Proceedings of the conference on fairness, accountability, and transparency. 220–229

  85. [93]

    Delbert C Miller and Neil J Salkind. 2002. Handbook of research design and social measurement . Sage

  86. [94]

    Sendhil Mullainathan and Ziad Obermeyer. 2021. On the inequity of predicting A while hoping for B. In AEA Papers and Proceedings, Vol. 111. American Economic Association 2014 Broadway, Suite 305, Nashville, TN 37203, 37–42

  87. [95]

    Michael Muller, Ingrid Lange, Dakuo Wang, David Piorkowski, Jason Tsay, Q Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How data science workers work with data: Discovery, capture, curation, design, creation. InProceedings of the 2019 CHI conference on human factors in co...

  88. [96]

    A Mol. 2002. The body multiple: Ontology in medical practice . Duke University Press

  89. [97]

    Alexander Martin Mussgnug. 2022. The predictive reframing of machine learning applications: good predictions and bad measurements. European Journal for Philosophy of Science 12, 3 (2022), 55

  90. [98]

    Brad A Myers, Amy J Ko, Thomas D LaToza, and YoungSeok Yoon. 2016. Programmers are users too: Human-centered methods for improving programming tools. Computer 49, 7, 44–52

  91. [99]

    Michael Muller, Christine T Wolf, Josh Andres, Michael Desmond, Narendra Nath Joshi, Zahra Ashktorab, Aabhas Sharma, Kristina Brimijoin, Qian Pan, Evelyn Duesterwald, et al. 2021. Designing ground truth and the social life of labels. In Proceedings of the 2021 CHI conference o...

  92. [100]

    Hiroki Nakayama, Takahiro Kubo, Junya Kamura, Yasufumi Taniguchi, and Xu Liang. 2018. doccano: Text Annotation Tool for Human. https://github.com/doccano/doccano Software available from https://github.com/doccano/doccano. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Articl...

  93. [101]

    Fatemeh Nargesian, Erkang Zhu, Renée J Miller, Ken Q Pu, and Patricia C Arocena. 2019. Data lake management: challenges and opportunities. Proceedings of the VLDB Endowment 12, 12, 1986–1989

  94. [102]

    Nadia Nahar, Shurui Zhou, Grace Lewis, and Christian Kästner. 2022. Collaboration challenges in building ml-enabled systems: Communication, documentation, engineering, and process. In Proceedings of the 44th international conference on software engineering. 413–425

  95. [103]

    Jorge A Osorio, Michel RV Chaudron, and Werner Heijstek. 2011. Moving from waterfall to iterative development: An empirical evaluation of advantages, disadvantages and risks of RUP. In 2011 37th EUROMICRO Conference on Software Engineering and Advanced Applications . IEEE, 453–460

  96. [104]

    Qian Pan, Zahra Ashktorab, Michael Desmond, Martin Santillan Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-judge. arXiv preprint arXiv:2407.03479 (2024)

  97. [105]

    Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan. 2019. Dissecting racial bias in an algorithm used to manage the health of populations. Science 366, 6464, 447–453

  98. [106]

    Samir Passi and Solon Barocas. 2019. Problem formulation and fairness. In Proceedings of the conference on fairness, accountability, and transparency. 39–48

  99. [107]

    Samir Passi and Steven J Jackson. 2018. Trust in data science: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on human-computer interaction 2, CSCW, 1–28

  100. [108]

    Bo Pang, Erik Nijkamp, and Ying Nian Wu. 2020. Deep learning with tensorflow: A review. Journal of Educational and Behavioral Statistics 45, 2, 227–248

  101. [109]

    Juan Carlos Perdomo. 2023. The Relative Value of Prediction in Algorithmic Decision Making. arXiv preprint arXiv:2312.08511

  102. [110]

    Kathleen H Pine and Max Liboiron. 2015. The politics of measurement and action. In Proceedings of the 33rd annual ACM conference on human factors in computing systems . 3147–3156

  103. [111]

    Norman W Paton, Jiaoyan Chen, and Zhenyu Wu. 2023. Dataset discovery and exploration: A survey. Comput. Surveys 56, 4, 1–37

  104. [112]

    Mahima Pushkarna, Andrew Zaldivar, and Oddur Kjartansson. 2022. Data cards: Purposeful and transparent dataset documentation for responsible ai. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Trans- parency. 1776–1826

  105. [113]

    Iyad Rahwan, Manuel Cebrian, Nick Obradovich, Josh Bongard, Jean-François Bonnefon, Cynthia Breazeal, Jacob W Crandall, Nicholas A Christakis, Iain D Couzin, Matthew O Jackson, et al. 2019. Machine behaviour. Nature 568, 7753, 477–486

  106. [114]

    Theodore M Porter. 1996. Trust in numbers: The pursuit of objectivity in science and public life . Princeton University Press

  107. [115]

    Inioluwa Deborah Raji, I Elizabeth Kumar, Aaron Horowitz, and Andrew Selbst. 2022. The fallacy of AI functionality. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 959–972

  108. [116]

    O’Reilly Media, Inc

    Tye Rattenbury, Joseph M Hellerstein, Jeffrey Heer, Sean Kandel, and Connor Carreras. 2017. Principles of data wrangling: Practical techniques for data preparation . " O’Reilly Media, Inc. "

  109. [117]

    Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. 2021. AI and the everything in the whole wide world benchmark. arXiv preprint arXiv:2111.15366 (2021)

  110. [118]

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Anchors: High-precision model-agnostic explanations. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32

  111. [119]

    Pedro Saleiro, Benedict Kuester, Loren Hinkson, Jesse London, Abby Stevens, Ari Anisfeld, Kit T Rodolfa, and Rayid Ghani. 2018. Aequitas: A bias and fairness audit toolkit. arXiv preprint arXiv:1811.05577

  112. [120]

    Why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. " Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135–1144

  113. [121]

    Devansh Saxena, Ji-Youn Jung, Jodi Forlizzi, Kenneth Holstein, and John Zimmerman. 2025. AI Mismatches: Identifying Potential Algorithmic Harms Before AI Development. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–23

  114. [122]

    James J Schlesselman. 1982. Case-control studies: design, conduct, analysis . Vol. 2. Oxford university press

  115. [123]

    Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in qualitative research: exploring its conceptualization and operationalization. Quality & quantity 52, 1893–1907

  116. [124]

    James C Scott. 2020. Seeing like a state: How certain schemes to improve the human condition have failed . yale university Press

  117. [125]

    2005.Human-centered software engineering-integrating usability in the software development lifecycle

    Ahmed Seffah, Jan Gulliksen, and Michel C Desmarais. 2005.Human-centered software engineering-integrating usability in the software development lifecycle . Vol. 8. Springer Science & Business Media. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW447. Publication ...

  118. [126]

    Moritz Schubotz, Ankit Satpute, André Greiner-Petter, Akiko Aizawa, and Bela Gipp. 2022. Caching and Repro- ducibility: Making Data Science Experiments Faster and FAIRer. Frontiers in Research Metrics and Analytics 7 (2022), 861944

  119. [127]

    Amit Sharma and Emre Kiciman. 2020. DoWhy: An end-to-end library for causal inference. arXiv preprint arXiv:2011.04216

  120. [128]

    Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Advances in Neural Information Processing Systems 36

  121. [129]

    Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024. Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software an...

  122. [130]

    Venkatesh Sivaraman, Anika Vaishampayan, Xiaotong Li, Brian R Buck, Ziyong Ma, Richard D Boyce, and Adam Perer. 2025. Tempo: Helping Data Scientists and Domain Experts Collaboratively Specify Predictive Modeling Tasks. In Proceedings of the 2025 CHI Conference on Human Factors...

  123. [131]

    Wesley G Skogan. 1999. Measuring what matters. In Proceedings from the policing research institute meetings . Department of Justice, 37–88

  124. [132]

    Kultar Singh. 2007. Quantitative social research methods . Sage

  125. [133]

    Lucy Suchman. 1993. Do categories have politics? The language/action perspective reconsidered. Computer supported cooperative work (CSCW) 2, 177–190

  126. [134]

    Annalisa Szymanski, Simret Araya Gebreegziabher, Oghenemaro Anuyah, Ronald A Metoyer, and Toby Jia-Jun Li

  127. [135]

    Jonathan Stray, Ivan Vendrov, Jeremy Nixon, Steven Adler, and Dylan Hadfield-Menell. 2021. What are you optimizing for? aligning recommender systems with human values. arXiv preprint arXiv:2107.10939

  128. [136]

    Sherry Turkle. 2011. Life on the Screen . Simon and Schuster

  129. [137]

    Bill Turque. 2012. Creative... motivating’and fired. The Washington Post 6

  130. [138]

    Anna Vallgårda and Ylva Fernaeus. 2015. Interaction design as a bricolage practice. In Proceedings of the ninth international conference on tangible, embedded, and embodied interaction . 173–180

  131. [139]

    Florian Tramer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu, Jean-Pierre Hubaux, Mathias Humbert, Ari Juels, and Huang Lin. 2017. Fairtest: Discovering unwarranted associations in data-driven applications. In 2017 IEEE European Symposium on Security and Privacy (EuroS&P) ....

  132. [140]

    Hanna Wallach, Meera Desai, Nicholas Pangakis, A Feder Cooper, Angelina Wang, Solon Barocas, Alexandra Choulde- chova, Chad Atalla, Su Lin Blodgett, Emily Corvi, et al. 2024. Evaluating Generative AI Systems is a Social Science Measurement Challenge. arXiv preprint arXiv:2411....

  133. [141]

    Jamelle Watson-Daniels, Solon Barocas, Jake M Hofman, and Alexandra Chouldechova. 2023. Multi-Target Multiplicity: Flexibility and Fairness in Target Specification under Resource Constraints. InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparenc...

  134. [142]

    Hilde Weerts, Miroslav Dudík, Richard Edgar, Adrin Jalali, Roman Lutz, and Michael Madaio. 2023. Fairlearn: Assessing and improving fairness of ai systems. Journal of Machine Learning Research 24, 257, 1–8

  135. [143]

    Anna Elisabeth Van’t Veer and Roger Giner-Sorolla. 2016. Pre-registration in social psychology—A discussion and suggested template. Journal of experimental social psychology 67, 2–12

  136. [144]

    Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019. Goals, process, and challenges of exploratory data analysis: An interview study. arXiv preprint arXiv:1911.00568

  137. [145]

    Kanit Wongsuphasawat, Dominik Moritz, Anushka Anand, Jock Mackinlay, Bill Howe, and Jeffrey Heer. 2015. Voyager: Exploratory analysis via faceted browsing of visualization recommendations. IEEE transactions on visualization and computer graphics 22, 1, 649–658

  138. [146]

    Feng Xie, Jun Zhou, Jin Wee Lee, Mingrui Tan, Siqi Li, Logasan S/O Rajnthern, Marcel Lucas Chee, Bibhas Chakraborty, An-Kwok Ian Wong, Alon Dagan, et al. 2022. Benchmarking emergency department prediction models with machine learning and public electronic health records. Scien...

  139. [147]

    James Wexler, Mahima Pushkarna, Tolga Bolukbasi, Martin Wattenberg, Fernanda Viégas, and Jimbo Wilson. 2019. The what-if tool: Interactive probing of machine learning models. IEEE transactions on visualization and computer graphics 26, 1, 56–65

  140. [148]

    medical acuity

    Amy X Zhang, Michael Muller, and Dakuo Wang. 2020. How do data science workers collaborate? roles, workflows, and tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1, 1–23. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW447. Publication date: Nov...

  141. [151]

    Jing Nathan Yan, Ziwei Gu, and Jeffrey M Rzeszotarski. 2021. Tessera: Discretizing data analysis workflows on a task level. In Proceedings of the 2021 CHI conference on human factors in computing systems . 1–15

  142. [2018]

    In 2018 IEEE 34th International Conference on Data Engineering (ICDE)

    Aurum: A data discovery system. In 2018 IEEE 34th International Conference on Data Engineering (ICDE) . IEEE, 1001–1012

  143. [2019]

    In 2019 IEEE Conference on Visual Analytics Science and Technology (V AST)

    FairVis: Visual analytics for discovering intersectional bias in machine learning. In 2019 IEEE Conference on Visual Analytics Science and Technology (V AST). IEEE, 46–56

  144. [2024]

    arXiv preprint arXiv:2410.02054 (2024)

    Comparing Criteria Development Across Domain Experts, Lay Users, and Models in Large Language Model Evaluation. arXiv preprint arXiv:2410.02054 (2024)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.