Pith. sign in

REVIEW 3 major objections 4 minor 5 cited by

The Evolution of LLM Adoption in Industry Data Curation Practices

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Industry data curation is shifting from manual, bottom-up inspection to LLM-first, top-down analysis, as practitioners layer LLM-generated 'silver' datasets under expert 'golden' ones.

desk verdict Useful new dataset taxonomy and workflow framing, but the 'evolution/paradigm shift' claim outruns the cross-sectional evidence. read the letter →

arxiv 2412.16089 v1 pith:5X3QGFQL submitted 2024-12-20 cs.HC cs.AI

classification cs.HCcs.AI
keywords DatacurationLargelanguagemodelsqualityanalysisworkflowsExploratoryTextpractitioners
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tracks how data practitioners at a large technology company adopted LLMs for curating unstructured text data across roughly eighteen months, from a spring 2023 survey through late-2024 interviews and a design-probe user study. Its central claim is that LLM adoption is transforming data understanding from a heuristic-first, bottom-up process, where people label and aggregate individual rows, into an insights-first, top-down process, where people ask an LLM for high-level summaries and drill into specifics only when needed. The paper also claims that this shift is producing a multi-tier dataset hierarchy: expert-labeled 'golden' datasets, LLM-generated 'silver' datasets, and 'super-golden' datasets validated by diverse expert teams for benchmarking against human performance. If true, the burden of data-quality evaluation moves from individual manual inspection toward shared frameworks for validating LLM-produced labels and summaries.

What carries the argument

The argument is carried by two linked mechanisms. The first is the inverted analysis pipeline, a documented shift from bottom-up data aggregation, in which practitioners label individual rows and then count them up, to top-down extraction, in which an LLM is asked directly for themes, summaries, or outlier candidates and the practitioner drills into raw examples only to verify or extract evidence. The second is the multi-tier dataset hierarchy named 'golden,' 'silver,' and 'super-golden': golden sets are expert labels, silver sets are predominantly LLM-generated labels used to complement them, and super-golden sets are small, painstakingly validated collections created by diverse expert teams to serve as a higher-authority ground truth for comparing LLMs with humans. These two mechanisms do the explanatory work: the pipeline inversion explains the reported efficiency gains, and the hierarchy explains how practitioners manage quality and trust as LLM-produced labels enter the pipeline.

What would settle it

Instrument real data-curation pipelines to log whether practitioners, after adopting LLM summarization, still open and inspect raw rows before trusting an LLM-derived insight; if a majority continue to inspect raw data first in most tasks, the claimed inversion from bottom-up to top-down analysis does not hold.

Watch

Extended reading notes

Core claim

The authors set out to measure whether practitioners were using LLMs for data curation; within six months the question flipped from whether to how. The paper's core discovery is that LLMs are enabling practitioners to reverse the traditional order of data analysis: instead of building insights bottom-up by labeling and aggregating individual data points, practitioners now generate high-level, insight-first summaries with LLMs and return to the raw data only when a specific claim needs evidence. Alongside this workflow inversion, the authors observe a new tiered dataset economy. Golden datasets, the expert-labeled gold standard, are now being supplemented by 'silver' datasets, whose labels are produced largely by LLMs and used for high-traffic or initialization purposes, and by 'super-golden' datasets, small, rigorously validated collections assembled by diverse expert teams to serve as higher authority for benchmarking LLMs against humans. The authors interpret these changes as a transformation in how practitioners engage with their data, while noting persistent barriers of reliability, cost, unfamiliarity, and content-refusal responses.

Load-bearing premise

The narrative of a fundamental transformation rests on fusing three studies with different samples, methods, and timings into a single trajectory, and on assuming that what practitioners said and did during hour-long sessions with the authors' own prototype tools reflects their real, sustained workflows.

Editorial extensions

If this is right

  • If the top-down shift holds, tool builders should prioritize LLM summarization, explanation, and outlier detection inside spreadsheets and notebooks over manual labeling features, since those are the capabilities practitioners reach for first.
  • Silver datasets will become a standard layer in production data pipelines, which means validation, error analysis, and bias auditing for LLM-generated labels will be needed before those datasets are used for training or evaluation.
  • Super-golden datasets, being small, expensive, and expert-assembled, will grow in importance as reference standards, shifting budgets and timelines for evaluation dataset construction.
  • Data quality will be defined collaboratively by safety teams, domain experts, and engineering managers rather than by a single metric, so tools must support consensus-building and inter-team sharing.
  • Reliability, latency, cost, and refusal behavior will determine which curation tasks are automated first and which remain human-only in the near term.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the inversion is real, a testable consequence is that practitioners who adopt LLM-first workflows may gradually lose the tacit, row-level familiarity that made them good at spotting when an insight is wrong; the paper's C3 quote gestures at this risk but does not develop it, so a longitudinal study of error-detection skill after top-down adoption would test it.
  • The golden/silver/super-golden hierarchy likely extends beyond text to images and audio, where the same pattern of LLM-generated labels under expert validation could appear; the authors list multimodal data only as future work.
  • The spreadsheet probe's broad appeal across technical and non-technical roles suggests that prompt-in-cell interfaces may become a default collaboration surface for data work, a trajectory the paper does not fully commit to because its probes were built by the authors.
  • A quantitative extension would measure the share of silver versus golden labels in production datasets over time; if silver is merely a stopgap rather than a durable tier, its share should plateau rather than grow.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper presents three studies conducted at a single large technology company to trace the evolution of LLM adoption in data curation: an exploratory survey of 84 employees (Q2 2023), interviews with 10 data practitioners and tool developers (late 2023/early 2024), and a user study with 12 practitioners using two author-built LLM-based design probes embedded in spreadsheets and notebooks (Q3 2024). The paper argues that practitioners moved from heuristic-first, bottom-up data analysis to insights-first, top-down workflows, that data quality is becoming a collaborative and subjective construct, and that multi-tiered dataset hierarchies ('golden,' 'silver,' 'super-golden') are emerging. It also reports perceived efficiency gains, barriers to adoption, and implications for future tools.

Significance. The paper offers a rare insider snapshot of LLM adoption in an industrial data-curation setting, with concrete quotes, role-specific participant tables, and a transparent description of the coding process; the design probes themselves are a useful vehicle for eliciting practitioner reactions. If the empirical claims are read as qualitative, contextual observations, the paper provides timely evidence of how practitioners perceive LLM capabilities. However, the headline conclusions—'fundamental transformation' and 'paradigm shift'—go beyond what the study design can establish. The three studies are cross-sectional, non-overlapping in participants and instruments, and the 2024 session is a reaction-to-probe study rather than an observational or longitudinal study. The paper's contribution is therefore strongest as a design-oriented snapshot of perceptions and anticipated workflows, and weaker as a demonstration of actual behavioral change.

major comments (3)
  1. [Section 8; see also Sections 6.2 and 7.1.2] The central conclusion that LLMs are causing a 'fundamental transformation' and 'paradigm shift' (Section 8) is not supported by the study design. The three empirical studies are cross-sectional snapshots of different, non-overlapping samples with different instruments; no participant was measured at two time points. The top-down workflow evidence in Section 7.1.2 comes from the 2024 user study, where participants were given the authors' LLM-centric spreadsheet and notebook probes and asked to try them (Section 6.2), so the observed behavior may reflect the probes' affordances rather than a pre-existing change in practice. The paper should either reframe the conclusion as evidence of participants' receptivity to, and reported use of, LLM-based workflows, or provide observational or baseline evidence showing the shift outside the probe context. The limitation paragraph in Section 7.2 acknowledges the snapshot character of the data, but Section 8 does not carry those caveats.
  2. [Sections 6.3.1 and 6.7] The claims of 'transformative efficiency gains' and 'emerging dataset hierarchies' are stated more strongly than the evidence allows. The efficiency result is based on participants' estimates and hypothetical projections (e.g., C2's 45 minutes to code 75 responses; T3's '500 data points per hour') rather than on measured before/after productivity. The silver and super-golden dataset hierarchy is supported mainly by two quotes from A1 and T4 within an N=12 study, with no indication of prevalence. These findings should be labeled as perceived and anticipated benefits and as emergent themes from a small sample, and the language in Section 8 ('rapidly growing reliance') should be correspondingly moderated.
  3. [Sections 3, 4, 6, and 7.1] The three studies measure different constructs, so the narrative of an 'evolution' is an analytic reconstruction rather than an observed trajectory. Section 3 surveys general development-task LLM adoption, Section 4 interviews practitioners about data needs and challenges, and Section 6 probes reactions to specific prototypes; the synthesis in Section 7.1 and the conclusion in Section 8 weave these together as a temporal progression. Since the samples and constructs do not line up, the paper should explicitly present the composite as an interpretive narrative across independent studies, and should identify which specific claims are supported by which study, rather than implying a continuous longitudinal measure.
minor comments (4)
  1. [Sections 4.1, 4.3.2, 4.4, 6.8, and 7.1.2] Participant identifiers are used inconsistently: Section 4.4 refers to a sample of 12 participants although Section 4.1 and Table 1 report N=10; Section 4.3.2 quotes 'T4' where Section 4 participants are coded U/D; Section 6.8 quotes 'S1' though Table 3 uses T/A/C codes; and Section 7.1.2 refers to 'R2 and R3' that never appear in any participant table.
  2. [Section 4.3.2] There is a typo in Section 4.3.2: 'development of language language models' should read 'large language models.'
  3. [Abstract and Sections 4, 5, and 8] The timeline is misaligned: the abstract dates the expert interviews to Q3 2023, but Section 4 states they were conducted between November 2023 and January 2024; the abstract dates the user study to Q2 2024, while Section 5 says Q3 2024; Section 8's 'just six months after' should be reconciled with these dates.
  4. [Section 5.2] Section 5.2 contains a sentence fragment: 'Since Python notebooks offer greater flexibility than spreadsheets. We provide two additional features:' should be joined into one sentence or rephrased.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity was found; the paper's findings rest on quoted practitioner reports and reproduced study data, and the design-probe goal evaluation is not a reduction.

full rationale

This paper is a qualitative multi-study report rather than a formal derivation, and no load-bearing step reduces by construction to its inputs. The central 'evolution' and 'paradigm shift' claims are supported by distinct cross-sectional studies whose evidence is reproduced in the paper: the survey statistics in Section 3, the interview themes and quotes in Section 4, and the user-study quotes and reported workflows in Sections 6 and 7. The self-citations to the authors' prior work are contextual (e.g., footnote 2 refers to a supplement; footnote 3 says the formative study was introduced in an earlier paper), but the relevant data and analysis are described in this manuscript, so those citations are not load-bearing. The design probes were built from challenges identified in the formative interviews and were then evaluated against those same design goals (Sections 5 and 6.3). That is a normal design-evaluation loop, not a circular reduction: the user study provides new evidence from participants, including Section 6.5's report that many teams had already independently built similar LLM tooling. A possible confound—that the probes may elicit top-down descriptions rather than reveal pre-existing practice—is a validity or external-evidence concern, not a circularity reduction. No equation, fitted parameter, imported uniqueness theorem, or definitional equivalence can be quoted to show a prediction that is forced by construction. The paper itself acknowledges its snapshot design and single-company scope in Section 7.2, which further supports the absence of a circular derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 2 invented entities

No free parameters; the paper does not fit any numerical model. The scientific contribution rests on assumptions about self-report validity, cross-study comparability, and representativeness of a single company, all acknowledged only partially.

assumptions (3)
  • domain assumption Participant self-reports during interviews and think-aloud sessions accurately reflect real data curation practices.
    Used throughout Sections 4 and 6 to infer actual workflows and adoption patterns from what practitioners said.
  • domain assumption The three different studies (survey, interviews, user study) can be jointly interpreted as an evolution over time.
    Section 7 fuses results from N=84, N=10, and N=12 samples recruited differently into one narrative; no cohort was followed longitudinally.
  • domain assumption Google is sufficiently representative of 'industry' data curation to support the paper's generalized claims.
    Acknowledged as a limitation in Section 7.2, but the Discussion and Conclusions still make broad claims about industry practice.
invented entities (2)
  • silver datasets
    purpose: LLM-generated labeled datasets used alongside expert golden datasets for high-traffic classification.
    Defined in Section 6.7 from participant quotes; no external validation that the term or practice generalizes beyond the interviewed teams.
  • super-golden datasets
    purpose: Small, expert-diverse datasets used as ground truth when comparing LLM and human performance.
    Introduced in Section 6.7; evidence is self-reported and resource costs are anecdotal (for example, weeks to label 500 examples).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Evolution of LLM Adoption in Industry Data Curation Practices." pith.science (2026). https://pith.science/paper/5X3QGFQL

@misc{pith2026241216089,
  author       = {Pith},
  title        = {Pith review of: The Evolution of LLM Adoption in Industry Data Curation Practices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5X3QGFQL}},
  note         = {Machine review of arXiv:2412.16089}
}
read the original abstract

As large language models (LLMs) grow increasingly adept at processing unstructured text data, they offer new opportunities to enhance data curation workflows. This paper explores the evolution of LLM adoption among practitioners at a large technology company, evaluating the impact of LLMs in data curation tasks through participants' perceptions, integration strategies, and reported usage scenarios. Through a series of surveys, interviews, and user studies, we provide a timely snapshot of how organizations are navigating a pivotal moment in LLM evolution. In Q2 2023, we conducted a survey to assess LLM adoption in industry for development tasks (N=84), and facilitated expert interviews to assess evolving data needs (N=10) in Q3 2023. In Q2 2024, we explored practitioners' current and anticipated LLM usage through a user study involving two LLM-based prototypes (N=12). While each study addressed distinct research goals, they revealed a broader narrative about evolving LLM usage in aggregate. We discovered an emerging shift in data understanding from heuristic-first, bottom-up approaches to insights-first, top-down workflows supported by LLMs. Furthermore, to respond to a more complex data landscape, data practitioners now supplement traditional subject-expert-created 'golden datasets' with LLM-generated 'silver' datasets and rigorously validated 'super golden' datasets curated by diverse experts. This research sheds light on the transformative role of LLMs in large-scale analysis of unstructured data and highlights opportunities for further tool development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A synthetic table-to-text pipeline that generates and validates key-value extraction benchmarks, revealing that LLM-generated reports keep numerical facts intact but are poorly machine-extractable.

  2. When Incentives Backfire, Data Stops Being Human

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.

  3. Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline

    cs.HC 2025-01 conditional novelty 6.0 of 10

    Twenty-nine interviews show AI practitioners rely on synthetic data across nearly every pipeline stage while validation remains mostly manual spot-checking.

  4. Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning

    cs.HC 2025-01 conditional novelty 6.0 of 10

    Gensors lets everyday users define personalized visual sensors by decomposing their sensing goal into testable criteria, and a user study shows improved perceived control and understanding over prompt-only authoring.

  5. Language Games as the Pathway to Artificial Superhuman Intelligence

    cs.AI 2025-01 conditional novelty 4.0 of 10

    A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.

Reference graph

Works this paper leans on

8 extracted references · 2 canonical work pages · cited by 5 Pith papers

  1. [8]

    Curran Associates Inc., Red Hook, NY, USA, 46595–46623

    JudgingLLM-as-a-judgewithMT-benchandChatbotArena.In Proceedingsofthe37thInternational Conference on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, 46595–46623. 28

  2. [2012]

    WIREs Data Mining and 19 The Evolution of LLM Adoption in Industry Data Curation Practices Knowledge Discovery 2, 6 (2012), 476–492

    Seeing beyond reading: a survey on visual text analytics. WIREs Data Mining and 19 The Evolution of LLM Adoption in Industry Data Curation Practices Knowledge Discovery 2, 6 (2012), 476–492. https://doi.org/10.1002/widm.1071 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/widm.1071. [Amershi et al.(2015)]Saleema Amershi, Max Chickering, Steven M....

  3. [2018]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18)

    The Story in the Notebook: Exploratory Data Science using a Literate Programming Tool. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–11.https://doi.org/10.1145/3173574.3173748 [Kim et al.(2024)]Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, ...

  4. [2019]

    InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19)

    Towards Effective Foraging by Data Scientists to Find Past Analysis Choices. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–13.https://doi.org/10.1145/3290605.3300322 [Kery et al.(2018)]Mary Beth Kery, Marissa Radensky, Mahima Arya, Bonnie E. John, and Bra...

  5. [2020]

    InAdvances in Neural Information Processing Systems, Vol

    Language Models are Few-Shot Learners. InAdvances in Neural Information Processing Systems, Vol. 33. 1877–1901. https://proceedings.neurips.cc/paper_files/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf [Caliskan et al.(2017)]Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora cont...

  6. [2023]

    Everyone wants to do the model work, not the data work

    Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! https://arxiv.org/abs/2310.03693 _eprint: 2310.03693. [Qian et al.(2024)]CrystalQian, EmilyReif, andMinsukKahng.2024. UnderstandingtheDatasetPractitioners Behind Large Language Model Development. InExtended Abstracts of the 2024 CHI Conference on Human Factors in Com...

  7. [2024]

    We Need Structured Output

    Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–28.https://doi.org/10.1145/3613904.3642830 [Lara and Tiwari(2022)]Harsh Lara and Manoj Tiwari. 2022. Evaluation of Synthetic Dat...

  8. [2504]

    [LMSYS(2024)] LMSYS

    https://doi.org/10.1109/TVCG.2018.2834341 Conference Name: IEEE Transactions on Visualization and Computer Graphics. [LMSYS(2024)] LMSYS. 2024. Chatbot Arena Conversations Dataset. https://huggingface.co/ datasets/lmsys/chatbot_arena_conversations [Longpre et al.(2023)]Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, D...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.