REVIEW 3 major objections 4 minor 5 cited by
The Evolution of LLM Adoption in Industry Data Curation Practices
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Industry data curation is shifting from manual, bottom-up inspection to LLM-first, top-down analysis, as practitioners layer LLM-generated 'silver' datasets under expert 'golden' ones.
desk verdict Useful new dataset taxonomy and workflow framing, but the 'evolution/paradigm shift' claim outruns the cross-sectional evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by two linked mechanisms. The first is the inverted analysis pipeline, a documented shift from bottom-up data aggregation, in which practitioners label individual rows and then count them up, to top-down extraction, in which an LLM is asked directly for themes, summaries, or outlier candidates and the practitioner drills into raw examples only to verify or extract evidence. The second is the multi-tier dataset hierarchy named 'golden,' 'silver,' and 'super-golden': golden sets are expert labels, silver sets are predominantly LLM-generated labels used to complement them, and super-golden sets are small, painstakingly validated collections created by diverse expert teams to serve as a higher-authority ground truth for comparing LLMs with humans. These two mechanisms do the explanatory work: the pipeline inversion explains the reported efficiency gains, and the hierarchy explains how practitioners manage quality and trust as LLM-produced labels enter the pipeline.
What would settle it
Instrument real data-curation pipelines to log whether practitioners, after adopting LLM summarization, still open and inspect raw rows before trusting an LLM-derived insight; if a majority continue to inspect raw data first in most tasks, the claimed inversion from bottom-up to top-down analysis does not hold.
Extended reading notes
Core claim
The authors set out to measure whether practitioners were using LLMs for data curation; within six months the question flipped from whether to how. The paper's core discovery is that LLMs are enabling practitioners to reverse the traditional order of data analysis: instead of building insights bottom-up by labeling and aggregating individual data points, practitioners now generate high-level, insight-first summaries with LLMs and return to the raw data only when a specific claim needs evidence. Alongside this workflow inversion, the authors observe a new tiered dataset economy. Golden datasets, the expert-labeled gold standard, are now being supplemented by 'silver' datasets, whose labels are produced largely by LLMs and used for high-traffic or initialization purposes, and by 'super-golden' datasets, small, rigorously validated collections assembled by diverse expert teams to serve as higher authority for benchmarking LLMs against humans. The authors interpret these changes as a transformation in how practitioners engage with their data, while noting persistent barriers of reliability, cost, unfamiliarity, and content-refusal responses.
Load-bearing premise
The narrative of a fundamental transformation rests on fusing three studies with different samples, methods, and timings into a single trajectory, and on assuming that what practitioners said and did during hour-long sessions with the authors' own prototype tools reflects their real, sustained workflows.
Editorial extensions
If this is right
- If the top-down shift holds, tool builders should prioritize LLM summarization, explanation, and outlier detection inside spreadsheets and notebooks over manual labeling features, since those are the capabilities practitioners reach for first.
- Silver datasets will become a standard layer in production data pipelines, which means validation, error analysis, and bias auditing for LLM-generated labels will be needed before those datasets are used for training or evaluation.
- Super-golden datasets, being small, expensive, and expert-assembled, will grow in importance as reference standards, shifting budgets and timelines for evaluation dataset construction.
- Data quality will be defined collaboratively by safety teams, domain experts, and engineering managers rather than by a single metric, so tools must support consensus-building and inter-team sharing.
- Reliability, latency, cost, and refusal behavior will determine which curation tasks are automated first and which remain human-only in the near term.
Reading between the lines
- If the inversion is real, a testable consequence is that practitioners who adopt LLM-first workflows may gradually lose the tacit, row-level familiarity that made them good at spotting when an insight is wrong; the paper's C3 quote gestures at this risk but does not develop it, so a longitudinal study of error-detection skill after top-down adoption would test it.
- The golden/silver/super-golden hierarchy likely extends beyond text to images and audio, where the same pattern of LLM-generated labels under expert validation could appear; the authors list multimodal data only as future work.
- The spreadsheet probe's broad appeal across technical and non-technical roles suggests that prompt-in-cell interfaces may become a default collaboration surface for data work, a trajectory the paper does not fully commit to because its probes were built by the authors.
- A quantitative extension would measure the share of silver versus golden labels in production datasets over time; if silver is merely a stopgap rather than a durable tier, its share should plateau rather than grow.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents three studies conducted at a single large technology company to trace the evolution of LLM adoption in data curation: an exploratory survey of 84 employees (Q2 2023), interviews with 10 data practitioners and tool developers (late 2023/early 2024), and a user study with 12 practitioners using two author-built LLM-based design probes embedded in spreadsheets and notebooks (Q3 2024). The paper argues that practitioners moved from heuristic-first, bottom-up data analysis to insights-first, top-down workflows, that data quality is becoming a collaborative and subjective construct, and that multi-tiered dataset hierarchies ('golden,' 'silver,' 'super-golden') are emerging. It also reports perceived efficiency gains, barriers to adoption, and implications for future tools.
Significance. The paper offers a rare insider snapshot of LLM adoption in an industrial data-curation setting, with concrete quotes, role-specific participant tables, and a transparent description of the coding process; the design probes themselves are a useful vehicle for eliciting practitioner reactions. If the empirical claims are read as qualitative, contextual observations, the paper provides timely evidence of how practitioners perceive LLM capabilities. However, the headline conclusions—'fundamental transformation' and 'paradigm shift'—go beyond what the study design can establish. The three studies are cross-sectional, non-overlapping in participants and instruments, and the 2024 session is a reaction-to-probe study rather than an observational or longitudinal study. The paper's contribution is therefore strongest as a design-oriented snapshot of perceptions and anticipated workflows, and weaker as a demonstration of actual behavioral change.
major comments (3)
- [Section 8; see also Sections 6.2 and 7.1.2] The central conclusion that LLMs are causing a 'fundamental transformation' and 'paradigm shift' (Section 8) is not supported by the study design. The three empirical studies are cross-sectional snapshots of different, non-overlapping samples with different instruments; no participant was measured at two time points. The top-down workflow evidence in Section 7.1.2 comes from the 2024 user study, where participants were given the authors' LLM-centric spreadsheet and notebook probes and asked to try them (Section 6.2), so the observed behavior may reflect the probes' affordances rather than a pre-existing change in practice. The paper should either reframe the conclusion as evidence of participants' receptivity to, and reported use of, LLM-based workflows, or provide observational or baseline evidence showing the shift outside the probe context. The limitation paragraph in Section 7.2 acknowledges the snapshot character of the data, but Section 8 does not carry those caveats.
- [Sections 6.3.1 and 6.7] The claims of 'transformative efficiency gains' and 'emerging dataset hierarchies' are stated more strongly than the evidence allows. The efficiency result is based on participants' estimates and hypothetical projections (e.g., C2's 45 minutes to code 75 responses; T3's '500 data points per hour') rather than on measured before/after productivity. The silver and super-golden dataset hierarchy is supported mainly by two quotes from A1 and T4 within an N=12 study, with no indication of prevalence. These findings should be labeled as perceived and anticipated benefits and as emergent themes from a small sample, and the language in Section 8 ('rapidly growing reliance') should be correspondingly moderated.
- [Sections 3, 4, 6, and 7.1] The three studies measure different constructs, so the narrative of an 'evolution' is an analytic reconstruction rather than an observed trajectory. Section 3 surveys general development-task LLM adoption, Section 4 interviews practitioners about data needs and challenges, and Section 6 probes reactions to specific prototypes; the synthesis in Section 7.1 and the conclusion in Section 8 weave these together as a temporal progression. Since the samples and constructs do not line up, the paper should explicitly present the composite as an interpretive narrative across independent studies, and should identify which specific claims are supported by which study, rather than implying a continuous longitudinal measure.
minor comments (4)
- [Sections 4.1, 4.3.2, 4.4, 6.8, and 7.1.2] Participant identifiers are used inconsistently: Section 4.4 refers to a sample of 12 participants although Section 4.1 and Table 1 report N=10; Section 4.3.2 quotes 'T4' where Section 4 participants are coded U/D; Section 6.8 quotes 'S1' though Table 3 uses T/A/C codes; and Section 7.1.2 refers to 'R2 and R3' that never appear in any participant table.
- [Section 4.3.2] There is a typo in Section 4.3.2: 'development of language language models' should read 'large language models.'
- [Abstract and Sections 4, 5, and 8] The timeline is misaligned: the abstract dates the expert interviews to Q3 2023, but Section 4 states they were conducted between November 2023 and January 2024; the abstract dates the user study to Q2 2024, while Section 5 says Q3 2024; Section 8's 'just six months after' should be reconciled with these dates.
- [Section 5.2] Section 5.2 contains a sentence fragment: 'Since Python notebooks offer greater flexibility than spreadsheets. We provide two additional features:' should be joined into one sentence or rephrased.
Circularity Check
No significant circularity was found; the paper's findings rest on quoted practitioner reports and reproduced study data, and the design-probe goal evaluation is not a reduction.
full rationale
This paper is a qualitative multi-study report rather than a formal derivation, and no load-bearing step reduces by construction to its inputs. The central 'evolution' and 'paradigm shift' claims are supported by distinct cross-sectional studies whose evidence is reproduced in the paper: the survey statistics in Section 3, the interview themes and quotes in Section 4, and the user-study quotes and reported workflows in Sections 6 and 7. The self-citations to the authors' prior work are contextual (e.g., footnote 2 refers to a supplement; footnote 3 says the formative study was introduced in an earlier paper), but the relevant data and analysis are described in this manuscript, so those citations are not load-bearing. The design probes were built from challenges identified in the formative interviews and were then evaluated against those same design goals (Sections 5 and 6.3). That is a normal design-evaluation loop, not a circular reduction: the user study provides new evidence from participants, including Section 6.5's report that many teams had already independently built similar LLM tooling. A possible confound—that the probes may elicit top-down descriptions rather than reveal pre-existing practice—is a validity or external-evidence concern, not a circularity reduction. No equation, fitted parameter, imported uniqueness theorem, or definitional equivalence can be quoted to show a prediction that is forced by construction. The paper itself acknowledges its snapshot design and single-company scope in Section 7.2, which further supports the absence of a circular derivation chain.
Assumptions & free parameters
assumptions (3)
- domain assumption Participant self-reports during interviews and think-aloud sessions accurately reflect real data curation practices.
- domain assumption The three different studies (survey, interviews, user study) can be jointly interpreted as an evolution over time.
- domain assumption Google is sufficiently representative of 'industry' data curation to support the paper's generalized claims.
invented entities (2)
-
silver datasets
-
super-golden datasets
Cite this review
Pith. "Pith review of The Evolution of LLM Adoption in Industry Data Curation Practices." pith.science (2026). https://pith.science/paper/5X3QGFQL
@misc{pith2026241216089,
author = {Pith},
title = {Pith review of: The Evolution of LLM Adoption in Industry Data Curation Practices},
year = {2026},
howpublished = {\url{https://pith.science/paper/5X3QGFQL}},
note = {Machine review of arXiv:2412.16089}
}
read the original abstract
As large language models (LLMs) grow increasingly adept at processing unstructured text data, they offer new opportunities to enhance data curation workflows. This paper explores the evolution of LLM adoption among practitioners at a large technology company, evaluating the impact of LLMs in data curation tasks through participants' perceptions, integration strategies, and reported usage scenarios. Through a series of surveys, interviews, and user studies, we provide a timely snapshot of how organizations are navigating a pivotal moment in LLM evolution. In Q2 2023, we conducted a survey to assess LLM adoption in industry for development tasks (N=84), and facilitated expert interviews to assess evolving data needs (N=10) in Q3 2023. In Q2 2024, we explored practitioners' current and anticipated LLM usage through a user study involving two LLM-based prototypes (N=12). While each study addressed distinct research goals, they revealed a broader narrative about evolving LLM usage in aggregate. We discovered an emerging shift in data understanding from heuristic-first, bottom-up approaches to insights-first, top-down workflows supported by LLMs. Furthermore, to respond to a more complex data landscape, data practitioners now supplement traditional subject-expert-created 'golden datasets' with LLM-generated 'silver' datasets and rigorously validated 'super golden' datasets curated by diverse experts. This research sheds light on the transformative role of LLMs in large-scale analysis of unstructured data and highlights opportunities for further tool development.
Forward citations
Cited by 5 Pith papers
-
StructText: A Synthetic Table-to-Text Approach for Benchmark Generation with Multi-Dimensional Evaluation
A synthetic table-to-text pipeline that generates and validates key-value extraction benchmarks, revealing that LLM-generated reports keep numerical facts intact but are poorly machine-extractable.
-
When Incentives Backfire, Data Stops Being Human
Incentive-driven crowdwork erodes intrinsic motivation and data quality, so data collection should be redesigned around intrinsic motivation, with games as a promising template.
-
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
Twenty-nine interviews show AI practitioners rely on synthetic data across nearly every pipeline stage while validation remains mostly manual spot-checking.
-
Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
Gensors lets everyday users define personalized visual sensors by decomposing their sensing goal into testable criteria, and a user study shows improved perceived control and understanding over prompt-only authoring.
-
Language Games as the Pathway to Artificial Superhuman Intelligence
A position paper arguing that open-ended language games with fluid roles, varied rewards, and evolving rules can drive expanded data reproduction and thus a path to artificial superhuman intelligence.
Reference graph
Works this paper leans on
-
[8]
Curran Associates Inc., Red Hook, NY, USA, 46595–46623
JudgingLLM-as-a-judgewithMT-benchandChatbotArena.In Proceedingsofthe37thInternational Conference on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA, 46595–46623. 28
-
[2012]
Seeing beyond reading: a survey on visual text analytics. WIREs Data Mining and 19 The Evolution of LLM Adoption in Industry Data Curation Practices Knowledge Discovery 2, 6 (2012), 476–492. https://doi.org/10.1002/widm.1071 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/widm.1071. [Amershi et al.(2015)]Saleema Amershi, Max Chickering, Steven M....
arXiv 2012
-
[2018]
In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18)
The Story in the Notebook: Exploratory Data Science using a Literate Programming Tool. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI ’18). Association for Computing Machinery, New York, NY, USA, 1–11.https://doi.org/10.1145/3173574.3173748 [Kim et al.(2024)]Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, ...
arXiv 2024
-
[2019]
InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19)
Towards Effective Foraging by Data Scientists to Find Past Analysis Choices. InProceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19). Association for Computing Machinery, New York, NY, USA, 1–13.https://doi.org/10.1145/3290605.3300322 [Kery et al.(2018)]Mary Beth Kery, Marissa Radensky, Mahima Arya, Bonnie E. John, and Bra...
arXiv 2018
-
[2020]
InAdvances in Neural Information Processing Systems, Vol
Language Models are Few-Shot Learners. InAdvances in Neural Information Processing Systems, Vol. 33. 1877–1901. https://proceedings.neurips.cc/paper_files/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf [Caliskan et al.(2017)]Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora cont...
-
[2023]
Everyone wants to do the model work, not the data work
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To! https://arxiv.org/abs/2310.03693 _eprint: 2310.03693. [Qian et al.(2024)]CrystalQian, EmilyReif, andMinsukKahng.2024. UnderstandingtheDatasetPractitioners Behind Large Language Model Development. InExtended Abstracts of the 2024 CHI Conference on Human Factors in Com...
arXiv 2024
-
[2024]
Concept Induction: Analyzing Unstructured Text with High-Level Concepts Using LLooM. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). Association for Computing Machinery, New York, NY, USA, 1–28.https://doi.org/10.1145/3613904.3642830 [Lara and Tiwari(2022)]Harsh Lara and Manoj Tiwari. 2022. Evaluation of Synthetic Dat...
arXiv 2022
-
[2504]
https://doi.org/10.1109/TVCG.2018.2834341 Conference Name: IEEE Transactions on Visualization and Computer Graphics. [LMSYS(2024)] LMSYS. 2024. Chatbot Arena Conversations Dataset. https://huggingface.co/ datasets/lmsys/chatbot_arena_conversations [Longpre et al.(2023)]Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, D...
arXiv 2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.