{"id":"48da97de-4696-4bec-876d-53ec80797d71","arxiv_id":"2412.16089","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Over 2023-2024, Google data practitioners moved from heuristics-based, bottom-up data exploration toward LLM-first, top-down analysis, while supplementing golden datasets with LLM-generated silver and expert super-golden datasets.","lead":"This paper surveys and interviews data practitioners at Google over 2023-2024 to trace how they adopt large language models in data curation. It describes a shift from manual, bottom-up data inspection to LLM-first top-down analysis, and new tiers of datasets including LLM-generated 'silver' sets and expert 'super-golden' sets.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central 'evolution/paradigm shift' claim rests on non-comparable cross-sectional snapshots, not a longitudinal trajectory; the 2024 top-down workflows may be elicited by the author-built design probes rather than observed in practice.","rationale":"The reader's weakest assumption is essentially the same as the load-bearing concern identified here: the three studies use different samples, different methods, and different time points, and the design-probe session may reveal reactions to the probes rather than actual practice. This concern targets the central 'evolution' and 'fundamental transformation' claims in Section 8, not the more modest and well-supported observations about silver and super-golden datasets or the qualitative snapshots of practitioner challenges. The paper's own Section 7.2 acknowledges single-company scope and small samples, but it does not address the deeper internal-validity issue that the 'before' and 'after' populations are non-overlapping and that the 'after' measurement is confounded with exposure to the authors' prototypes. Because the paper is framed as a snapshot with honest limitations, a conditional accept remains appropriate: the descriptive findings about emerging dataset tiers and LLM-assisted workflows are plausible and grounded in quotes, but the strong temporal narrative should be revised or explicitly labeled as a hypothesis. No change to the reader's verdict is needed; the concern reinforces the conditionality rather than overturning the paper.","tokens_in":27394,"tokens_out":3818,"duration_ms":33984,"concrete_test":"Conduct a matched longitudinal follow-up: re-interview the 10 Q3 2023 expert interviewees (Section 4.1) in Q3/Q4 2024 using the same semi-structured protocol but without introducing any design probes, coding transcripts for (a) self-reported LLM use in data curation and (b) descriptions of bottom-up versus top-down analysis. If the 2024 interviews do not show a clear increase in LLM use or top-down workflow descriptions relative to the 2023 transcripts (Table 2, Section 4.3), the evolution claim is unsupported. As a cheaper proximate check, re-analyze the existing 2024 user-study transcripts to separate statements about current practice made before participants touched the probes from statements made after probe exposure; if top-down descriptions appear only after probe usage, the conclusion should be recast as probe-elicited potential, not observed evolution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 8: 'rapidly growing reliance... fundamental transformation') depends on a before/after trajectory assembled from Section 3 (Q2 2023 survey, N=84), Section 4 (Q3 2023 interviews, N=10), and Section 6 (2024 user study, N=12). These are cross-sectional snapshots of different, non-overlapping samples with different instruments; no participant is measured twice. The 2024 session was not an observational study of practice. Participants were handed author-built spreadsheet and notebook probes that centrally feature LLM prompting, summarization, and classification (Sections 5.1-5.2, 6.2), and the main evidence for the top-down shift (Section 7.1.2) is participants' reported use of LLM summaries during a session in which summative LLM queries were explicitly demonstrated. Section 6.5 reports that some teams had independently built similar LLM tooling, which supports real adoption, but the paper provides no baseline for these 12 participants and no measure of whether their workflows were heuristic-first before LLMs. The 'evolution' claim therefore rests on an implicit confound: the design probes may elicit top-down descriptions rather than reveal a pre-existing shift. The conclusion's 'paradigm shift' language outruns the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents three studies conducted at a single large technology company to trace the evolution of LLM adoption in data curation: an exploratory survey of 84 employees (Q2 2023), interviews with 10 data practitioners and tool developers (late 2023/early 2024), and a user study with 12 practitioners using two author-built LLM-based design probes embedded in spreadsheets and notebooks (Q3 2024). The paper argues that practitioners moved from heuristic-first, bottom-up data analysis to insights-first, top-down workflows, that data quality is becoming a collaborative and subjective construct, and that multi-tiered dataset hierarchies ('golden,' 'silver,' 'super-golden') are emerging. It also reports perceived efficiency gains, barriers to adoption, and implications for future tools.","tokens_in":27763,"tokens_out":7934,"duration_ms":67815,"significance":"The paper offers a rare insider snapshot of LLM adoption in an industrial data-curation setting, with concrete quotes, role-specific participant tables, and a transparent description of the coding process; the design probes themselves are a useful vehicle for eliciting practitioner reactions. If the empirical claims are read as qualitative, contextual observations, the paper provides timely evidence of how practitioners perceive LLM capabilities. However, the headline conclusions—'fundamental transformation' and 'paradigm shift'—go beyond what the study design can establish. The three studies are cross-sectional, non-overlapping in participants and instruments, and the 2024 session is a reaction-to-probe study rather than an observational or longitudinal study. The paper's contribution is therefore strongest as a design-oriented snapshot of perceptions and anticipated workflows, and weaker as a demonstration of actual behavioral change.","major_comments":[{"comment":"The central conclusion that LLMs are causing a 'fundamental transformation' and 'paradigm shift' (Section 8) is not supported by the study design. The three empirical studies are cross-sectional snapshots of different, non-overlapping samples with different instruments; no participant was measured at two time points. The top-down workflow evidence in Section 7.1.2 comes from the 2024 user study, where participants were given the authors' LLM-centric spreadsheet and notebook probes and asked to try them (Section 6.2), so the observed behavior may reflect the probes' affordances rather than a pre-existing change in practice. The paper should either reframe the conclusion as evidence of participants' receptivity to, and reported use of, LLM-based workflows, or provide observational or baseline evidence showing the shift outside the probe context. The limitation paragraph in Section 7.2 acknowledges the snapshot character of the data, but Section 8 does not carry those caveats.","section":"Section 8; see also Sections 6.2 and 7.1.2"},{"comment":"The claims of 'transformative efficiency gains' and 'emerging dataset hierarchies' are stated more strongly than the evidence allows. The efficiency result is based on participants' estimates and hypothetical projections (e.g., C2's 45 minutes to code 75 responses; T3's '500 data points per hour') rather than on measured before/after productivity. The silver and super-golden dataset hierarchy is supported mainly by two quotes from A1 and T4 within an N=12 study, with no indication of prevalence. These findings should be labeled as perceived and anticipated benefits and as emergent themes from a small sample, and the language in Section 8 ('rapidly growing reliance') should be correspondingly moderated.","section":"Sections 6.3.1 and 6.7"},{"comment":"The three studies measure different constructs, so the narrative of an 'evolution' is an analytic reconstruction rather than an observed trajectory. Section 3 surveys general development-task LLM adoption, Section 4 interviews practitioners about data needs and challenges, and Section 6 probes reactions to specific prototypes; the synthesis in Section 7.1 and the conclusion in Section 8 weave these together as a temporal progression. Since the samples and constructs do not line up, the paper should explicitly present the composite as an interpretive narrative across independent studies, and should identify which specific claims are supported by which study, rather than implying a continuous longitudinal measure.","section":"Sections 3, 4, 6, and 7.1"}],"minor_comments":[{"comment":"Participant identifiers are used inconsistently: Section 4.4 refers to a sample of 12 participants although Section 4.1 and Table 1 report N=10; Section 4.3.2 quotes 'T4' where Section 4 participants are coded U/D; Section 6.8 quotes 'S1' though Table 3 uses T/A/C codes; and Section 7.1.2 refers to 'R2 and R3' that never appear in any participant table.","section":"Sections 4.1, 4.3.2, 4.4, 6.8, and 7.1.2"},{"comment":"There is a typo in Section 4.3.2: 'development of language language models' should read 'large language models.'","section":"Section 4.3.2"},{"comment":"The timeline is misaligned: the abstract dates the expert interviews to Q3 2023, but Section 4 states they were conducted between November 2023 and January 2024; the abstract dates the user study to Q2 2024, while Section 5 says Q3 2024; Section 8's 'just six months after' should be reconciled with these dates.","section":"Abstract and Sections 4, 5, and 8"},{"comment":"Section 5.2 contains a sentence fragment: 'Since Python notebooks offer greater flexibility than spreadsheets. We provide two additional features:' should be joined into one sentence or rephrased.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript synthesizes material from the authors' prior publications (Qian et al. 2024; Qian and Wexler 2024) with a new user study. The authors should clearly delineate which findings are new to this submission; as written, the 'evolution' framing may be inadvertently created by concatenating three separately motivated studies rather than by a designed longitudinal comparison. If the editor values the design-probe findings on their own, the paper would be stronger with a more modest framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the arXiv paper. The takeaway: the silver/super-golden dataset taxonomy and the bottom-up-to-top-down framing are the real contributions, and they are grounded in quotes from a small but well-chosen set of practitioners. The paper is an honest qualitative snapshot, not a longitudinal study, and the conclusion's 'paradigm shift' language overstates what the evidence can support.\n\nWhat's actually new: the idea that LLMs are creating tiered dataset hierarchies—LLM-generated 'silver' labels complementing expert golden sets, and 'super-golden' sets for benchmarking—is a useful vocabulary for tool builders. The shift from proxy heuristics to direct LLM queries is also a concrete observation with supporting quotes. The paper does a decent job of documenting that some teams already built similar tooling independently (Section 6.5), which strengthens the adoption claim.\n\nThe soft spots are real but not fatal. The three studies are cross-sectional snapshots of different samples with different instruments; no participant is measured twice, so 'evolution' is an inference, not an observation. The 2024 user study gave participants probes that centrally featured LLM prompting, so the top-down descriptions may reflect the probe rather than pre-existing practice. There's no baseline for the 12 participants, and the paper itself acknowledges the single-company context and small sample. The conclusion's 'fundamental transformation' and 'paradigm shift' language should be toned down to 'emerging shift' or 'early evidence.'\n\nThe citations look appropriate; the self-citations are to related prior work and are not a problem. The paper doesn't ship code or data—raw interview data isn't available—but for a qualitative study that's not unusual.\n\nWho this is for: HCI researchers and tool builders working on LLM-assisted data analysis, especially those designing spreadsheet/notebook integrations. It deserves a serious referee: the taxonomy is novel and likely influential, and the paper is clearly written. It needs revision to align claims with evidence, but it's not a desk reject.\n\nRecommendation: send to peer review, with the expectation of major revisions to reframe the 'evolution' narrative as a multi-snapshot exploratory study and to add a limitations section that explicitly addresses the probe-elicitation confound. I'd also ask the authors to provide the raw anonymized quotes or an appendix with the coding scheme.","headline":"Useful new dataset taxonomy and workflow framing, but the 'evolution/paradigm shift' claim outruns the cross-sectional evidence.","tokens_in":28179,"tokens_out":2862,"would_cite":true,"duration_ms":23368,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Industry data curation is shifting from manual, bottom-up inspection to LLM-first, top-down analysis, as practitioners layer LLM-generated 'silver' datasets under expert 'golden' ones.","keywords":["Data curation","Large language models","Data quality","Data analysis workflows","Exploratory data analysis","Text analysis","Data practitioners"],"falsifier":"Instrument real data-curation pipelines to log whether practitioners, after adopting LLM summarization, still open and inspect raw rows before trusting an LLM-derived insight; if a majority continue to inspect raw data first in most tasks, the claimed inversion from bottom-up to top-down analysis does not hold.","tokens_in":27195,"feed_emoji":"🤖","tokens_out":8475,"duration_ms":64684,"temperature":0.7,"pith_summary":"The paper tracks how data practitioners at a large technology company adopted LLMs for curating unstructured text data across roughly eighteen months, from a spring 2023 survey through late-2024 interviews and a design-probe user study. Its central claim is that LLM adoption is transforming data understanding from a heuristic-first, bottom-up process, where people label and aggregate individual rows, into an insights-first, top-down process, where people ask an LLM for high-level summaries and drill into specifics only when needed. The paper also claims that this shift is producing a multi-tier dataset hierarchy: expert-labeled 'golden' datasets, LLM-generated 'silver' datasets, and 'super-golden' datasets validated by diverse expert teams for benchmarking against human performance. If true, the burden of data-quality evaluation moves from individual manual inspection toward shared frameworks for validating LLM-produced labels and summaries.","feed_headline":"LLMs flip data curation from row-level checks to insights first","feed_subtitle":"Surveys, interviews, and a design probe show teams layering LLM-labeled silver data under expert golden sets.","key_machinery":"The argument is carried by two linked mechanisms. The first is the inverted analysis pipeline, a documented shift from bottom-up data aggregation, in which practitioners label individual rows and then count them up, to top-down extraction, in which an LLM is asked directly for themes, summaries, or outlier candidates and the practitioner drills into raw examples only to verify or extract evidence. The second is the multi-tier dataset hierarchy named 'golden,' 'silver,' and 'super-golden': golden sets are expert labels, silver sets are predominantly LLM-generated labels used to complement them, and super-golden sets are small, painstakingly validated collections created by diverse expert teams to serve as a higher-authority ground truth for comparing LLMs with humans. These two mechanisms do the explanatory work: the pipeline inversion explains the reported efficiency gains, and the hierarchy explains how practitioners manage quality and trust as LLM-produced labels enter the pipeline.","core_discovery":"The authors set out to measure whether practitioners were using LLMs for data curation; within six months the question flipped from whether to how. The paper's core discovery is that LLMs are enabling practitioners to reverse the traditional order of data analysis: instead of building insights bottom-up by labeling and aggregating individual data points, practitioners now generate high-level, insight-first summaries with LLMs and return to the raw data only when a specific claim needs evidence. Alongside this workflow inversion, the authors observe a new tiered dataset economy. Golden datasets, the expert-labeled gold standard, are now being supplemented by 'silver' datasets, whose labels are produced largely by LLMs and used for high-traffic or initialization purposes, and by 'super-golden' datasets, small, rigorously validated collections assembled by diverse expert teams to serve as higher authority for benchmarking LLMs against humans. The authors interpret these changes as a transformation in how practitioners engage with their data, while noting persistent barriers of reliability, cost, unfamiliarity, and content-refusal responses.","pith_inferences":["If the inversion is real, a testable consequence is that practitioners who adopt LLM-first workflows may gradually lose the tacit, row-level familiarity that made them good at spotting when an insight is wrong; the paper's C3 quote gestures at this risk but does not develop it, so a longitudinal study of error-detection skill after top-down adoption would test it.","The golden/silver/super-golden hierarchy likely extends beyond text to images and audio, where the same pattern of LLM-generated labels under expert validation could appear; the authors list multimodal data only as future work.","The spreadsheet probe's broad appeal across technical and non-technical roles suggests that prompt-in-cell interfaces may become a default collaboration surface for data work, a trajectory the paper does not fully commit to because its probes were built by the authors.","A quantitative extension would measure the share of silver versus golden labels in production datasets over time; if silver is merely a stopgap rather than a durable tier, its share should plateau rather than grow."],"forward_implications":["If the top-down shift holds, tool builders should prioritize LLM summarization, explanation, and outlier detection inside spreadsheets and notebooks over manual labeling features, since those are the capabilities practitioners reach for first.","Silver datasets will become a standard layer in production data pipelines, which means validation, error analysis, and bias auditing for LLM-generated labels will be needed before those datasets are used for training or evaluation.","Super-golden datasets, being small, expensive, and expert-assembled, will grow in importance as reference standards, shifting budgets and timelines for evaluation dataset construction.","Data quality will be defined collaboratively by safety teams, domain experts, and engineering managers rather than by a single metric, so tools must support consensus-building and inter-team sharing.","Reliability, latency, cost, and refusal behavior will determine which curation tasks are automated first and which remain human-only in the near term."],"supporting_citations":[{"why":"Supplies the prior interview-study taxonomy of enterprise data analysis that the paper uses as the baseline for comparing practitioners' tools, tasks, and challenges.","marker":"[Kandel et al.(2012a)]"},{"why":"The internal survey this paper extends, providing the Q2 2023 adoption and trust baseline for the evolution narrative.","marker":"[Qian and Wexler(2024)]"},{"why":"Companion paper reporting the expert interviews; its methods and findings are the formative study that motivates the design probes.","marker":"[Qian et al.(2024)]"},{"why":"External evidence that LLM usage was rising before the Q3 2024 user study, justifying the design-probe follow-up.","marker":"[Liao et al.(2024)]"},{"why":"Provides the SPACE productivity framework (accuracy, efficiency, satisfaction) used to evaluate the two design probes in the user study.","marker":"[Forsgren et al.(2021)]"},{"why":"Cited basis for the notion of 'silver' datasets as LLM-generated synthetic labels complementing human golden labels.","marker":"[Liu et al.(2023a)]"},{"why":"Supports the paper's claim that small, high-quality datasets can drive strong model performance, motivating the emphasis on super-golden datasets.","marker":"[Abdin et al.(2024)]"},{"why":"Establishes the data-cascades argument that data quality is critical for reliable AI systems, anchoring the paper's focus on quality as the central challenge.","marker":"[Sambasivan et al.(2021)]"}],"fun_headline_variants":["LLMs flip data curation: insights first, raw rows on demand","Golden, silver, super-golden: LLMs build tiered data sets","Curators invert workflow: LLMs summarize first, verify later","How LLMs turned data curation from bottom-up to top-down","Industry shifts to insight-first data work as LLMs mature"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The narrative of a fundamental transformation rests on fusing three studies with different samples, methods, and timings into a single trajectory, and on assuming that what practitioners said and did during hour-long sessions with the authors' own prototype tools reflects their real, sustained workflows.","fun_headline_variants_meta":{"raw":{"variants":["LLMs flip data curation: insights first, raw rows on demand","Golden, silver, super-golden: LLMs build tiered data sets","Curators invert workflow: LLMs summarize first, verify later","How LLMs turned data curation from bottom-up to top-down","Industry shifts to insight-first data work as LLMs mature"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000283,"raw_usage":{"total_tokens":1707,"prompt_tokens":1019,"completion_tokens":688,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":598}},"tokens_in":635,"tokens_out":688,"duration_ms":6006,"temperature":1.0,"reasoning_tokens":598,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:46:56.353565+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Instrument real data-curation pipelines to log whether practitioners, after adopting LLM summarization, still open and inspect raw rows before trusting an LLM-derived insight; if a majority continue to inspect raw data first in most tasks, the claimed inversion from bottom-up to top-down analysis does not hold.","supporting_citations":[],"review_version":1}