{"id":"901105bc-53a6-4c07-afa7-1ab84fb945fe","arxiv_id":"2505.18458","paper_version":3,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive survey of the bidirectional links between LLMs and data management, organized as DATA4LLM and LLM4DATA with a new 'IaaS' data-quality framework.","lead":"This paper surveys how large language models and data management feed each other: data pipelines for LLM training and retrieval, and LLMs used to clean, analyze, and tune data systems. It proposes the 'IaaS' lens (inclusiveness, abundance, articulation, sanitization) for judging LLM dataset quality.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The IaaS lens is asserted, not validated: Section 2.1 derives a four-dimensional quality taxonomy from 'over 400 papers' without a coding protocol or any reliability/predictive test, and the survey's comprehensiveness claim rests on this unverified corpus.","rationale":"The paper is a large survey rather than a formal derivation, so its correctness risk is about representativeness and taxonomy validity rather than a provable theorem. The reader's weakest assumption (unverified 400+ paper corpus) is real and directly affects the comprehensiveness claim. My stress-test identifies a closely related but distinct load-bearing concern: even if the corpus were perfectly representative, the IaaS lens is introduced without a coding protocol, reliability check, or predictive validation. The paper says \"Based on our investigation of over 400 papers, we introduce the IaaS concept,\" which is an empirical derivation claim, but no derivation evidence is presented. This matters because the survey's organization, tables, and claims of principled assessment all depend on IaaS being a useful partition of the literature. The concrete audit I propose would settle both concerns at once: it tests whether an independent annotator can reliably reproduce the taxonomy labels and whether fresh relevant papers fit without forcing. If the audit passes, the conditional verdict can be upgraded; if it fails, the central contribution is weakened. I do not see a more load-bearing issue: the survey's factual errors are mostly ordinal numbering mistakes and typos, and its novelty claims are not internally inconsistent with the cited prior surveys. The GitHub repository is a credit to the authors and provides partial reproducibility, but it does not by itself establish a selection protocol or inter-annotator agreement. Therefore I recommend keeping the reader's CONDITIONAL verdict unchanged: the survey is plausible and useful, but the IaaS principle and corpus representativeness need an explicit validation step before the strongest claims should be taken as established.","tokens_in":47043,"tokens_out":3567,"duration_ms":33402,"concrete_test":"Perform an annotation audit on 100 papers: 50 randomly drawn from the paper's GitHub list and 50 from an independent systematic search (DBLP/arXiv, 2023-2025, queries including 'LLM data', 'data management', 'data-centric LLM', plus backward/forward citation snowballing) with eligibility criteria fixed in advance. Two annotators blind to the survey's taxonomy independently assign each paper to (a) DATA4LLM vs. LLM4DATA vs. both/neither, and (b) the primary IaaS dimension(s) the paper addresses. Measure coverage (fraction of fresh-systematic papers that fit the taxonomy without forcing) and inter-annotator agreement (Cohen's kappa on taxonomy and IaaS labels). If coverage is below 90% or kappa is below 0.6, the comprehensiveness/IaaS claims are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central value is its claimed contribution: a lifecycle-based taxonomy organized by the new \"IaaS\" concept, plus the claim of being the first comprehensive bidirectional review. The load-bearing condition is that IaaS is a genuine organizing structure rather than a post-hoc acronym. Section 2.1 states: \"Based on our investigation of over 400 papers, we introduce the IaaS concept for evaluating the quality of LLM datasets,\" but the surrounding text gives no derivation, no definition of the space from which the four dimensions were selected, no independence or exhaustiveness argument, and no inter-annotator reliability. The subsequent organization (Figure 4, Tables 1-7) is built on this lens, so if the dimensions overlap (e.g., \"articulation\" vs. filtering; \"abundance\" vs. data mixing) or omit important aspects (e.g., temporal freshness, licensing constraints), the taxonomy's coverage is not established. The second load-bearing element is the comprehensiveness claim itself: \"over 400 papers\" is supported only by a GitHub link, with no search strategy, inclusion/exclusion criteria, deduplication procedure, or validation against a second sample. The reader's weakest assumption correctly identifies this representativeness risk. Together, these are not matters of opinion about the field's direction; they are unsupported empirical assertions about how the literature clusters, and the paper's own text does not provide the evidence needed to test them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews the bidirectional relationship between large language models and data management, organizing the literature into DATA4LLM (data processing, storage, and serving across pre-training, continual pre-training, SFT, RL, RAG, agents, and evaluation) and LLM4DATA (data manipulation, data analysis, and system optimization). The authors introduce an \"IaaS\" concept with four quality dimensions (inclusiveness, abundance, articulation, sanitization), claim to be the first lifecycle-based bidirectional survey, and base the taxonomy on an investigation of over 400 papers. The paper covers a wide range of techniques and systems, with a companion GitHub repository of cited works.","tokens_in":47308,"tokens_out":2775,"duration_ms":25352,"significance":"If the IaaS lens and the claimed comprehensiveness hold, this survey would provide a useful shared vocabulary and map for the rapidly growing LLM-data intersection, spanning both training-side data management and LLM-driven data management. The breadth is substantial, and the companion GitHub repository is a valuable community resource. However, the paper's central contributions are empirical and organizational rather than formal: there are no machine-checked proofs or parameter-free derivations, and the value rests on whether the taxonomy is a genuine, non-redundant decomposition of the literature and whether the 400+ paper sample is representative. These conditions are asserted rather than demonstrated, so the significance is conditional on methodological support that the manuscript does not currently provide.","major_comments":[{"comment":"The derivation of the IaaS concept is not supported. The text says \"Based on our investigation of over 400 papers, we introduce the IaaS concept,\" but it does not describe how the four dimensions were selected, what the candidate space was, how overlap among dimensions was resolved (e.g., articulation vs. filtering, abundance vs. data mixing), or how exhaustiveness was assessed. Since Figures 1 and 4 and Tables 1-7 are organized around IaaS, this is load-bearing. The authors should add a methodology paragraph describing the coding procedure, or at minimum a falsifiable consistency test, such as reporting inter-annotator agreement on a random sample of cited papers classified by the four dimensions.","section":"Section 2.1"},{"comment":"The comprehensiveness claim is not verifiable from the manuscript. The statement \"over 400 papers\" is supported only by a GitHub link; there is no search strategy, inclusion/exclusion criteria, deduplication procedure, or validation against a second literature sample. Because the survey distinguishes itself from prior surveys on the basis of comprehensiveness and lifecycle coverage, this is a load-bearing methodological gap. I recommend adding a short survey-methodology subsection or, failing that, explicitly tempering the \"first comprehensive\" claim to \"covering the selected 400+ papers.\"","section":"Section 1.3 and Section 2.1"},{"comment":"There is an internal inconsistency in the lifecycle taxonomy. Table 1 marks RAG-stage filtering as \"×\" (no relevant work reported), yet Section 2.5.1 explicitly discusses RAG knowledge filtering and re-ranking with citations [280], [114], [87], [128], [12], [318], [47]. This contradiction undermines the reliability of the table as a summary of the field. The table should either correct the RAG filtering entry or explain what kind of filtering \"×\" refers to, and it should include a legend distinguishing \"×\" from \"N/A.\"","section":"Table 1 and Section 2.5.1"},{"comment":"The subsection numbering is erroneous: the text labels pipeline subsections as 2.2.7.1, 2.2.7.2, and 2.2.7.3 when they appear under Section 2.3.7. This is an editorial defect that makes the internal references in the introduction (e.g., Section 2.3.7) and the table of contents inconsistent, and it should be fixed before publication.","section":"Section 2.3.7"}],"minor_comments":[{"comment":"\"Slimpajama\" should be \"SlimPajama\" in the description of [347].","section":"Section 2.3.5"},{"comment":"In the description of AutoDS, the second prompt question is rendered as \"It it educationally valuable\"; this should be \"Is it educationally valuable.\"","section":"Section 2.3.4"},{"comment":"The word \"strctured\" should be \"structured\" in the description of the data processing framework from [368].","section":"Section 2.3.7.1"},{"comment":"The opening sentence says the storage techniques are \"categorized accroding to the tasks they address\"; \"accroding\" should be \"according.\"","section":"Section 2.4"},{"comment":"The sentence \"if a data isn't cached locally\" is ungrammatical; it should be \"if data is not cached locally.\"","section":"Section 2.4.4"},{"comment":"The statement that RAG data \"must be strictly reviewed to ensure authenticity and validity\" is not connected to a specific processing step or table entry; consider adding a cross-reference to Section 2.5.1 or clarifying how authenticity is achieved in the RAG data pipeline.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be useful to the community once the methodological support for the IaaS taxonomy and the comprehensiveness claim is added and the editorial inconsistencies are fixed. The concerns raised in the major comments are about verifiability rather than about the paper's direction, so I see major revision as the appropriate outcome rather than rejection. The GitHub repository may contain enough detail to support the \"over 400 papers\" claim, but that detail needs to be in the manuscript itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear [colleague],\n\nThis survey is a solid reference map for the LLM×data area, and it is broader than earlier surveys I've seen. The lifecycle structure — pre-training, SFT, RL, RAG, agents — plus the cross-cutting DATA4LLM/LLM4DATA split is genuinely useful, and the figures and tables are dense but mostly well organized. The IaaS acronym (Inclusiveness, Abundance, Articulation, Sanitization) is a good mnemonic, and it gives the survey a spine. If you want a starting point for where data management and LLMs actually meet, this fills a gap.\n\nThe soft spots are real but not fatal. The biggest one is the \"over 400 papers\" claim: no search protocol, inclusion criteria, or deduplication/validation procedure is given, only a GitHub link. That weakens the comprehensiveness claim, though not the day-to-day usefulness. Second, IaaS is asserted rather than validated. Section 2.1 says it is based on an investigation of those papers, but there is no coding protocol or inter-annotator reliability, and the four dimensions do overlap a bit (abundance and mixing; articulation and filtering). That is a framing weakness, not a reason to toss the survey. The paper would be stronger with a paragraph explaining how the dimensions were chosen and where their boundaries are.\n\nEditorial quality also needs a pass: I hit a numbering error (2.2.7.1 under Section 2.3.7), typos like 'Slimpajama' for SlimPajama, and some ungrammatical sentences. Minor, but they add up.\n\nOn the stress test: I agree with the concern that IaaS is asserted, not derived. But I would not call it load-bearing for the whole survey. The map still works if IaaS is treated as a mnemonic framing device rather than a validated instrument. The field does not have a well-tested taxonomy for LLM data quality, and this paper at least gives people a vocabulary.\n\nWho is it for? A grad student or researcher entering LLM×data, or a practitioner wanting a quick lay of the land. It is not a methods paper; no new empirical results. It deserves a serious referee: a reviewer should push for the corpus-selection methods paragraph and tighter IaaS definitions, but the core survey is worth publishing.\n\nRecommendation: accept after minor/moderate revision, with the methods note and copyediting addressed.","headline":"A useful but under-documented survey map; the IaaS lens is a framing device rather than a validated instrument, yet the paper deserves serious referee time.","tokens_in":47883,"tokens_out":3012,"would_cite":true,"duration_ms":27517,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims the first comprehensive, lifecycle-based map of the two-way relationship between LLMs and data management.","keywords":["large language models","data management","data-centric AI","data quality","data processing pipelines","retrieval-augmented generation","LLM agents","IaaS concept"],"falsifier":"Take 100 recent papers on LLM data management from an independent, broad literature search and try to place each one into the survey's stage-task and direction taxonomies; if more than about 10 percent fall outside the named categories or are described in ways their own abstracts contradict, the survey's comprehensiveness and accuracy claims would be falsified.","tokens_in":46834,"feed_emoji":"📊","tokens_out":7705,"duration_ms":64434,"temperature":0.7,"pith_summary":"The paper tries to establish that LLMs and data management are one bidirectionally connected field and to give that field a shared map. It organizes more than 400 papers into two directions: DATA4LLM, where data processing, storage, and serving supply models across pre-training, fine-tuning, reinforcement learning, RAG, agents, and evaluation; and LLM4DATA, where LLMs serve as general-purpose engines for data cleaning, analysis, and system optimization. To make data quality discussable, the authors introduce the 'IaaS' lens, which says good LLM data should be inclusive, abundant, articulated, and sanitized. A sympathetic reader would care because the map turns scattered techniques into a common vocabulary and exposes which stage-technique combinations are still empty.","feed_headline":"New survey charts the two-way street between LLMs and data","feed_subtitle":"It adds a four-part data-quality lens and a stage-by-stage map spanning pretraining, RAG, and agents.","key_machinery":"The carrying mechanism is the IaaS taxonomy, a four-part definition of what makes LLM data good: inclusiveness (broad, diverse coverage across domains, tasks, sources, languages, styles, and modalities), abundance (sufficient volume with balanced composition), articulation (well-formatted, clean, instructive, step-by-step data), and sanitization (privacy-compliant, toxicity-free, ethically consistent, risk-mitigated content). Around this lens the survey builds a lifecycle-based taxonomy that maps data processing, storage, and serving techniques onto LLM stages, from pre-training through evaluation, RAG, and agents, in a stage-task table, and pairs it with an LLM4DATA taxonomy for data manipulation, analysis, and system optimization. The taxonomy does the argumentative work: it turns scattered papers into comparable cells, exposes where techniques exist or are missing, and gives the field a shared reference structure.","core_discovery":"The paper's central claim is that the intersection of LLMs and data management is genuinely bidirectional and should be studied as one field, not two. On the DATA4LLM side, the paper identifies three families of techniques—data processing (acquisition, deduplication, filtering, selection, mixing, synthesis, and end-to-end pipelines), data storage (formats, distribution, organization, movement, fault tolerance, and KV caches), and data serving (shuffling, compression, packing, and provenance)—and shows how each applies differently across pre-training, continual pre-training, SFT, RL, RAG, agents, and evaluation. On the LLM4DATA side, it argues that LLMs are becoming general-purpose engines for data manipulation, analysis over structured, semi-structured, and unstructured data, and system optimization such as configuration tuning, query rewriting, and anomaly diagnosis. The authors position the survey as the first to cover this full lifecycle, and they anchor the data-quality discussion in the IaaS concept: good LLM data should be inclusive, abundant, articulated, and sanitized.","pith_inferences":["One natural next step is to turn the IaaS dimensions into a quantitative scoring rubric and test whether datasets that score higher on inclusiveness, abundance, articulation, and sanitization consistently produce better downstream model performance.","The same taxonomy could be reused as a living map: with explicit inclusion criteria and versioned updates, later surveys could track how quickly the empty cells in the stage-task table fill in.","If LLMs really are general-purpose data engines, then data-management products may converge on natural-language interfaces where users describe a cleaning or analysis task and the system composes the operations, an outcome the survey's LLM4DATA section points toward but does not itself predict.","Because the survey's literature selection is not documented, an immediate extension is to test whether the taxonomy is stable under a different, independently chosen corpus of LLM-data papers."],"forward_implications":["Dataset quality can now be assessed along four named axes, so a data pipeline can be audited for what it lacks rather than only for what it filters.","Every LLM stage has its own data profile; what works for pre-training is not what works for SFT, RAG, or agents, and the survey makes those differences explicit.","Data management should be treated as a first-class component of LLM development, with storage, movement, and serving as important as model architecture.","LLMs can take over classical data tasks such as cleaning, schema matching, and query tuning, shifting the bottleneck from handcrafted rules to prompt design and retrieval-augmented reasoning.","The survey's stage-task table marks combinations with no reported work, giving researchers a direct list of open technique-stage gaps."],"supporting_citations":[{"why":"Sets up the gap this survey fills: existing LLM-data surveys center on pre-training rather than the full lifecycle.","marker":"[405]"},{"why":"Cited as a prior survey focused narrowly on data selection, helping justify the broader processing taxonomy.","marker":"[55]"},{"why":"Cited as prior work on classical machine learning for data management, establishing the need for an LLM-specific account.","marker":"[488]"},{"why":"Supplies a large-scale web-crawling preprocessing pipeline used to illustrate end-to-end data processing for pre-training.","marker":"[311]"},{"why":"Provides a widely used data processing framework that anchors the survey's discussion of data processing and orchestration.","marker":"[90]"},{"why":"Provides the KV-cache management approach that grounds the data-storage-for-inference category.","marker":"[220]"},{"why":"Supplies the large-scale model example (671B parameters) used to motivate storage and fault-tolerance challenges.","marker":"[162]"},{"why":"Provides a graph-based RAG indexing approach that grounds the RAG data-organization discussion.","marker":"[127]"}],"fun_headline_variants":["Bidirectional survey: DATA4LLM and LLM4DATA in one map","IaaS data lens: inclusive, abundant, articulated, sanitized","Two-way street: LLMs shaped by data, data managed by LLMs","Stage-by-stage map of LLM–data synergy from pretraining to agents","Full lifecycle survey: data for LLMs, LLMs for data management"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 400-plus papers examined, selected without a documented protocol, accurately represent the LLM-data literature and are each characterized correctly enough to support the taxonomy.","fun_headline_variants_meta":{"raw":{"variants":["Bidirectional survey: DATA4LLM and LLM4DATA in one map","IaaS data lens: inclusive, abundant, articulated, sanitized","Two-way street: LLMs shaped by data, data managed by LLMs","Stage-by-stage map of LLM–data synergy from pretraining to agents","Full lifecycle survey: data for LLMs, LLMs for data management"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1858,"prompt_tokens":1032,"completion_tokens":826,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":727}},"tokens_in":648,"tokens_out":826,"duration_ms":6847,"temperature":1.0,"reasoning_tokens":727,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:29:59.232243+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take 100 recent papers on LLM data management from an independent, broad literature search and try to place each one into the survey's stage-task and direction taxonomies; if more than about 10 percent fall outside the named categories or are described in ways their own abstracts contradict, the survey's comprehensiveness and accuracy claims would be falsified.","supporting_citations":[],"review_version":1}