{"id":"1765f607-80a3-4728-b3e6-968824748e45","arxiv_id":"2501.18648","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A literature review cataloging LLM-based augmentation methods across image, text, and speech, with a taxonomy of techniques, limitations, and suggested fixes.","lead":"This paper surveys recent uses of multimodal large language models to create synthetic training data for images, text, and speech. It organizes dozens of methods and known limitations into a single reference for researchers choosing augmentation strategies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's own inclusion criteria are nearly vacuous, and several table entries are not data-augmentation methods, so the claimed comprehensive coverage is not internally supported.","rationale":"The reader's weakest assumption concerns corpus representativeness and the lack of a reproducible search protocol. That concern is valid, but the more load-bearing issue is internal: the survey's own tables contain papers that are not data-augmentation studies, and its inclusion criterion Q3 is broad enough to admit them. This is verifiable from the manuscript alone, without needing an external literature search. If the corpus is inflated with non-augmentation applications, then every downstream product of the survey, the modality counts, the technique taxonomies, the limitation lists, and the firstness claim, is weakened. The reader's proposed fix (exact queries, date ranges, reproducible screening) is necessary but may not be sufficient: even a perfectly reproducible search would reproduce a corpus that includes off-topic entries. The concrete audit I propose would settle whether the misalignment is substantial. If the audit shows few false positives, the concern does not land and the survey's framework remains useful; if it shows many, the central claim is unsupported until the corpus and criteria are corrected. Either way, the reader's CONDITIONAL verdict is the right call, so I do not recommend changing it.","tokens_in":44494,"tokens_out":3287,"duration_ms":36968,"concrete_test":"Independently code all rows in Tables 1-3 (104 papers) using only each paper's abstract: (1) Does the paper propose or use data augmentation, defined as generating or transforming training samples for a downstream model? (2) Is an LLM or multimodal LLM central to the method? (3) Is the target modality image, text, or speech? Count the rows that fail criterion (1). If more than roughly 15-20% fail, the corpus does not support the central claim and the survey must be re-scoped or the inclusion criteria tightened. As a secondary reproducibility check, reconstruct the Section 2 search with explicit query strings and date ranges and compare the retrieved counts to the reported 24/45/35.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first comprehensive survey of LLM-based data augmentation across image, text, and speech. For that claim to hold, the corpus must actually consist of data-augmentation studies. Section 2.4 defines eligibility with three questions, the third of which is 'Does the article propose a framework, tool, or methodology?' Since almost every paper proposes some method, this criterion admits essentially any paper, making the screening non-restrictive. The tables confirm the problem: Table 1 lists DeepDR-LLM [106] and Med-MLLM [107], which are diagnostic/representation-learning systems rather than augmentation methods; Table 2 lists OphGLM [176] (ophthalmology assistant) and Forged-GAN-BERT [175] (authorship attribution); Table 3 lists LLM-Commentator [212] (football commentary), LAMB [214] (LMS assistant), LaMini-Flan-T5 [215] (video summarization), and MMed-Llama 3 [217] (multilingual medical corpus). None of these generates or modifies training samples for a downstream model, which is the paper's own definition of data augmentation from Section 1. If a substantial fraction of the 104 selected papers are off-topic, then the 24/45/35 counts, the fifteen-technique taxonomies, and the limitation analyses are built on work outside the stated scope. This is an internal consistency problem: the first-comprehensive-survey claim fails on the manuscript's own evidence even before external completeness is assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of recent work that uses multimodal large language models for data augmentation in image, text, and speech modalities. It proposes an eight-step (image, text) or seven-step (speech) process view of LLM-based augmentation, organizes the literature into technique taxonomies (fifteen techniques per modality in the figures), tabulates 24 image, 45 text, and 35 speech papers, and discusses limitations and potential solutions for each modality. The central claim, stated in the abstract and in Section 1, is that this is the first survey to comprehensively cover LLM-based data augmentation across all three modalities.","tokens_in":44660,"tokens_out":3734,"duration_ms":38293,"significance":"If the survey's corpus and taxonomies were reliable, the paper would be a useful entry point for researchers wanting a cross-modal view of LLM-based augmentation, and the organized lists of limitations and proposed solutions would have practical value. The paper does gather a substantial number of relevant recent works, includes a dedicated discussion of 3D point-cloud augmentation, and provides a public GitHub repository. However, the survey's central claims of comprehensiveness and firstness are not supported by the evidence in the manuscript: the screening protocol is non-reproducible, several tabulated entries are not data-augmentation methods under the paper's own definition, the historical speech section contains non-augmentation works, and the claimed fifteen-technique taxonomies are contradicted by the enumerated lists. These are load-bearing issues for a survey whose stated contribution is comprehensive coverage.","major_comments":[{"comment":"The inclusion criteria are effectively non-restrictive. The third eligibility question, 'Does the article propose a framework, tool, or methodology?', is satisfied by almost any paper, and the first two questions are redundant with the survey's topic. The screening is also described as consensus-based without any protocol for disagreements, and the search keywords, exact queries, date ranges, and per-database results are not reported. Without a reproducible search and screening protocol, the claim of comprehensive coverage in Section 1 and the abstract cannot be verified.","section":"Section 2.4"},{"comment":"Several tabulated papers do not perform data augmentation under the definition given in Section 1 (generating or modifying training samples for a downstream model). Table 1 lists DeepDR-LLM [106] and Med-MLLM [107], which are diagnostic and representation-learning systems; Table 2 lists OphGLM [176], an ophthalmology assistant, and Forged-GAN-BERT [175], an authorship-attribution method; Table 3 lists LLM-Commentator [212], LAMB [214], LaMini-Flan-T5 [215], and MMed-Llama 3 [217], which are commentary generation, LMS assistant, video summarization, and multilingual medical corpus efforts. These entries need to be removed or accompanied by an explicit explanation of the augmentation mechanism they propose; otherwise the 24/45/35 corpus counts and the fifteen-technique taxonomies are built on out-of-scope work.","section":"Tables 1-3"},{"comment":"The historical speech augmentation discussion includes works that are not data augmentation. Brandenburg et al. [66] is a surgical vocal-cord augmentation procedure, Watanabe et al. [67] is a teleconferencing system, Schmandt et al. [68] adds speech input to window systems, and Adams and Lang [69] studies the Lombard effect in Parkinson's patients. Presenting these as traditional speech data augmentation methods indicates that the screening criteria were not applied consistently and further weakens the internal consistency of the survey's scope.","section":"Section 3.2"},{"comment":"Both the text and speech subsections state that Figure 6 and Figure 7 outline 'fifteen diverse techniques,' but the enumerated lists contain only ten techniques in each case. The text subsection lists Paraphrasing, Back-Translation, Text Expansion, Role Playing, Synonym Replacement, Text Simplification, Textual Entailment Generation, Noise Injection, Contextual Variation, and Controlled Generation; the speech subsection lists ten similarly. The taxonomy counts are therefore internally inconsistent with the text, and the 'fifteen' number appears to be carried over from the image section without verification.","section":"Sections 4.2.2 and 4.3.2"},{"comment":"The 'first comprehensive survey' claim is not substantiated. The manuscript does not compare its coverage with existing LLM-augmentation surveys such as Ding et al. [19], does not report a completeness analysis, and the non-reproducible protocol and out-of-scope entries described above prevent the reader from assessing whether the corpus is representative. The claim should be softened or supported by a systematic, reproducible selection process and a comparison with prior surveys.","section":"Section 1 (Key Contributions)"}],"minor_comments":[{"comment":"The manuscript contains numerous typographical errors, including 'augmmentation,' 'Challanges,' 'Amplititude,' 'uch as,' 'random swapm,' 'outweight,' and 'theroy'; the paper needs a careful proofreading pass.","section":"General"},{"comment":"The database is referred to as 'DBSL' and 'DataBase systems and Logic Programming platform,' but the link points to dblp.uni-trier.de; the correct name is DBLP (Digital Bibliography and Library Project).","section":"Section 2.1"},{"comment":"The figure caption says keywords are color-coded in red, green, and blue, but the figure appears to be in grayscale or with colors that are not clearly distinguishable; the color coding should be made accessible or replaced with labels.","section":"Figure 2"},{"comment":"The phrase 'the study by [106]' and similar citations are used informally; several citations in the text refer to bracketed numbers that are not consistently tied to the Tables, making it hard to trace which paper supports which step.","section":"Section 4.1.1"},{"comment":"The discussion of DeepSeek R1 and reinforcement-learning-based self-augmentation is speculative and is not part of the reviewed corpus; it should be clearly labeled as a future outlook rather than a surveyed method.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a useful scope and some genuinely relevant content, but the corpus integrity and the reproducibility of the screening are the core issues. The off-topic entries in Tables 1-3 and in the historical speech section are factual problems that the authors can address by removing or reclassifying those papers, but the 'first comprehensive survey' claim will require either a substantially tightened protocol or a more modest framing. I would also ask the editor to verify the GitHub repository and the completeness of the reference list, since several cited items appear only in truncated form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful core here is real: a three-modality organization of LLM-based augmentation, with concrete tables and a limitation-solution mapping that a newcomer could actually navigate. The process overview figures (image, text, speech) are a decent tutorial scaffold. Giving credit where it is due, the authors did assemble a large corpus and tried to abstract common failure modes across modalities, which is more than most modality-specific surveys bother to do.\n\nThe soft spots are proportionate to how much the paper claims. The 'first comprehensive survey' claim is not supported by the methodology. Section 2.4's third inclusion question ('Does the article propose a framework, tool, or methodology?') admits essentially any paper, and the tables confirm the leak: DeepDR-LLM and Med-MLLM are diagnostic systems, OphGLM is an ophthalmology assistant, LLM-Commentator generates football commentary, LAMB is an LMS assistant, and Forged-GAN-BERT is authorship attribution. None of these generates or modifies training samples for a downstream model, which is the paper's own definition of augmentation from Section 1. That is an internal consistency problem, not just an external completeness concern. If a meaningful fraction of the 104 entries are off-topic, the 24/45/35 counts and the fifteen-technique taxonomies are built on shaky ground.\n\nThe historical background also mixes in non-augmentation work (vocal cord augmentation, teleconferencing) and the 'multimodal LLM' label is stretched to cover GANs and diffusion models without much discussion. The search protocol is not reproducible: no exact queries, date ranges, or screening logs, despite the reproducibility claim in the contributions. On citation patterns, the authors do reuse their own prior work in tables, but those are legitimately relevant to the topic and I would not flag that as misconduct.\n\nWho is this for? A graduate student who wants a quick map of the area and is willing to check the primary sources. The paper is not yet a reliable reference because the corpus needs cleaning and the firstness claim needs qualification. I would not cite it in my own work yet.\n\nRecommendation: send it to peer review, but with a heavy revision request: tighten the inclusion criteria, re-screen the tables, and either drop or carefully scope the 'first comprehensive' claim. A serious referee could turn this into a useful survey.","headline":"A useful three-modality framing undercut by a non-reproducible screening process and off-topic table entries; worth engaging after revision, not as it stands.","tokens_in":45287,"tokens_out":897,"would_cite":false,"duration_ms":12336,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey attempts to establish the first full map of LLM-based data augmentation across image, text, and speech, organized as per-modality pipelines, technique taxonomies, and matched limitations and solutions.","keywords":["data augmentation","large language models","multimodal LLMs","image augmentation","text augmentation","speech augmentation","synthetic data generation","survey"],"falsifier":"Re-run the search described in Section 2 with exact keyword families, date ranges, and inclusion criteria across the eight databases listed there, and check whether any earlier peer-reviewed survey already covered image, text, and speech augmentation together; finding one before 2025 would refute the firstness claim.","tokens_in":44203,"feed_emoji":"🔁","tokens_out":7132,"duration_ms":66114,"temperature":0.7,"pith_summary":"Data augmentation has shifted from hand-crafted transformations and LSTM-based generation toward context-aware synthetic data produced by multimodal large language models. This survey tries to establish that the shift can be mapped as one coherent field across image, text, and speech, and it claims to be the first review to cover all three modalities together. It organizes the recent literature into per-modality pipelines, named techniques, and paired lists of limitations and literature-sourced solutions, based on 104 studies published from 2020 onward. A sympathetic reader would care because the resulting map lets practitioners in computer vision, NLP, and audio research locate their methods, compare failure modes, and see where one modality's fixes might transfer to another.","feed_headline":"First survey maps LLM data augmentation across image, text, speech","feed_subtitle":"It organizes 104 recent studies into per-modality pipelines, technique lists, and matched limitations and solutions.","key_machinery":"The central organizing device is a three-modality taxonomy, one branch each for image, text, and speech. Each branch pairs a step-by-step augmentation pipeline (image encoding → prompt generation → instruction generation → natural-language-to-code translation → code execution → quality assessment → metadata generation → dataset integration; text encoding → prompt generation → instruction generation → transformation → execution → quality assessment → metadata → dataset integration; speech preprocessing → feature extraction → initial augmentation → multimodal embedding/contextual understanding → synthetic speech generation → refinement/filtering → dataset integration) with a list of named techniques and a set of limitations matched to literature-sourced solutions. This three-part structure is what lets the paper treat heterogeneous, modality-specific methods as instances of one LLM-driven data augmentation practice.","core_discovery":"The paper's central claim is that LLM-based data augmentation has become a genuinely cross-modal practice and that it can be captured in a single survey covering image, text, and speech together—a scope the authors state has not been attempted before. It argues that post-2020 multimodal LLMs produce context-aware synthetic data through language-mediated steps, replacing manual transformations and LSTM-era automation. On the basis of 104 curated studies (24 image, 45 text, 35 speech), it catalogs distinct augmentation techniques and common limitations for each modality, and it collects proposed remedies from the same literature. It concludes that the field is moving toward self-augmenting systems, citing reinforcement-learning approaches as a next step.","pith_inferences":["Editorial inference: the taxonomy suggests that evaluation across modalities is the missing piece; if each technique reported gains normalized by data size and compute, cross-modal comparison would become possible.","Editorial inference: the recurring limitations—LLMs lacking native acoustic and visual structure—point toward hybrid architectures combining LLMs with signal-processing or vision-specific modules, a direction the paper sketches but does not assert as its own finding.","Editorial inference: the firstness claim is empirically checkable by bibliographic date-mapping; a reader could reconstruct the search with explicit queries and verify whether any earlier survey already covered all three modalities.","Editorial inference: the proposed solutions are culled from the surveyed papers rather than validated here, so their effectiveness is an open empirical question rather than a settled result."],"forward_implications":["Researchers can place any new augmentation method into the appropriate pipeline and compare it with the named techniques already catalogued for that modality.","The limitation lists provide concrete design targets: image methods must guard against semantic misalignment, text methods against semantic drift and redundancy, speech methods against temporal distortion, timbre loss, and synthetic unrealism.","The literature-sourced solutions give practitioners ready-made starting points, such as natural-language-inference filtering for generated text and joint timbre-content modeling for speech.","If the firstness claim holds, later work on LLM-based augmentation will cite this survey as the reference map for the pre-2025 landscape."],"supporting_citations":[{"why":"It supplies a prior survey and taxonomy for time-series augmentation, one of the single-modality reviews the paper positions itself as extending.","marker":"[1]"},{"why":"It is the standard image data augmentation survey that the paper uses to set up the image-modality baseline and its limitations.","marker":"[2]"},{"why":"It provides a modern survey of augmentation approaches, used as evidence that prior reviews concentrate on ML/DL rather than LLMs.","marker":"[4]"},{"why":"It reviews augmentation for object detection specifically, illustrating the single-modality scope the paper claims to transcend.","marker":"[8]"},{"why":"It reviews medical image augmentation techniques, another single-modality prior survey the paper builds on.","marker":"[15]"},{"why":"It surveys text data augmentation, establishing the NLP-only scope the paper contrasts with its tri-modal coverage.","marker":"[25]"},{"why":"It surveys data augmentation for text classification, another NLP-only review that supports the claimed gap.","marker":"[27]"},{"why":"It supplies the systematic literature review methodology the paper follows for database search, screening, and selection.","marker":"[34]"}],"fun_headline_variants":["LLM augmentation goes multimodal in 104-study survey","Survey: LLMs augment image, text, and speech data","First cross-modal survey of LLM data augmentation","Survey of LLM augmentation spans three modalities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 104 curated studies are representative and complete enough to support a first comprehensive survey covering image, text, and speech; the search protocol leaves exact queries and date ranges unspecified, so a missed prior tri-modal survey would undermine the claim.","fun_headline_variants_meta":{"raw":{"variants":["LLM augmentation goes multimodal in 104-study survey","Survey: LLMs augment image, text, and speech data","First cross-modal survey of LLM data augmentation","Survey of LLM augmentation spans three modalities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001554,"raw_usage":{"total_tokens":6228,"prompt_tokens":977,"completion_tokens":5251,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":5189}},"tokens_in":593,"tokens_out":5251,"duration_ms":31821,"temperature":1.0,"reasoning_tokens":5189,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:33:05.644046+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the search described in Section 2 with exact keyword families, date ranges, and inclusion criteria across the eight databases listed there, and check whether any earlier peer-reviewed survey already covered image, text, and speech augmentation together; finding one before 2025 would refute the firstness claim.","supporting_citations":[],"review_version":1}