Pith. sign in

REVIEW 4 cited by

Controllable Data Augmentation for Few-Shot Text Mining with Chain-of-Thought Attribute Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.07099 v3 pith:WMRDAUZO submitted 2023-07-14 cs.CL

classification cs.CL
keywords attributemanipulationtextaugmentationchain-of-thoughtdataanalysisapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Prompting large language models (LLMs) for data augmentation has recently become a common practice in few-shot NLP tasks. In this paper, we propose Chain-of-Thought Attribute Manipulation (CoTAM), a novel approach that generates new data from existing examples by only tweaking in the user-provided, task-specific attribute, e.g., sentiment polarity or topic in movie reviews. Instead of conventional latent representation controlling, we leverage the chain-of-thought prompting to directly edit the text in three steps, (1) attribute decomposition, (2) manipulation proposal, and (3) sentence reconstruction. Extensive results on various tasks, such as text (pair) classification, aspect-based sentiment analysis, and conditional text generation, verify the superiority of CoTAM over other LLM-based augmentation methods with the same number of training examples for both fine-tuning and in-context learning. Remarkably, the 2D visualization of the augmented dataset using principal component analysis revealed a human-recognizable decision boundary that is likely hinted by the attribute manipulation, demonstrating the potential of our proposed approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-based Semantic Augmentation for Harmful Content Detection

    cs.CL 2025-04 conditional novelty 6.0 of 10

    LLM-generated explanations and trigger words, added to training text, improve harmful-content classifiers and can approach human-annotation performance at lower cost, with caveats about uneven comparisons.

  2. DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease

    cs.CL 2025-02 reject novelty 5.0 of 10

    DECT, an LLM-based transcript-distillation and synthetic-data augmentation pipeline, reports 90.48% accuracy on ADReSSo Alzheimer's detection, but with an unspecified evaluation split.

  3. Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration

    cs.NI 2025-07 conditional novelty 4.0 of 10

    A survey of multi-LLM systems in edge computing, covering architectures, enabling technologies, trust mechanisms, applications, and open datasets for edge general intelligence.

  4. Enhancing Granular Sentiment Classification with Chain-of-Thought Prompting in Large Language Models

    cs.CL 2025-05 reject novelty 3.0 of 10

    Chain-of-thought prompting raised GPT-4's granular sentiment classification accuracy on 2,000 Amazon app reviews from 84% to 93%, though the paper's own numbers and example data are inconsistent.

Pith tools