REVIEW 4 cited by
Controllable Data Augmentation for Few-Shot Text Mining with Chain-of-Thought Attribute Manipulation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Prompting large language models (LLMs) for data augmentation has recently become a common practice in few-shot NLP tasks. In this paper, we propose Chain-of-Thought Attribute Manipulation (CoTAM), a novel approach that generates new data from existing examples by only tweaking in the user-provided, task-specific attribute, e.g., sentiment polarity or topic in movie reviews. Instead of conventional latent representation controlling, we leverage the chain-of-thought prompting to directly edit the text in three steps, (1) attribute decomposition, (2) manipulation proposal, and (3) sentence reconstruction. Extensive results on various tasks, such as text (pair) classification, aspect-based sentiment analysis, and conditional text generation, verify the superiority of CoTAM over other LLM-based augmentation methods with the same number of training examples for both fine-tuning and in-context learning. Remarkably, the 2D visualization of the augmented dataset using principal component analysis revealed a human-recognizable decision boundary that is likely hinted by the attribute manipulation, demonstrating the potential of our proposed approach.
Forward citations
Cited by 4 Pith papers
-
LLM-based Semantic Augmentation for Harmful Content Detection
LLM-generated explanations and trigger words, added to training text, improve harmful-content classifiers and can approach human-annotation performance at lower cost, with caveats about uneven comparisons.
-
DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
DECT, an LLM-based transcript-distillation and synthetic-data augmentation pipeline, reports 90.48% accuracy on ADReSSo Alzheimer's detection, but with an unspecified evaluation split.
-
Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration
A survey of multi-LLM systems in edge computing, covering architectures, enabling technologies, trust mechanisms, applications, and open datasets for edge general intelligence.
-
Enhancing Granular Sentiment Classification with Chain-of-Thought Prompting in Large Language Models
Chain-of-thought prompting raised GPT-4's granular sentiment classification accuracy on 2,000 Amazon app reviews from 84% to 93%, though the paper's own numbers and example data are inconsistent.
Discussion (0). Continue with ORCID to comment.