Pith. sign in

REVIEW 2 cited by

An Effective Deployment of Diffusion LM for Data Augmentation in Low-Resource Sentiment Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03203 v2 pith:4XZNRWTO submitted 2024-09-05 cs.CL cs.AI

classification cs.CLcs.AI
keywords diffusionlow-resourcemodelsentimenttokensaugmentationbalanceclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Sentiment classification (SC) often suffers from low-resource challenges such as domain-specific contexts, imbalanced label distributions, and few-shot scenarios. The potential of the diffusion language model (LM) for textual data augmentation (DA) remains unexplored, moreover, textual DA methods struggle to balance the diversity and consistency of new samples. Most DA methods either perform logical modifications or rephrase less important tokens in the original sequence with the language model. In the context of SC, strong emotional tokens could act critically on the sentiment of the whole sequence. Therefore, contrary to rephrasing less important context, we propose DiffusionCLS to leverage a diffusion LM to capture in-domain knowledge and generate pseudo samples by reconstructing strong label-related tokens. This approach ensures a balance between consistency and diversity, avoiding the introduction of noise and augmenting crucial features of datasets. DiffusionCLS also comprises a Noise-Resistant Training objective to help the model generalize. Experiments demonstrate the effectiveness of our method in various low-resource scenarios including domain-specific and domain-general problems. Ablation studies confirm the effectiveness of our framework's modules, and visualization studies highlight optimal deployment conditions, reinforcing our conclusions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A framework called DIFND generates debunking evidence via conditional diffusion and uses multi-agent MLLM reasoning to detect fake news videos, outperforming baselines on FakeSV and FVC.

  2. ProtoConNet: Prototypical Augmentation and Alignment for Open-Set Few-Shot Image Classification

    cs.CV 2025-07 reject novelty 4.0 of 10

    ProtoConNet improves open-set few-shot classification by combining clustering-based sample selection, contextual feature fusion, and prototypical alignment with a threshold-based unknown-class detector.

Pith tools