Pith. sign in

REVIEW 2 cited by

Data Augmentation for Neural Machine Translation using Generative Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.16833 v2 pith:NGK3NLJD submitted 2023-07-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords dataaugmentationmodelsyntheticlanguagemachinemethodsmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the rapid growth in model architecture, the scarcity of large parallel corpora remains the main bottleneck in Neural Machine Translation. Data augmentation is a technique that enhances the performance of data-hungry models by generating synthetic data instead of collecting new ones. We explore prompt-based data augmentation approaches that leverage large-scale language models such as ChatGPT. To create a synthetic parallel corpus, we compare 3 methods using different prompts. We employ two assessment metrics to measure the diversity of the generated synthetic data. This approach requires no further model training cost, which is mandatory in other augmentation methods like back-translation. The proposed method improves the unaugmented baseline by 0.68 BLEU score.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Step-Opt: Boosting Optimization Modeling in LLMs through Iterative Data Synthesis and Structured Validation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Step-Opt, a LLaMA-3-8B model fine-tuned on iteratively evolved and stepwise-validated data, reports state-of-the-art accuracy on NL4OPT, MAMO, and IndustryOR.

  2. Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

    cs.CL 2025-01 conditional novelty 1.0 of 10

    A literature review that classifies LLM text data augmentation into simple, prompt-based, retrieval-based, and hybrid techniques, with post-processing and evaluation notes.

Pith tools