Pith. sign in

REVIEW 3 cited by

Neural Machine Translation for Low-Resource Languages: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.15115 v1 pith:ZVCIIB4V submitted 2021-06-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords researchlow-resourcelanguagelrl-nmtmachinesurveytranslationneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Machine Translation (NMT) has seen a tremendous spurt of growth in less than ten years, and has already entered a mature phase. While considered as the most widely used solution for Machine Translation, its performance on low-resource language pairs still remains sub-optimal compared to the high-resource counterparts, due to the unavailability of large parallel corpora. Therefore, the implementation of NMT techniques for low-resource language pairs has been receiving the spotlight in the recent NMT research arena, thus leading to a substantial amount of research reported on this topic. This paper presents a detailed survey of research advancements in low-resource language NMT (LRL-NMT), along with a quantitative analysis aimed at identifying the most popular solutions. Based on our findings from reviewing previous work, this survey paper provides a set of guidelines to select the possible NMT technique for a given LRL data setting. It also presents a holistic view of the LRL-NMT research landscape and provides a list of recommendations to further enhance the research efforts on LRL-NMT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

    cs.CL 2026-07 conditional novelty 6.5 of 10

    CKTN is the first Cham–Khmer–Tay-Nung corpus; vocabulary augmentation plus script-calibrated replaced-token pretraining yields the strongest classification encoder and exposes misleading adaptation metrics.

  2. Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A new multilingual corpus for Cham, Khmer, and Tay-Nung, plus a script-aware ELECTRA-style recipe, achieves the best topic-classification accuracy among the tested encoders — and shows why lexical-overlap retrieval ca...

  3. BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning

    cs.LG 2026-08 reject novelty 4.0 of 10

    A 90%-pruned few-shot Bengali model is reported to rival larger baselines on some tasks, but the reported F1 scores contradict the paper's own precision and recall values.

Pith tools