Pith. sign in

REVIEW 2 cited by

Mergen: The First Manchu-Korean Machine Translation Model Trained on Augmented Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17492 v2 pith:TOJ7LCFR submitted 2023-11-29 cs.CL

classification cs.CL
keywords manchu-koreanmodeltranslationmachinedatahistoricallanguagemanchu
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Manchu language, with its roots in the historical Manchurian region of Northeast China, is now facing a critical threat of extinction, as there are very few speakers left. In our efforts to safeguard the Manchu language, we introduce Mergen, the first-ever attempt at a Manchu-Korean Machine Translation (MT) model. To develop this model, we utilize valuable resources such as the Manwen Laodang(a historical book) and a Manchu-Korean dictionary. Due to the scarcity of a Manchu-Korean parallel dataset, we expand our data by employing word replacement guided by GloVe embeddings, trained on both monolingual and parallel texts. Our approach is built around an encoder-decoder neural machine translation model, incorporating a bi-directional Gated Recurrent Unit (GRU) layer. The experiments have yielded promising results, showcasing a significant enhancement in Manchu-Korean translation, with a remarkable 20-30 point increase in the BLEU score.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QueEn: A Large Language Model for Quechua-English Translation

    cs.CL 2024-12 reject novelty 3.0 of 10

    QueEn reports BLEU 17.6 for Quechua-English translation, but its internal table shows BLEU 0.235, the method is not reproducible, and the translation direction is inconsistent.

  2. Transcending Language Boundaries: Harnessing LLMs for Low-Resource Language Translation

    cs.CL 2024-11 reject novelty 3.0 of 10

    A retrieval-augmented GPT-4o pipeline improves surface-level metrics for low-resource translation, but the gains are small, semantic scores are inconsistent, and human acceptance remains near zero.

Pith tools