Pith. sign in

REVIEW 1 cited by

Shortcomings of LLMs for Low-Resource Translation: Retrieval and Understanding are Both the Problem

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15625 v3 pith:Q3TNZZIO submitted 2024-06-21 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords contextllmstranslationtypelanguagelow-resourcemodelretrieval
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work investigates the in-context learning abilities of pretrained large language models (LLMs) when instructed to translate text from a low-resource language into a high-resource language as part of an automated machine translation pipeline. We conduct a set of experiments translating Southern Quechua to Spanish and examine the informativity of various types of context retrieved from a constrained database of digitized pedagogical materials (dictionaries and grammar lessons) and parallel corpora. Using both automatic and human evaluation of model output, we conduct ablation studies that manipulate (1) context type (morpheme translations, grammar descriptions, and corpus examples), (2) retrieval methods (automated vs. manual), and (3) model type. Our results suggest that even relatively small LLMs are capable of utilizing prompt context for zero-shot low-resource translation when provided a minimally sufficient amount of relevant linguistic information. However, the variable effects of context type, retrieval method, model type, and language-specific factors highlight the limitations of using even the best LLMs as translation systems for the majority of the world's 7,000+ languages and their speakers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMsAgainstHate @ NLU of Devanagari Script Languages 2025: Hate Speech Detection and Target Identification in Devanagari Languages via Parameter Efficient Fine-Tuning of LLMs

    cs.CL 2024-12 conditional novelty 3.0 of 10

    LoRA fine-tuning of Nemo-Instruct achieves the best F1 among four LLMs on Devanagari hate speech detection (90.05%) and target identification (71.47%), but without baseline comparisons the approach's efficacy is unproven.

Pith tools