Pith. sign in

REVIEW 8 cited by

Low-resource Languages: A Review of Past Work and Future Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.07264 v1 pith:LB6LTBFH submitted 2020-06-12 cs.CL

classification cs.CL
keywords futurelanguageslow-resourceproblemreviewachievementsanalyzesattributes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A current problem in NLP is massaging and processing low-resource languages which lack useful training attributes such as supervised data, number of native speakers or experts, etc. This review paper concisely summarizes previous groundbreaking achievements made towards resolving this problem, and analyzes potential improvements in the context of the overall future research direction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark shows state-of-the-art LLMs perform poorly on three Taiwanese indigenous languages across MT, ASR, and summarization.

  2. L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 15,000-sentence Marathi emotion benchmark shows GPT-4 and Llama3-405B outperform fine-tuned Marathi BERT and MuRIL, while BERT trained on GPT-4-generated labels still trails GPT-4.

  3. An End-to-End Approach for Child Reading Assessment in the Xhosa Language

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Fine-tuned speech models classify correct versus incorrect Xhosa child pronunciations on 10 EGRA reading items with about 91% diagnostic efficiency.

  4. Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

    cs.CL 2026-07 conditional novelty 5.5 of 10

    First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.

  5. QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval

    cs.IR 2025-06 conditional novelty 4.0 of 10

    A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.

  6. A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy

    cs.CL 2026-08 conditional novelty 3.0 of 10

    Off-the-shelf spaCy pipelines trained from scratch on the ARCOSG corpus reach 88.6% (fine-grained) and 93.7% (coarse-grained) POS tagging accuracy for Scottish Gaelic, comparable to prior custom-built taggers.

  7. Bridging the Gap with Retrieval-Augmented Generation: Making Prosthetic Device User Manuals Available in Marginalised Languages

    cs.LG 2025-06 reject novelty 2.0 of 10

    A proposed RAG-based framework for translating prosthetic device manuals into marginalised languages is described, but no results are presented.

  8. Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

    cs.CL 2025-06

Pith tools