REVIEW 8 cited by
Low-resource Languages: A Review of Past Work and Future Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A current problem in NLP is massaging and processing low-resource languages which lack useful training attributes such as supervised data, number of native speakers or experts, etc. This review paper concisely summarizes previous groundbreaking achievements made towards resolving this problem, and analyzes potential improvements in the context of the overall future research direction.
Forward citations
Cited by 8 Pith papers
-
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models
A new benchmark shows state-of-the-art LLMs perform poorly on three Taiwanese indigenous languages across MT, ASR, and summarization.
-
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
A new 15,000-sentence Marathi emotion benchmark shows GPT-4 and Llama3-405B outperform fine-tuned Marathi BERT and MuRIL, while BERT trained on GPT-4-generated labels still trails GPT-4.
-
An End-to-End Approach for Child Reading Assessment in the Xhosa Language
Fine-tuned speech models classify correct versus incorrect Xhosa child pronunciations on 10 EGRA reading items with about 91% diagnostic efficiency.
-
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
First end-to-end Efik TTS baseline: a 3-hour single-speaker corpus and four fine-tuned models, with MMS-TTS best at MOS 3.80±0.63 but residual tonal errors.
-
QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval
A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.
-
A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy
Off-the-shelf spaCy pipelines trained from scratch on the ARCOSG corpus reach 88.6% (fine-grained) and 93.7% (coarse-grained) POS tagging accuracy for Scottish Gaelic, comparable to prior custom-built taggers.
-
Bridging the Gap with Retrieval-Augmented Generation: Making Prosthetic Device User Manuals Available in Marginalised Languages
A proposed RAG-based framework for translating prosthetic device manuals into marginalised languages is described, but no results are presented.
- Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages
Discussion (0). Sign in to comment.