REVIEW 5 cited by
A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deep neural networks and huge language models are becoming omnipresent in natural language applications. As they are known for requiring large amounts of training data, there is a growing body of work to improve the performance in low-resource settings. Motivated by the recent fundamental changes towards neural models and the popular pre-train and fine-tune paradigm, we survey promising approaches for low-resource natural language processing. After a discussion about the different dimensions of data availability, we give a structured overview of methods that enable learning when training data is sparse. This includes mechanisms to create additional labeled data like data augmentation and distant supervision as well as transfer learning settings that reduce the need for target supervision. A goal of our survey is to explain how these methods differ in their requirements as understanding them is essential for choosing a technique suited for a specific low-resource setting. Further key aspects of this work are to highlight open issues and to outline promising directions for future research.
Forward citations
Cited by 5 Pith papers
-
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models
A new benchmark shows state-of-the-art LLMs perform poorly on three Taiwanese indigenous languages across MT, ASR, and summarization.
-
Language verY Rare for All
A single-GPU pipeline mixing LLM fine-tuning, RAG, and French-Italian transfer learning produces a French-Monégasque translator that matches or exceeds NLLB-200 on BLEU and METEOR.
-
Low-Resource Fast Text Classification Based on Intra-Class and Inter-Class Distance Calculation
LFTC is a two-stage compression-based text classifier that builds per-class Zstd dictionaries and then applies NCD with KNN, beating gzip and several pre-trained models on low-resource benchmarks.
-
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
Fine-tuning Wav2Vec 2.0 on a custom Kurdish corpus is reported to cut speaker diarization error by 7.2 percentage points and raise cluster purity by about 13 percentage points, though the paper contains conflicting numbers.
-
Task-Oriented Dialog Systems for the Senegalese Wolof Language
A Rasa-based Wolof task-oriented dialog system, trained on French MASSIVE data projected through an in-house French-Wolof machine translation system, achieves near-French intent classification but weaker slot filling.
Discussion (0). Continue with ORCID to comment.