Pith. sign in

REVIEW 5 cited by

A Survey on Recent Approaches for Natural Language Processing in Low-Resource Scenarios

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.12309 v3 pith:AGPVMLT5 submitted 2020-10-23 cs.CL cs.LG

classification cs.CLcs.LG
keywords datalanguagelow-resourcenaturalsurveyapproacheslearningmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep neural networks and huge language models are becoming omnipresent in natural language applications. As they are known for requiring large amounts of training data, there is a growing body of work to improve the performance in low-resource settings. Motivated by the recent fundamental changes towards neural models and the popular pre-train and fine-tune paradigm, we survey promising approaches for low-resource natural language processing. After a discussion about the different dimensions of data availability, we give a structured overview of methods that enable learning when training data is sparse. This includes mechanisms to create additional labeled data like data augmentation and distant supervision as well as transfer learning settings that reduce the need for target supervision. A goal of our survey is to explain how these methods differ in their requirements as understanding them is essential for choosing a technique suited for a specific low-resource setting. Further key aspects of this work are to highlight open issues and to outline promising directions for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark shows state-of-the-art LLMs perform poorly on three Taiwanese indigenous languages across MT, ASR, and summarization.

  2. Language verY Rare for All

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A single-GPU pipeline mixing LLM fine-tuning, RAG, and French-Italian transfer learning produces a French-Monégasque translator that matches or exceeds NLLB-200 on BLEU and METEOR.

  3. Low-Resource Fast Text Classification Based on Intra-Class and Inter-Class Distance Calculation

    cs.CL 2024-12 conditional novelty 5.0 of 10

    LFTC is a two-stage compression-based text classifier that builds per-class Zstd dictionaries and then applies NCD with KNN, beating gzip and several pre-trained models on low-resource benchmarks.

  4. Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning

    cs.SD 2025-04 reject novelty 4.0 of 10

    Fine-tuning Wav2Vec 2.0 on a custom Kurdish corpus is reported to cut speaker diarization error by 7.2 percentage points and raise cluster purity by about 13 percentage points, though the paper contains conflicting numbers.

  5. Task-Oriented Dialog Systems for the Senegalese Wolof Language

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A Rasa-based Wolof task-oriented dialog system, trained on French MASSIVE data projected through an in-house French-Wolof machine translation system, achieves near-French intent classification but weaker slot filling.

Pith tools