Pith. sign in

REVIEW 1 cited by

MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12958 v1 pith:DJ35YR6N submitted 2024-09-19 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesdatasetsinstructionlow-resourcemodelsmurituninginstructions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Instruction tuning enhances large language models (LLMs) by aligning them with human preferences across diverse tasks. Traditional approaches to create instruction tuning datasets face serious challenges for low-resource languages due to their dependence on data annotation. This work introduces a novel method, Multilingual Reverse Instructions (MURI), which generates high-quality instruction tuning datasets for low-resource languages without requiring human annotators or pre-existing multilingual models. Utilizing reverse instructions and a translation pipeline, MURI produces instruction-output pairs from existing human-written texts in low-resource languages. This method ensures cultural relevance and diversity by sourcing texts from different native domains and applying filters to eliminate inappropriate content. Our dataset, MURI-IT, includes more than 2 million instruction-output pairs across 200 languages. Evaluation by native speakers and fine-tuning experiments with mT5 models demonstrate the approach's effectiveness for both NLU and open-ended generation. We publicly release datasets and models at https://github.com/akoksal/muri.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish

    cs.CL 2025-10 conditional novelty 6.0 of 10

    LuxInstruct is the first native-output instruction-tuning dataset for Luxembourgish (537k samples, instructions in English/French/German); its evidence that cross-lingual tuning beats monolingual tuning is directional...

Pith tools