A seed-free, Wikipedia-backed synthetic data pipeline lets a 5,000-example Thai fine-tune reach BERTScore close to Thai LLMs trained on tens of thousands to hundreds of thousands of instructions.
In Findings of the Association for Com- putational Linguistics: EMNLP 2023, pages 12365– 12394, Singapore
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai
A seed-free, Wikipedia-backed synthetic data pipeline lets a 5,000-example Thai fine-tune reach BERTScore close to Thai LLMs trained on tens of thousands to hundreds of thousands of instructions.