A seed-free, Wikipedia-backed synthetic data pipeline lets a 5,000-example Thai fine-tune reach BERTScore close to Thai LLMs trained on tens of thousands to hundreds of thousands of instructions.
In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 3029–3051, Singapore
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai
A seed-free, Wikipedia-backed synthetic data pipeline lets a 5,000-example Thai fine-tune reach BERTScore close to Thai LLMs trained on tens of thousands to hundreds of thousands of instructions.