Finetuning open LMs on ChatGPT outputs creates models that mimic style and fool human raters but fail to close the performance gap to proprietary systems on tasks not well-represented in the imitation data.
arXiv preprint arXiv:2302.10724 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2representative citing papers
A clustering-based synthetic data distillation framework enables compact models to match or exceed a large teacher on financial sentiment analysis using only a small set of real labeled examples.
citing papers explorer
-
The False Promise of Imitating Proprietary LLMs
Finetuning open LMs on ChatGPT outputs creates models that mimic style and fool human raters but fail to close the performance gap to proprietary systems on tasks not well-represented in the imitation data.
-
Efficient Financial Language Understanding via Distillation with Synthetic Data
A clustering-based synthetic data distillation framework enables compact models to match or exceed a large teacher on financial sentiment analysis using only a small set of real labeled examples.