On a balanced benchmark built from the OSDG community dataset, fine-tuned LLaMa-2 13B achieves the highest macro F1 (92.4%), and small models such as Flan-T5-base (220M) reach 90.5%, close to fine-tuned GPT-3.5 (91.4%).
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
On a balanced benchmark built from the OSDG community dataset, fine-tuned LLaMa-2 13B achieves the highest macro F1 (92.4%), and small models such as Flan-T5-base (220M) reach 90.5%, close to fine-tuned GPT-3.5 (91.4%).