On a balanced benchmark built from the OSDG community dataset, fine-tuned LLaMa-2 13B achieves the highest macro F1 (92.4%), and small models such as Flan-T5-base (220M) reach 90.5%, close to fine-tuned GPT-3.5 (91.4%).
Ziegler et al., ‘‘How to Measure Sustainability? An Open-Data Ap- proach’’,Sustainability, vol
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
A Comparative Study of Task Adaptation Techniques of Large Language Models for Identifying Sustainable Development Goals
On a balanced benchmark built from the OSDG community dataset, fine-tuned LLaMa-2 13B achieves the highest macro F1 (92.4%), and small models such as Flan-T5-base (220M) reach 90.5%, close to fine-tuned GPT-3.5 (91.4%).