A two-stage data curation pipeline (GloVe similarity plus BERT educational-value scoring) produces an astronomy dataset that, after fine-tuning LLaMA-3-8B on 1B tokens, lifts MMLU astronomy from 69.08% to 76.3%.
B.3.2 Recruitment and Voluntary Participation All annotators were graduate students specializ- ing in astronomy
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study
A two-stage data curation pipeline (GloVe similarity plus BERT educational-value scoring) produces an astronomy dataset that, after fine-tuning LLaMA-3-8B on 1B tokens, lifts MMLU astronomy from 69.08% to 76.3%.