REFINE-AF generates instruction-tuning data with small open LLMs and adds reinforcement learning from automated feedback, beating a Self-Instruct baseline on 63-66% of SUPER-NI tasks at 15k instructions.
Guess the instruction! flipped learning makes language models stronger zero-shot learners,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
REFINE-AF: A Task-Agnostic Framework to Align Language Models via Self-Generated Instructions using Reinforcement Learning from Automated Feedback
REFINE-AF generates instruction-tuning data with small open LLMs and adds reinforcement learning from automated feedback, beating a Self-Instruct baseline on 63-66% of SUPER-NI tasks at 15k instructions.