RAFT aligns generative models by ranking samples with a reward model and fine-tuning only on the top-ranked outputs, reporting gains on reward scores and automated metrics for LLMs and diffusion models.
arXiv preprint arXiv:2202.06417 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2verdicts
UNVERDICTED 2representative citing papers
SCENIC framework reports up to 99% exact match on structured IoT command generation using sub-0.2B models, with pruned INT8 versions retaining 91% EM@1 after 25% size reduction.
citing papers explorer
-
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
RAFT aligns generative models by ranking samples with a reward model and fine-tuning only on the top-ranked outputs, reporting gains on reward scores and automated metrics for LLMs and diffusion models.
-
SCENIC: Semantic-Conditioned Edge-Aware Neural Framework for Structured IoT Command Generation
SCENIC framework reports up to 99% exact match on structured IoT command generation using sub-0.2B models, with pruned INT8 versions retaining 91% EM@1 after 25% size reduction.