A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.
https://huggingface.co/datasets/Intel/orca_dpo_pairs Intel/orca\_dpo\_pairs · Datasets at Hugging Face
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
A shuffle-based token classifier plus group-level loss reweighting improves supervised fine-tuning of LLM agents on tool-use benchmarks.