Reward-filtered distillation from 24 labeled examples improves structured reasoning F1 for fine-tuned Llama-3-8B agents on the LLMSR@XLLM25 test sets.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
LLMSR@XLLM25: Less is More: Enhancing Structured Multi-Agent Reasoning via Quality-Guided Distillation
Reward-filtered distillation from 24 labeled examples improves structured reasoning F1 for fine-tuned Llama-3-8B agents on the LLMSR@XLLM25 test sets.