A single-head cross-attention router that jointly encodes the query and each candidate model improves cost-quality trade-offs on RouterBench by a few percent over KNN, MLP, and SVM baselines, though the headline gains are smaller than the abstract suggests.
Optllm: Optimal assignment of queries to large language models
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
A single-head cross-attention router that jointly encodes the query and each candidate model improves cost-quality trade-offs on RouterBench by a few percent over KNN, MLP, and SVM baselines, though the headline gains are smaller than the abstract suggests.