{"paper":{"title":"Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"LQM-ContextRoute routes LLM agents to equivalent tool providers by expected answer quality per service cycle rather than additive rewards.","cross_cats":[],"primary_cat":"cs.LG","authors_text":"Dawei Xiang, Kexin Chu, Wei Zhang","submitted_at":"2026-05-14T01:14:13Z","abstract_excerpt":"Tool-augmented LLM agents increasingly access the same tool type through multiple functionally equivalent providers, such as web-search APIs, retrievers, or LLM backends exposed behind a shared interface. This creates a provider-routing problem under runtime load: the router must choose among providers that differ in latency, reliability, and answer quality, often without gold labels at deployment time. We introduce LQM-ContextRoute, a contextual bandit router for same-function tool providers. Its key design is latency-quality matching: instead of letting low latency offset poor answers in an "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"On the main web-search load benchmark, LQM-ContextRoute improves F1 by +2.18 pp over SW-UCB while staying on the latency-quality frontier; in high-heterogeneity StrategyQA it improves accuracy by up to +18 pp.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That LLM-as-judge feedback provides a sufficiently reliable and unbiased signal for online adaptation without gold labels at deployment time.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"LQM-ContextRoute routes tool calls by expected quality per service cycle using contextual bandits and LLM-as-judge feedback, yielding +2.18 pp F1, up to +18 pp accuracy, and +2.91-3.22 pp NDCG gains over SW-UCB on web-search, StrategyQA, and retriever benchmarks.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"LQM-ContextRoute routes LLM agents to equivalent tool providers by expected answer quality per service cycle rather than additive rewards.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"c0e416f7ce9dc3729df14509f26d09644135acc93ba2c0c592020be26145a31c"},"source":{"id":"2605.14241","kind":"arxiv","version":1},"verdict":{"id":"b169b46d-03f5-4522-bcf1-5aa034ea2e84","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-15T02:55:32.026976Z","strongest_claim":"On the main web-search load benchmark, LQM-ContextRoute improves F1 by +2.18 pp over SW-UCB while staying on the latency-quality frontier; in high-heterogeneity StrategyQA it improves accuracy by up to +18 pp.","one_line_summary":"LQM-ContextRoute routes tool calls by expected quality per service cycle using contextual bandits and LLM-as-judge feedback, yielding +2.18 pp F1, up to +18 pp accuracy, and +2.91-3.22 pp NDCG gains over SW-UCB on web-search, StrategyQA, and retriever benchmarks.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That LLM-as-judge feedback provides a sufficiently reliable and unbiased signal for online adaptation without gold labels at deployment time.","pith_extraction_headline":"LQM-ContextRoute routes LLM agents to equivalent tool providers by expected answer quality per service cycle rather than additive rewards."},"references":{"count":25,"sample":[{"doi":"","year":1966,"title":"1966.Lectures on Functional Equations and Their Applications, volume 19 ofMathematics in Science and Engineering","work_id":"27c26f18-8219-41f7-9192-2c59b704d04e","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2025,"title":"Learning to route llms from bandit feedback","work_id":"5eaea69f-5b22-46cf-9070-203ac88610b9","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2026,"title":"https://modelcontextprotocol","work_id":"9b1b64ec-1af7-4d58-8599-04e1eb1239ca","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2026,"title":"https://docs.litellm","work_id":"99201e17-0dee-4515-aae0-4ffa2795d4d3","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":null,"title":"Thompson Sampling contextual bandit over heterogeneous tools (PubMed, drug DBs, calculator, web) with composite reward including latency","work_id":"881a01cc-71f8-46d0-a82d-e58650ee4a92","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":25,"snapshot_sha256":"0aa13764d4de90d2c09043284e4c48ab0e3a1e9ecd5f5ae57d3ec273f72b3df5","internal_anchors":7},"formal_canon":{"evidence_count":2,"snapshot_sha256":"d414e39d04f3affc032a4abf605015c36256497c24cbd0a48917bae5e24e21a2"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}