TwinRouterBench supplies 970 execution-verified router prefixes across five datasets plus a live harness for 100 held-out SWE-bench cases, scoring routers on tier accuracy, trajectory success, and realized token cost without LLM judges.
2601.07206 , archivePrefix=
11 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 11representative citing papers
LQM-ContextRoute routes LLM tool calls via latency-quality matching in a contextual bandit, improving F1 by 2.18 pp, accuracy by up to 18 pp, and NDCG by 2.91-3.22 pp over SW-UCB on web-search, StrategyQA, and retriever benchmarks.
Switchcraft routes agentic tool-calling queries to the lowest-cost model that preserves correctness, reaching 82.9% accuracy and 84% cost reduction on five benchmarks.
The paper introduces route receipts as a portable runtime record of routing decisions to make adaptive AI systems more transparent and trustworthy.
An online resample-or-reroute policy that allocates each unit of a per-query budget by estimated marginal correctness per unit cost attains a better cost–quality Pareto front than single-route, cascade, and best-of-K baselines on four multi-draw LLM pools.
Agent-as-a-Router turns static LLM routing into an iterative C-A-F loop that accumulates execution feedback to lower cumulative regret on coding tasks.
LLM routers across 21 methods on 5 benchmarks converge to similar accuracy below oracle due to learning global performance trends rather than fine-grained query signals.
A learned orchestration policy for LLM agents that jointly optimizes task decomposition and selective routing to (model, primitive) pairs, delivering 77% macro pass@1 at 10x lower cost than strong baselines across 13 benchmarks.
The Workload-Router-Pool architecture is a 3D framework for LLM inference optimization that synthesizes prior vLLM work into a 3x3 interaction matrix and proposes 21 research directions at the intersections.
RealRoute uses parallel source-agnostic retrieval followed by dynamic verification to improve accuracy over predictive LLM routers in heterogeneous multi-hop RAG tasks.
Aleena is an open-source AI agent that ingests multi-modal research software collaboration artifacts and transforms them into structured GitHub records to maintain continuous stakeholder alignment across the project lifecycle.
citing papers explorer
-
TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing
TwinRouterBench supplies 970 execution-verified router prefixes across five datasets plus a live harness for 100 held-out SWE-bench cases, scoring routers on tier accuracy, trajectory success, and realized token cost without LLM judges.
-
Latency-Quality Routing for Functionally Equivalent Tools in LLM Agents
LQM-ContextRoute routes LLM tool calls via latency-quality matching in a contextual bandit, improving F1 by 2.18 pp, accuracy by up to 18 pp, and NDCG by 2.91-3.22 pp over SW-UCB on web-search, StrategyQA, and retriever benchmarks.
-
Switchcraft: AI Model Router for Agentic Tool Calling
Switchcraft routes agentic tool-calling queries to the lowest-cost model that preserves correctness, reaching 82.9% accuracy and 84% cost reduction on five benchmarks.
-
Model Routing as a Trust Problem: Route Receipts for Adaptive AI Systems
The paper introduces route receipts as a portable runtime record of routing decisions to make adaptive AI systems more transparent and trustworthy.
-
Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models
An online resample-or-reroute policy that allocates each unit of a per-query budget by estimated marginal correctness per unit cost attains a better cost–quality Pareto front than single-route, cascade, and best-of-K baselines on four multi-draw LLM pools.
-
Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Agent-as-a-Router turns static LLM routing into an iterative C-A-F loop that accumulates execution feedback to lower cumulative regret on coding tasks.
-
The Routing Plateau: Understanding and Breaking the Accuracy Limits of LLM Routers
LLM routers across 21 methods on 5 benchmarks converge to similar accuracy below oracle due to learning global performance trends rather than fine-grained query signals.
-
Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation
A learned orchestration policy for LLM agents that jointly optimizes task decomposition and selective routing to (model, primitive) pairs, delivering 77% macro pass@1 at 10x lower cost than strong baselines across 13 benchmarks.
-
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
The Workload-Router-Pool architecture is a 3D framework for LLM inference optimization that synthesizes prior vLLM work into a 3x3 interaction matrix and proposes 21 research directions at the intersections.
-
RealRoute: Dynamic Query Routing System via Retrieve-then-Verify Paradigm
RealRoute uses parallel source-agnostic retrieval followed by dynamic verification to improve accuracy over predictive LLM routers in heterogeneous multi-hop RAG tasks.
-
Aleena: Alignment Agent for Research Software Engineering Collaborations
Aleena is an open-source AI agent that ingests multi-modal research software collaboration artifacts and transforms them into structured GitHub records to maintain continuous stakeholder alignment across the project lifecycle.