Pith. sign in

REVIEW 29 cited by

AutoMix: Automatically Mixing Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12963 v5 pith:B2RBQZSD submitted 2023-10-19 cs.CL cs.AI

AutoMix: Automatically Mixing Language Models

classification cs.CL cs.AI
keywords automixlanguagemodelschallengingcomputationalcosteffectivelyfive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present Automix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to Automix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50% for comparable performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models

    cs.AI 2026-06 unverdicted novelty 7.0

    Any single-output LLM ensemble is accuracy-capped at 1-beta where beta is the all-models-wrong rate, a quantity not captured by pairwise correlations and frequently underestimated by copula models.

  2. DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

    cs.AI 2026-05 unverdicted novelty 7.0

    DecisionBench supplies a fixed task suite, model pool, delegation interface, and multi-axis metrics to evaluate emergent delegation, showing similar quality across awareness conditions but 15-31 point headroom under p...

  3. Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models

    eess.AS 2026-04 unverdicted novelty 7.0

    Semantic-level and verification-based uncertainty methods outperform token-level baselines for audio reasoning in ALLMs, but their relative performance on hallucination and unanswerable-question benchmarks is model- a...

  4. Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

    cs.CR 2026-04 unverdicted novelty 7.0

    R²A uses a hybrid ensemble surrogate router and suffix optimization to significantly increase black-box LLM router selection of expensive models across query distributions.

  5. A Workflow-Aware Serving Layer for Agentic Applications

    cs.DC 2026-07 conditional novelty 6.5

    A workflow-aware serving layer compiles per-node model-verifier-backend plans with an ILP and adapts only uncommitted work via pre-solved pressure rungs and residual re-solves.

  6. CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

    cs.AI 2026-07 conditional novelty 6.0

    A trained router plus conformal budget calibration lets coding agents choose cheap recovery vs. escalation after a failed attempt, producing a cost–quality frontier from a single model.

  7. CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents

    cs.AI 2026-07 unverdicted novelty 6.0

    Budget-calibrated recovery routing with conformal risk control lets coding agents match always-escalate solve rates at about 35% of the cost.

  8. SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks

    cs.SE 2026-06 unverdicted novelty 6.0

    SWE-Router introduces trajectory-conditioned value-based routing for LLM agents on SWE tasks, with a Bayes-optimality theorem and empirical cost savings while retaining most strong-model performance.

  9. Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management

    cs.LG 2026-06 unverdicted novelty 6.0

    NSR reframes continual learning as retrieval-based subspace memory management with SVD compression and similarity retrieval from a TaskKnowledgeBank, showing that the memory mechanism itself drives performance gains o...

  10. Selective Ensemble Based on Preference-Directed Multi-Objective Bandits

    cs.LG 2026-06 unverdicted novelty 6.0

    Introduces Pareto C-optimality and PrefUCB algorithm for PDMOB with instance-dependent logarithmic regret bounds, validated on selective ensemble and asset allocation tasks.

  11. DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?

    cs.RO 2026-06 unverdicted novelty 6.0

    DIRECT is a multimodal-context router that allocates test-time compute across chain-of-thought depth, model size, and memory history for VLM embodied planners, improving the success-cost Pareto frontier and matching s...

  12. HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools

    cs.CL 2026-05 unverdicted novelty 6.0

    HyDRA routes queries to cost-effective LLMs by predicting multi-dimensional capability requirements with a multi-head encoder and applying shortfall matching against configuration-defined model profiles, delivering up...

  13. LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?

    cs.AI 2026-05 unverdicted novelty 6.0

    LatentRouter routes image-question queries to the best MLLM by predicting counterfactual performance via latent communication between learned query capsules and model capability tokens.

  14. Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge

    cs.AI 2026-05 unverdicted novelty 6.0

    RACER routes between reasoning and non-reasoning LLM judges via constrained distributionally robust optimization to achieve better accuracy-cost trade-offs under distribution shift.

  15. Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation

    cs.AI 2026-05 unverdicted novelty 6.0

    A learned orchestration policy for LLM agents that jointly optimizes task decomposition and selective routing to (model, primitive) pairs, delivering 77% macro pass@1 at 10x lower cost than strong baselines across 13 ...

  16. AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?

    cs.AI 2026-05 unverdicted novelty 6.0

    Small open-weight models match GPT-5 on routine agent tool-use tasks but lag on long-horizon planning, supporting tiered routing to reduce costs in agentic systems.

  17. Privacy-Preserving LLMs Routing

    cs.CR 2026-04 unverdicted novelty 6.0

    PPRoute achieves plaintext-level LLM routing quality with MPC-based privacy and a 20x speedup over naive encrypted implementations via MPC-friendly encoders, multi-step training, and O(1) communication Top-k search.

  18. RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving

    cs.NI 2026-04 unverdicted novelty 6.0

    Joint resource allocation and routing for multi-model LLM serving can produce up to 87% variation in achievable output quality across setups on the same GPU cluster.

  19. R2-Router: A New Paradigm for LLM Routing with Reasoning

    cs.CL 2026-02 conditional novelty 6.0

    R2-Router jointly selects the LLM and an output-token budget, modeling each model as a quality-cost curve rather than a fixed point, and reports 4-5x cost savings on its new R2-Bench.

  20. RouteLLM: Learning to Route LLMs with Preference Data

    cs.LG 2024-06 unverdicted novelty 6.0

    Router models trained on preference data dynamically select between strong and weak LLMs, cutting inference costs by more than 2x on benchmarks with no quality loss and showing transfer to new model pairs.

  21. RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving

    cs.NI 2026-04 conditional novelty 5.5

    Joint resource allocation and routing for multi-model LLM serving can raise quality under a latency SLO by up to 87% versus fixed-setup routing on the same GPUs.

  22. How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness

    cs.AI 2026-07 conditional novelty 5.0

    Value-weighted LLM routing matches difficulty-only recall while raising precision, exposes within-category calibration collapse, and an elastic value-scaled budget absorbs a synthetic Black Friday surge.

  23. ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries

    cs.LG 2026-06 unverdicted novelty 5.0

    ComplianceGate places a fast encoder classifier before any LLM inference to route PII queries locally and simple queries to small models, reporting 39% latency reduction and 33-52% cost savings on 600 queries.

  24. ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries

    cs.LG 2026-06 unverdicted novelty 5.0

    A classifier before any LLM inference routes PII queries to local endpoints and simple queries to small models, reporting 39% latency reduction and 33-52% cost savings on 600 queries with 99.2% classifier accuracy.

  25. LLMs Show No Signs Of Individuated Metacognition

    cs.LG 2026-05 unverdicted novelty 5.0

    LLM confidence judgments are dominated by a shared difficulty factor across models, with the confidence-performance link collapsing after removing agreed items, yielding no evidence for individuated metacognition.

  26. The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project

    cs.LG 2026-03 unverdicted novelty 5.0

    The Workload-Router-Pool architecture is a 3D framework for LLM inference optimization that synthesizes prior vLLM work into a 3x3 interaction matrix and proposes 21 research directions at the intersections.

  27. vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

    cs.NI 2026-02 conditional novelty 5.0

    vLLM Semantic Router routes LLM requests by composing thirteen signal types into Boolean decision policies, with safety, caching, and model-selection plugin chains.

  28. AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning

    cs.CL 2024-10 unverdicted novelty 5.0

    AdaSwitch improves small local LLM performance on reasoning tasks by adaptively switching to a large cloud LLM upon detected errors, sometimes matching cloud results with far less overhead.

  29. Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

    cs.CL 2025-02 unverdicted novelty 2.0

    A systematic survey of LLM ensemble methods organized into a taxonomy of ensemble-before-inference, ensemble-during-inference, and ensemble-after-inference stages, with review of benchmarks, applications, and future d...