{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2023:B2RBQZSDCFBPUQV5S5XF4NYYIZ","short_pith_number":"pith:B2RBQZSD","schema_version":"1.0","canonical_sha256":"0ea21866431142fa42bd976e5e3718466e3578c14027f3804b7a7753ce57bdc2","source":{"kind":"arxiv","id":"2310.12963","version":5},"attestation_state":"computed","paper":{"title":"AutoMix: Automatically Mixing Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Aditya Gupta, Aman Madaan, Ankit Anand, Dheeraj Rajagopal, Karthik Kappaganthu, Manaal Faruqui, Mausam, Pei Zhou, Pranjal Aggarwal, Shyam Upadhyay, Srividya Pranavi Potharaju, Swaroop Mishra, Yiming Yang","submitted_at":"2023-10-19T17:57:39Z","abstract_excerpt":"Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present Automix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to Automix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2310.12963","kind":"arxiv","version":5},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2023-10-19T17:57:39Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"c40aa915b32deab7f9d394fc461ade1fb403c3f34512fce459d325945b87c829","abstract_canon_sha256":"e019c9fe4e5ab20125dbabfb74bcd80a9033a03ac59c7e4c7f1935ed51ef3e42"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T10:02:41.795163Z","signature_b64":"rqMM3A7mzHLDzTaH/PowFWrvRQqRfPvnuaAVdz3SrMukf0arsegX7DKn6/3dVmELHzpM4aJvdxRyGukQRrZpCQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0ea21866431142fa42bd976e5e3718466e3578c14027f3804b7a7753ce57bdc2","last_reissued_at":"2026-07-05T10:02:41.794715Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T10:02:41.794715Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"AutoMix: Automatically Mixing Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Aditya Gupta, Aman Madaan, Ankit Anand, Dheeraj Rajagopal, Karthik Kappaganthu, Manaal Faruqui, Mausam, Pei Zhou, Pranjal Aggarwal, Shyam Upadhyay, Srividya Pranavi Potharaju, Swaroop Mishra, Yiming Yang","submitted_at":"2023-10-19T17:57:39Z","abstract_excerpt":"Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present Automix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to Automix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2310.12963","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2310.12963/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2310.12963","created_at":"2026-07-05T10:02:41.794772+00:00"},{"alias_kind":"arxiv_version","alias_value":"2310.12963v5","created_at":"2026-07-05T10:02:41.794772+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2310.12963","created_at":"2026-07-05T10:02:41.794772+00:00"},{"alias_kind":"pith_short_12","alias_value":"B2RBQZSDCFBP","created_at":"2026-07-05T10:02:41.794772+00:00"},{"alias_kind":"pith_short_16","alias_value":"B2RBQZSDCFBPUQV5","created_at":"2026-07-05T10:02:41.794772+00:00"},{"alias_kind":"pith_short_8","alias_value":"B2RBQZSD","created_at":"2026-07-05T10:02:41.794772+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":22,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.27288","citing_title":"When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models","ref_index":25,"is_internal_anchor":false},{"citing_arxiv_id":"2606.21929","citing_title":"Selective Ensemble Based on Preference-Directed Multi-Objective Bandits","ref_index":54,"is_internal_anchor":false},{"citing_arxiv_id":"2606.12402","citing_title":"DIRECT: When and Where Should You Allocate Test-Time Compute in Embodied Planners?","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31163","citing_title":"ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2607.00053","citing_title":"SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks","ref_index":55,"is_internal_anchor":false},{"citing_arxiv_id":"2606.31163","citing_title":"ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2605.24299","citing_title":"LLMs Show No Signs Of Individuated Metacognition","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2606.30067","citing_title":"Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management","ref_index":19,"is_internal_anchor":false},{"citing_arxiv_id":"2410.13181","citing_title":"AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning","ref_index":16,"is_internal_anchor":false},{"citing_arxiv_id":"2502.18036","citing_title":"Harnessing Multiple Large Language Models: A Survey on LLM Ensemble","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17106","citing_title":"HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19099","citing_title":"DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2603.21354","citing_title":"The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project","ref_index":93,"is_internal_anchor":false},{"citing_arxiv_id":"2605.11301","citing_title":"LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10805","citing_title":"Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.25591","citing_title":"Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models","ref_index":81,"is_internal_anchor":false},{"citing_arxiv_id":"2406.18665","citing_title":"RouteLLM: Learning to Route LLMs with Preference Data","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00334","citing_title":"AgentFloor: How Far Up the tool use Ladder Can Small Open-Weight Models Go?","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2604.10907","citing_title":"RouterWise: Joint Resource Allocation and Routing for Latency-Aware Multi-Model LLM Serving","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15022","citing_title":"Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization","ref_index":3,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15728","citing_title":"Privacy-Preserving LLMs Routing","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2605.05007","citing_title":"Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation","ref_index":2,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ","json":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ.json","graph_json":"https://pith.science/api/pith-number/B2RBQZSDCFBPUQV5S5XF4NYYIZ/graph.json","events_json":"https://pith.science/api/pith-number/B2RBQZSDCFBPUQV5S5XF4NYYIZ/events.json","paper":"https://pith.science/paper/B2RBQZSD"},"agent_actions":{"view_html":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ","download_json":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ.json","view_paper":"https://pith.science/paper/B2RBQZSD","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2310.12963&json=true","fetch_graph":"https://pith.science/api/pith-number/B2RBQZSDCFBPUQV5S5XF4NYYIZ/graph.json","fetch_events":"https://pith.science/api/pith-number/B2RBQZSDCFBPUQV5S5XF4NYYIZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ/action/storage_attestation","attest_author":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ/action/author_attestation","sign_citation":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ/action/citation_signature","submit_replication":"https://pith.science/pith/B2RBQZSDCFBPUQV5S5XF4NYYIZ/action/replication_record"}},"created_at":"2026-07-05T10:02:41.794772+00:00","updated_at":"2026-07-05T10:02:41.794772+00:00"}