{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:BW2G3CEV2MWUAJVNPW5CPZHPJH","short_pith_number":"pith:BW2G3CEV","schema_version":"1.0","canonical_sha256":"0db46d8895d32d4026ad7dba27e4ef49cda631a2f71680e1403d7d21798a4130","source":{"kind":"arxiv","id":"2404.11531","version":1},"attestation_state":"computed","paper":{"title":"Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Costas Mavromatis, George Karypis, Petros Karypis","submitted_at":"2024-04-17T16:24:07Z","abstract_excerpt":"Fusing knowledge from multiple Large Language Models (LLMs) can combine their diverse strengths to achieve improved performance on a given task. However, current fusion approaches either rely on learning-based fusers that do not generalize to new LLMs, or do not take into account how well each LLM understands the input. In this work, we study LLM fusion at test-time, which enables leveraging knowledge from arbitrary user-specified LLMs during inference. We introduce Pack of LLMs (PackLLM), an effective method for test-time fusion that leverages each LLM's expertise, given an input prompt. Pack"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2404.11531","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-04-17T16:24:07Z","cross_cats_sorted":[],"title_canon_sha256":"5e6c130cd55b7021419577f5c03abd681173fda0052017c56cdbdcf9a11a0144","abstract_canon_sha256":"e48d011cc94f94fff3c11e25ed5632e692c3ddadd63c0dc80619ff816e7fcbe9"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:09:10.109413Z","signature_b64":"oLxeg0mi05Ns1ez5JoaKAbBZjjb0u2K8tiPmJ4u+qTGdqmpwos78JNHmf4RnWDq+7ogXIhrl8rRO/LTSeLvgBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"0db46d8895d32d4026ad7dba27e4ef49cda631a2f71680e1403d7d21798a4130","last_reissued_at":"2026-07-05T08:09:10.108945Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:09:10.108945Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Costas Mavromatis, George Karypis, Petros Karypis","submitted_at":"2024-04-17T16:24:07Z","abstract_excerpt":"Fusing knowledge from multiple Large Language Models (LLMs) can combine their diverse strengths to achieve improved performance on a given task. However, current fusion approaches either rely on learning-based fusers that do not generalize to new LLMs, or do not take into account how well each LLM understands the input. In this work, we study LLM fusion at test-time, which enables leveraging knowledge from arbitrary user-specified LLMs during inference. We introduce Pack of LLMs (PackLLM), an effective method for test-time fusion that leverages each LLM's expertise, given an input prompt. Pack"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2404.11531","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2404.11531/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2404.11531","created_at":"2026-07-05T08:09:10.109001+00:00"},{"alias_kind":"arxiv_version","alias_value":"2404.11531v1","created_at":"2026-07-05T08:09:10.109001+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2404.11531","created_at":"2026-07-05T08:09:10.109001+00:00"},{"alias_kind":"pith_short_12","alias_value":"BW2G3CEV2MWU","created_at":"2026-07-05T08:09:10.109001+00:00"},{"alias_kind":"pith_short_16","alias_value":"BW2G3CEV2MWUAJVN","created_at":"2026-07-05T08:09:10.109001+00:00"},{"alias_kind":"pith_short_8","alias_value":"BW2G3CEV","created_at":"2026-07-05T08:09:10.109001+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":6,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.04378","citing_title":"DLLG: Dynamic Logit-Level Gating of LLM Experts","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00419","citing_title":"Rethinking LLM Ensembling from the Perspective of Mixture Models","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2502.18036","citing_title":"Harnessing Multiple Large Language Models: A Survey on LLM Ensemble","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2507.14200","citing_title":"A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement","ref_index":40,"is_internal_anchor":false},{"citing_arxiv_id":"2510.08592","citing_title":"Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2605.00419","citing_title":"Rethinking LLM Ensembling from the Perspective of Mixture Models","ref_index":11,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH","json":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH.json","graph_json":"https://pith.science/api/pith-number/BW2G3CEV2MWUAJVNPW5CPZHPJH/graph.json","events_json":"https://pith.science/api/pith-number/BW2G3CEV2MWUAJVNPW5CPZHPJH/events.json","paper":"https://pith.science/paper/BW2G3CEV"},"agent_actions":{"view_html":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH","download_json":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH.json","view_paper":"https://pith.science/paper/BW2G3CEV","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2404.11531&json=true","fetch_graph":"https://pith.science/api/pith-number/BW2G3CEV2MWUAJVNPW5CPZHPJH/graph.json","fetch_events":"https://pith.science/api/pith-number/BW2G3CEV2MWUAJVNPW5CPZHPJH/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH/action/timestamp_anchor","attest_storage":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH/action/storage_attestation","attest_author":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH/action/author_attestation","sign_citation":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH/action/citation_signature","submit_replication":"https://pith.science/pith/BW2G3CEV2MWUAJVNPW5CPZHPJH/action/replication_record"}},"created_at":"2026-07-05T08:09:10.109001+00:00","updated_at":"2026-07-05T08:09:10.109001+00:00"}