{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:FXUUOTXP4MLGOJCKWHD3SCVXUZ","short_pith_number":"pith:FXUUOTXP","schema_version":"1.0","canonical_sha256":"2de9474eefe31667244ab1c7b90ab7a67c9138f8f5dce852c9e2ba7556da3c53","source":{"kind":"arxiv","id":"2402.18158","version":2},"attestation_state":"computed","paper":{"title":"Evaluating Quantized Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Guohao Dai, Huazhong Yang, Luning Wang, Shengen Yan, Shiyao Li, Tengxuan Liu, Xiangsheng Shi, Xuefei Ning, Yu Wang","submitted_at":"2024-02-28T08:43:05Z","abstract_excerpt":"Post-training quantization (PTQ) has emerged as a promising technique to reduce the cost of large language models (LLMs). Specifically, PTQ can effectively mitigate memory consumption and reduce computational overhead in LLMs. To meet the requirements of both high efficiency and performance across diverse scenarios, a comprehensive evaluation of quantized LLMs is essential to guide the selection of quantization methods. This paper presents a thorough evaluation of these factors by evaluating the effect of PTQ on Weight, Activation, and KV Cache on 11 model families, including OPT, LLaMA2, Falc"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.18158","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-02-28T08:43:05Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"098e1e148300249dc17acb40a7f9253bfcac8bc6d53b05c621938a7f6b0c40b7","abstract_canon_sha256":"b8eebccade26fcb49b65451df3e9b99979b7d0dae97fb3f362bbd6bd7295a1f8"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:28:07.286822Z","signature_b64":"fsl6RpDP6FxzkgrTjBMaAxmAfW07y4XtK+oc08WoCOklsb+JGZC5sjZ3+wUk1c0yUpqXvXgrWGkTOF296xStBQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"2de9474eefe31667244ab1c7b90ab7a67c9138f8f5dce852c9e2ba7556da3c53","last_reissued_at":"2026-07-05T08:28:07.286275Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:28:07.286275Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Evaluating Quantized Large Language Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Guohao Dai, Huazhong Yang, Luning Wang, Shengen Yan, Shiyao Li, Tengxuan Liu, Xiangsheng Shi, Xuefei Ning, Yu Wang","submitted_at":"2024-02-28T08:43:05Z","abstract_excerpt":"Post-training quantization (PTQ) has emerged as a promising technique to reduce the cost of large language models (LLMs). Specifically, PTQ can effectively mitigate memory consumption and reduce computational overhead in LLMs. To meet the requirements of both high efficiency and performance across diverse scenarios, a comprehensive evaluation of quantized LLMs is essential to guide the selection of quantization methods. This paper presents a thorough evaluation of these factors by evaluating the effect of PTQ on Weight, Activation, and KV Cache on 11 model families, including OPT, LLaMA2, Falc"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.18158","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.18158/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.18158","created_at":"2026-07-05T08:28:07.286346+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.18158v2","created_at":"2026-07-05T08:28:07.286346+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.18158","created_at":"2026-07-05T08:28:07.286346+00:00"},{"alias_kind":"pith_short_12","alias_value":"FXUUOTXP4MLG","created_at":"2026-07-05T08:28:07.286346+00:00"},{"alias_kind":"pith_short_16","alias_value":"FXUUOTXP4MLGOJCK","created_at":"2026-07-05T08:28:07.286346+00:00"},{"alias_kind":"pith_short_8","alias_value":"FXUUOTXP","created_at":"2026-07-05T08:28:07.286346+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.03002","citing_title":"Perplexity Can Miss SAE Feature Damage Under Quantization","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2409.00084","citing_title":"Vision-Language and Large Language Model Performance in Gastroenterology: GPT, Claude, Llama, Phi, Mistral, Gemma, and Quantized Models","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2411.10656","citing_title":"Precision or Peril: A PoC of Python Code Quality from Quantized Large Language Models","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2504.12334","citing_title":"QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2404.14294","citing_title":"A Survey on Efficient Inference for Large Language Models","ref_index":214,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19342","citing_title":"Are Large Language Models Economically Viable for Industry Deployment?","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2604.19884","citing_title":"From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization","ref_index":3,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ","json":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ.json","graph_json":"https://pith.science/api/pith-number/FXUUOTXP4MLGOJCKWHD3SCVXUZ/graph.json","events_json":"https://pith.science/api/pith-number/FXUUOTXP4MLGOJCKWHD3SCVXUZ/events.json","paper":"https://pith.science/paper/FXUUOTXP"},"agent_actions":{"view_html":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ","download_json":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ.json","view_paper":"https://pith.science/paper/FXUUOTXP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.18158&json=true","fetch_graph":"https://pith.science/api/pith-number/FXUUOTXP4MLGOJCKWHD3SCVXUZ/graph.json","fetch_events":"https://pith.science/api/pith-number/FXUUOTXP4MLGOJCKWHD3SCVXUZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ/action/storage_attestation","attest_author":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ/action/author_attestation","sign_citation":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ/action/citation_signature","submit_replication":"https://pith.science/pith/FXUUOTXP4MLGOJCKWHD3SCVXUZ/action/replication_record"}},"created_at":"2026-07-05T08:28:07.286346+00:00","updated_at":"2026-07-05T08:28:07.286346+00:00"}