{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:GCASEVD4BCR2GUYXEE4RWCNQO5","short_pith_number":"pith:GCASEVD4","schema_version":"1.0","canonical_sha256":"308122547c08a3a3531721391b09b07761722e06a64c744e1386eab3ddef2c39","source":{"kind":"arxiv","id":"2410.04466","version":4},"attestation_state":"computed","paper":{"title":"Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.AR","authors_text":"Guohao Dai, Hao Zhou, Jiaming Xu, Jiayi Pan, Jinhao Li, Jun Liu, Li Ding, Shan Huang, Wen Li, Yaoxiu Lian, Yonghua Chen, Yu Wang","submitted_at":"2024-10-06T12:42:04Z","abstract_excerpt":"Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation. Compared to non-generative LLMs like BERT and DeBERTa, generative LLMs like GPT series and Llama series are currently the main focus due to their superior algorithmic performance. The advancements in generative LLMs are closely intertwined with the development of hardware capabilities. Various hardware platforms exhibit distinct hardware characteristics, which can help improve LLM inference performance. Therefore, this paper comprehensively surve"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2410.04466","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.AR","submitted_at":"2024-10-06T12:42:04Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"810adfd96f7a4c4b8ddc930ca36d5d8b410a9a85177d000fa84bde452b74558f","abstract_canon_sha256":"a559033a8732e3aa053656927b4db8d465beb818b2d56f45b534672f61749fff"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:20:46.841035Z","signature_b64":"KoW8E3XohzIoFChKrwe2P9xy1RSKX55eUv7giMrKKUmXvd9jcv2SS9XHQgf5rwYgQNzyiLukQE0xiWpWSypzBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"308122547c08a3a3531721391b09b07761722e06a64c744e1386eab3ddef2c39","last_reissued_at":"2026-07-05T11:20:46.840453Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:20:46.840453Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.LG"],"primary_cat":"cs.AR","authors_text":"Guohao Dai, Hao Zhou, Jiaming Xu, Jiayi Pan, Jinhao Li, Jun Liu, Li Ding, Shan Huang, Wen Li, Yaoxiu Lian, Yonghua Chen, Yu Wang","submitted_at":"2024-10-06T12:42:04Z","abstract_excerpt":"Large Language Models (LLMs) have demonstrated remarkable capabilities across various fields, from natural language understanding to text generation. Compared to non-generative LLMs like BERT and DeBERTa, generative LLMs like GPT series and Llama series are currently the main focus due to their superior algorithmic performance. The advancements in generative LLMs are closely intertwined with the development of hardware capabilities. Various hardware platforms exhibit distinct hardware characteristics, which can help improve LLM inference performance. Therefore, this paper comprehensively surve"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2410.04466","kind":"arxiv","version":4},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2410.04466/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2410.04466","created_at":"2026-07-05T11:20:46.840519+00:00"},{"alias_kind":"arxiv_version","alias_value":"2410.04466v4","created_at":"2026-07-05T11:20:46.840519+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2410.04466","created_at":"2026-07-05T11:20:46.840519+00:00"},{"alias_kind":"pith_short_12","alias_value":"GCASEVD4BCR2","created_at":"2026-07-05T11:20:46.840519+00:00"},{"alias_kind":"pith_short_16","alias_value":"GCASEVD4BCR2GUYX","created_at":"2026-07-05T11:20:46.840519+00:00"},{"alias_kind":"pith_short_8","alias_value":"GCASEVD4","created_at":"2026-07-05T11:20:46.840519+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":12,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.19119","citing_title":"PuDGhost: Experimental Analysis of Computation Result Corruption in Processing-using-DRAM Operations on Real DRAM Chips and Implications for Future Systems","ref_index":211,"is_internal_anchor":false},{"citing_arxiv_id":"2607.01313","citing_title":"Black-Box Inference of LLM Architectural Properties with Restrictive API Access","ref_index":26,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09080","citing_title":"Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy","ref_index":58,"is_internal_anchor":false},{"citing_arxiv_id":"2509.22244","citing_title":"FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing","ref_index":17,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19537","citing_title":"The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.19537","citing_title":"The Silent Hyperparameter: Quantifying the Impact of Inference Backends on LLM Reproducibility","ref_index":12,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10195","citing_title":"Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2605.10195","citing_title":"Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration","ref_index":29,"is_internal_anchor":false},{"citing_arxiv_id":"2604.22935","citing_title":"Secure eFPGA-Enabled Edge LLM Inference: Architectural and Hardware Countermeasures","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08749","citing_title":"A Little Rank Goes a Long Way: Random Scaffolds with LoRA Adapters Are All You Need","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2605.07750","citing_title":"Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling","ref_index":1,"is_internal_anchor":false},{"citing_arxiv_id":"2604.21952","citing_title":"Focus Session: Hardware and Software Techniques for Accelerating Multimodal Foundation Models","ref_index":47,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5","json":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5.json","graph_json":"https://pith.science/api/pith-number/GCASEVD4BCR2GUYXEE4RWCNQO5/graph.json","events_json":"https://pith.science/api/pith-number/GCASEVD4BCR2GUYXEE4RWCNQO5/events.json","paper":"https://pith.science/paper/GCASEVD4"},"agent_actions":{"view_html":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5","download_json":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5.json","view_paper":"https://pith.science/paper/GCASEVD4","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2410.04466&json=true","fetch_graph":"https://pith.science/api/pith-number/GCASEVD4BCR2GUYXEE4RWCNQO5/graph.json","fetch_events":"https://pith.science/api/pith-number/GCASEVD4BCR2GUYXEE4RWCNQO5/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5/action/timestamp_anchor","attest_storage":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5/action/storage_attestation","attest_author":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5/action/author_attestation","sign_citation":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5/action/citation_signature","submit_replication":"https://pith.science/pith/GCASEVD4BCR2GUYXEE4RWCNQO5/action/replication_record"}},"created_at":"2026-07-05T11:20:46.840519+00:00","updated_at":"2026-07-05T11:20:46.840519+00:00"}