{"work":{"id":"1b10f2a9-a178-4d23-97fb-8db2354c7e6c","openalex_id":"https://openalex.org/W4387321503","doi":"10.1145/3600006.3613163","arxiv_id":"0006.361316","raw_key":null,"title":"Efficient memory management for large language model serving with pagedattention,","authors":null,"authors_text":"W","year":2023,"venue":null,"abstract":null,"external_url":"https://arxiv.org/abs/0006.361316","cited_by_count":31,"metadata_source":"arxiv_reference","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":null,"created_at":"2026-05-09T19:05:10.258697+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":false,"display_title":"Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =","render_title":"Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle ="},"hub":{"state":{"work_id":"1b10f2a9-a178-4d23-97fb-8db2354c7e6c","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":158,"external_cited_by_count":31,"distinct_field_count":20,"first_pith_cited_at":"2025-04-14T00:29:49+00:00","last_pith_cited_at":"2026-07-09T08:12:15+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T17:49:17.959126+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":17},{"context_role":"method","n":5},{"context_role":"baseline","n":1},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":16},{"context_polarity":"use_method","n":5},{"context_polarity":"baseline","n":1},{"context_polarity":"unclear","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =","claims":[{"claim_text":"evaluated using regular and small-sized models to assess the impact of compressor size on latency and to ensure feasibility on a consumer-grade MacBook Pro with an M1 Pro processor. For the target models, we chose the 7B parameter model from Mistral [9] as well as the 70B LLaMA 3.1 model [4]. Both models were served either locally using the inference frameworks Hugging Face Transformers (HF-TF) [22] or vLLM [13], or were accessed through an API persistent model server, both hosted on our Nvidia ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"throughput, communication overhead, and GPU load balance. For latency-sensitive online workloads (e.g., LLM serving), Tessera employs a latency-oriented scheduling policy (§III-B). Meanwhile, Tessera incorporates a lightweight online monitor to dynamically adjust scheduling policies under changing request loads (§III-D). We integrate Tessera into popular vLLM [16] and PyTorch [17] frameworks, and evaluate Tessera across five heteroge- neous GPUs and four model architectures, scaling up to 16 GPU","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"is another dimension of tensor program optimizations, where optimizers exploit algebraic identities to transform compute expressions in the program and improve performance. Ex- pression rewriting often happens at graph level, such as in TASO [17], TenSat [44], and XLA's algebraic simplifier [28], where expressions describe computation over whole tensors. Existing compilers and languages such as TVM, MLIR [20] and Lift [36] allow expression rewriting at lower levels. Nau- tilus's IR design enable","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Existing weight transfer approaches fall under three main categories:collective communication,point-to-point commu- nication, anddistributed storage. Table 1 summarizes how well each meets the above goals. Collective Communication.NVIDIA Collective Communi- cation Library (NCCL) [23] is the most widely used approach for weight transfer, adopted by veRL [34], OpenRLHF [12], AReaL [7], vLLM [16], SGLang [43], Slime [48], among oth- ers. In these frameworks, trainers broadcast weights to roll- outs","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"erNorm toQandKafter the linear projection. The QK normalization controls the magnitude ofQandK, so that the dot-product attention logitsQK T remain numerically stable. Vdoes not participate in the normalization, becauseVis built by weighted-summation of the resulting attention scores, and normalizingVis meaningless. Rotary position embeddings (RoPE).RoPE [15] encodes positional information by rotatingQandKso that the dot productq T i kj depends on relative positioni−j, whereq i andk j are the qu","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"In Stage II, we use a lightweight proxy A2C trained for K2 = 8,000 steps across G2 = 16 refinement rounds. In Stage III, we train full policies for K3 = 10 7 steps over G3 = 3 rounds. All methods use identical CrowdNav policy architectures [ 28] with PPO optimization, implemented in PyTorch on NVIDIA A6000 GPUs. We use the open-source gpt-oss-120B [1] as the LLM served locally via vLLM [19]. Critically, EvoNav, Eureka, and CARD in Table 1 differ only in their reward functions-all use the same Cr","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle = because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":5,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-26T20:05:23.403432+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[]},"error":null,"updated_at":"2026-06-26T20:05:23.397793+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T16:12:13.391726+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":18},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":14},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":6},{"title":"Splitwise: Eﬀicient generative llm inference using phase splitting","work_id":"6a69794e-e9cd-48c9-9b7d-502b131ff6db","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Triton: an intermediate language and compiler for tiled neural network computations , year =","work_id":"36502ac3-807e-4ab8-ac44-c87017160e32","shared_citers":5},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":4},{"title":"Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference","work_id":"b41eb6db-193f-48ba-8278-6b1698524755","shared_citers":3},{"title":"Aegaeon: Effective gpu pooling for concurrent llm serving on the market","work_id":"bcb0b297-69ef-45fb-96b9-add34313508c","shared_citers":3},{"title":"Compressing context to enhance inference efficiency of large language models","work_id":"6929d93f-92e7-4fc0-9681-7de781846d42","shared_citers":3},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":3},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":3},{"title":"Efficient Streaming Language Models with Attention Sinks","work_id":"a8d25452-c237-48c9-88a4-682717c3979a","shared_citers":3},{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":3},{"title":"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness","work_id":"efa96825-0830-4cfc-a250-fdaf6af302ab","shared_citers":3},{"title":"Flashin- fer: Efficient and customizable attention engine for llm inference serving","work_id":"a153885b-2460-4177-9053-8d0011adfcb9","shared_citers":3},{"title":"Flux: Fast software-based communication overlap on gpus through kernel fusion.arXiv preprint arXiv:2406.06858","work_id":"5d0e6adc-6dd2-49cc-8551-dc00433ed79f","shared_citers":3},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":3},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":3},{"title":"HybridFlow: A Flexible and Efficient RLHF Framework , url=","work_id":"4909736c-3fe1-4820-b489-cca51669c6d2","shared_citers":3}],"time_series":[{"n":44,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T16:22:26.099843+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T16:12:08.013428+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =","claims":[{"claim_text":"evaluated using regular and small-sized models to assess the impact of compressor size on latency and to ensure feasibility on a consumer-grade MacBook Pro with an M1 Pro processor. For the target models, we chose the 7B parameter model from Mistral [9] as well as the 70B LLaMA 3.1 model [4]. Both models were served either locally using the inference frameworks Hugging Face Transformers (HF-TF) [22] or vLLM [13], or were accessed through an API persistent model server, both hosted on our Nvidia ","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"throughput, communication overhead, and GPU load balance. For latency-sensitive online workloads (e.g., LLM serving), Tessera employs a latency-oriented scheduling policy (§III-B). Meanwhile, Tessera incorporates a lightweight online monitor to dynamically adjust scheduling policies under changing request loads (§III-D). We integrate Tessera into popular vLLM [16] and PyTorch [17] frameworks, and evaluate Tessera across five heteroge- neous GPUs and four model architectures, scaling up to 16 GPU","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"is another dimension of tensor program optimizations, where optimizers exploit algebraic identities to transform compute expressions in the program and improve performance. Ex- pression rewriting often happens at graph level, such as in TASO [17], TenSat [44], and XLA's algebraic simplifier [28], where expressions describe computation over whole tensors. Existing compilers and languages such as TVM, MLIR [20] and Lift [36] allow expression rewriting at lower levels. Nau- tilus's IR design enable","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Existing weight transfer approaches fall under three main categories:collective communication,point-to-point commu- nication, anddistributed storage. Table 1 summarizes how well each meets the above goals. Collective Communication.NVIDIA Collective Communi- cation Library (NCCL) [23] is the most widely used approach for weight transfer, adopted by veRL [34], OpenRLHF [12], AReaL [7], vLLM [16], SGLang [43], Slime [48], among oth- ers. In these frameworks, trainers broadcast weights to roll- outs","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"erNorm toQandKafter the linear projection. The QK normalization controls the magnitude ofQandK, so that the dot-product attention logitsQK T remain numerically stable. Vdoes not participate in the normalization, becauseVis built by weighted-summation of the resulting attention scores, and normalizingVis meaningless. Rotary position embeddings (RoPE).RoPE [15] encodes positional information by rotatingQandKso that the dot productq T i kj depends on relative positioni−j, whereq i andk j are the qu","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"In Stage II, we use a lightweight proxy A2C trained for K2 = 8,000 steps across G2 = 16 refinement rounds. In Stage III, we train full policies for K3 = 10 7 steps over G3 = 3 rounds. All methods use identical CrowdNav policy architectures [ 28] with PPO optimization, implemented in PyTorch on NVIDIA A6000 GPUs. We use the open-source gpt-oss-120B [1] as the LLM served locally via vLLM [19]. Critically, EvoNav, Eureka, and CARD in Table 1 differ only in their reward functions-all use the same Cr","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle = because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (16 contexts).","role_counts":[{"n":16,"context_role":"background"},{"n":5,"context_role":"method"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-06-26T20:05:23.401194+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Eﬀicient memory management for large language model serving with pagedattention","claims":[],"why_cited":"Pith tracks Eﬀicient memory management for large language model serving with pagedattention because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T16:22:26.106281+00:00"}},"summary":{"title":"Eﬀicient memory management for large language model serving with pagedattention","claims":[],"why_cited":"Pith tracks Eﬀicient memory management for large language model serving with pagedattention because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":18},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":14},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":6},{"title":"Splitwise: Eﬀicient generative llm inference using phase splitting","work_id":"6a69794e-e9cd-48c9-9b7d-502b131ff6db","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Triton: an intermediate language and compiler for tiled neural network computations , year =","work_id":"36502ac3-807e-4ab8-ac44-c87017160e32","shared_citers":5},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":4},{"title":"Ada-KV: Optimizing KV cache eviction by adaptive budget allocation for efficient LLM inference","work_id":"b41eb6db-193f-48ba-8278-6b1698524755","shared_citers":3},{"title":"Aegaeon: Effective gpu pooling for concurrent llm serving on the market","work_id":"bcb0b297-69ef-45fb-96b9-add34313508c","shared_citers":3},{"title":"Compressing context to enhance inference efficiency of large language models","work_id":"6929d93f-92e7-4fc0-9681-7de781846d42","shared_citers":3},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":3},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":3},{"title":"Efficient Streaming Language Models with Attention Sinks","work_id":"a8d25452-c237-48c9-88a4-682717c3979a","shared_citers":3},{"title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":3},{"title":"FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness","work_id":"efa96825-0830-4cfc-a250-fdaf6af302ab","shared_citers":3},{"title":"Flashin- fer: Efficient and customizable attention engine for llm inference serving","work_id":"a153885b-2460-4177-9053-8d0011adfcb9","shared_citers":3},{"title":"Flux: Fast software-based communication overlap on gpus through kernel fusion.arXiv preprint arXiv:2406.06858","work_id":"5d0e6adc-6dd2-49cc-8551-dc00433ed79f","shared_citers":3},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":3},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":3},{"title":"HybridFlow: A Flexible and Efficient RLHF Framework , url=","work_id":"4909736c-3fe1-4820-b489-cca51669c6d2","shared_citers":3}],"time_series":[{"n":44,"year":2026}],"dependency_candidates":[]},"authors":[]}}