Pith. sign in

Paper Citation Record · LEDGER

Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.08311.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.08311 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:31.995104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.514524Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0bcb144d-aa00-4007-b461-f4a2a6e581ba · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.995104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.995104Z digest=sha256:c63e598950bc63ec910c27d767851561fa71b64a776cedf40e25511bfc72bd4a

Observation 3bdb5e06-bfca-4823-a4f2-4f0034616312 · inbound

Toward Robust and Efficient ML-Based GPU Caching for Modern Inference cites this paper.

Toward Robust and Efficient ML-Based GPU Caching for Modern Inference Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:32:40.079227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:29:20.060855Z digest=sha256:be6da505ad6458cfd6808c12dd60c574ec437e24a9957ba690aec3fa6e0cdd76

Observation 96cc46c2-7399-4291-a494-0e0026bc175d · inbound

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective cites this paper.

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:33.546662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T00:55:14.213746Z digest=sha256:c0877485626611bce776f8e0c4aadc15a9458198de23f4683a37aac8503a7997

Observation c2dc07d0-e70e-46c8-ba74-b5747fa3dc1f · inbound

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures cites this paper.

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.419062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:33:59.777818Z digest=sha256:8afb0e1f3fd4797101d0f7613b922edf0962878e205044c28de6eb6a5e5bcf18

Observation 609d1a52-0c12-493d-ab7d-61a783eeaf0b · inbound

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation cites this paper.

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T23:32:59.467653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:32:59.467653Z digest=sha256:1a779fdd7ed01b57d780cc7c0958e1418a147accb7fb6c080e58a6694ff48793

Observation f31a7009-429d-4790-a76f-b0aa71b88c9c · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:24.276597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:bf8a11de328d729a080b1409deb4e86fd4ee2b42f2a213f6a204e604c87deaf1

Observation e6bce191-8487-4393-8358-b218e8712af1 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.027448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:0b48f4dc7d41ce8ca6f4771662c41844bbaa91bba8155a75f86ad116918921b4

Observation 4e8573d9-622d-44c9-9017-aba2b769e68d · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.541383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:b6de695bb935a1acc0604e78f96b4bb698b9f644e15df95c03c0c8c1691d6116

Observation 8abd2e5b-9af7-4232-ab69-d537c1ddb71f · inbound

The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference cites this paper.

The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:31:12.527547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:31:12.527547Z digest=sha256:7c66f8316030ef59e336b2cf8ba4daebc701e0b4eb1910ab76447719a57d4895

Observation f14d4b7f-5e4a-4e19-8587-e289603eafc4 · inbound

Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs cites this paper.

Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T01:39:54.103496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:39:54.103496Z digest=sha256:4206f9780a699e8002d771d0151853859afaa2dac13b229c50e198ce98c69173