Pith. sign in

Paper Citation Record · LEDGER

Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2503.08311.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.08311 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:32:31.995104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.514524Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0bcb144d-aa00-4007-b461-f4a2a6e581ba · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.995104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.995104Z digest=sha256:47d32c61c21412dde83294b58252bf1dce3f3910b54e2605ec6692547f57fc78

Observation 3bdb5e06-bfca-4823-a4f2-4f0034616312 · inbound

Toward Robust and Efficient ML-Based GPU Caching for Modern Inference cites this paper.

Toward Robust and Efficient ML-Based GPU Caching for Modern Inference Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:32:40.079227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:29:20.060855Z digest=sha256:72e5c1f25043d848e561461db968c9239814c732cbf2f63be6506c8dc603f861

Observation 96cc46c2-7399-4291-a494-0e0026bc175d · inbound

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective cites this paper.

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:55:33.546662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T00:55:14.213746Z digest=sha256:4fe13c9218500c0971ce376c74a2d09495404b76bd554aacd77442abde2e0c9e

Observation c2dc07d0-e70e-46c8-ba74-b5747fa3dc1f · inbound

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures cites this paper.

Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:36:03.419062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:33:59.777818Z digest=sha256:a4540a9af9464672a2e211a09d61fc7cd46680e447392c43ca00a975b775e14a

Observation 609d1a52-0c12-493d-ab7d-61a783eeaf0b · inbound

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation cites this paper.

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T23:32:59.467653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:32:59.467653Z digest=sha256:0286ef11f6abf63934a899b2464bddbd01e5213d384f4f0473a05bffee8f18d4

Observation f31a7009-429d-4790-a76f-b0aa71b88c9c · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:24.276597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:7a16a0d90460b43e1c65c7104b33ab8ab92500ad5053e427ede3479936ef3eb3

Observation e6bce191-8487-4393-8358-b218e8712af1 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.027448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:f71fc9d4c56cf9fd3b12bd0093bc5e363318349334f17f309ca5ef546f166049

Observation 4e8573d9-622d-44c9-9017-aba2b769e68d · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.541383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:5ca1c4de57934e51e99517af238620ecfd7502df6ba771979cc280087aa0fa0c

Observation 8abd2e5b-9af7-4232-ab69-d537c1ddb71f · inbound

The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference cites this paper.

The Economics of AI Decoding Chips: Rebalancing Compute, Capacity, and Bandwidth for Efficient LLM Inference Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T07:31:12.527547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:31:12.527547Z digest=sha256:b5b7b4d0d0f8fdb2b4df58e3cee341b4b915984b6734197978e6ad2a8aee57d1

Observation f14d4b7f-5e4a-4e19-8587-e289603eafc4 · inbound

Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs cites this paper.

Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T01:39:54.103496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:39:54.103496Z digest=sha256:ddb3db6162ab604449e8a9be6ca0e8bc405923703fc0647c12df2fe394542a59