Pith. sign in

Paper Citation Record · LEDGER

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories

As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2508.08457.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.08457 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:36:25.895941Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f47553a0-a5f2-4537-8721-603c27270229 · outbound

This paper cites FLAT: An optimized dataflow for mitigating attention bottlenecks,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FLAT: An optimized dataflow for mitigating attention bottlenecks,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.793088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.793088Z digest=sha256:c93c89d0c143a3676aad9121c5d01239185aa679d84fe3c6d98cbd4aa65e369d

Observation 0f131db9-5af6-470f-8c10-a70d592c3874 · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.860506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.860506Z digest=sha256:55b93a5ee403c3a7ad9bfeefec33b3f199f7b564ea44fdd3e06607bf03625a73

Observation 5a701258-f509-428e-8b57-655d2572f0a7 · outbound

This paper cites PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:24.928403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:24.928403Z digest=sha256:8ca4a81ecd54b0e001ba6d01d67323e26c39302a9efaf9ee5f2109368d8a1e5f

Observation 089691a3-5ddb-47da-9315-a00db09fd2eb · outbound

This paper cites CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:36:26.158866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T21:36:24.993295Z digest=sha256:ec42d99c8291a04e094c681c36e6646e10b3350479422ffa0a4a604f23bdcc37

Observation d91127b3-f289-493d-af8e-e55852b8fb28 · outbound

This paper cites Timeloop: A systematic approach to DNN accelerator evaluation,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Timeloop: A systematic approach to DNN accelerator evaluation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.127584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.127584Z digest=sha256:cde1af9e400cf1b95cd0c98a74780636c9d8b6fc1f1bbd8ae5e40e6a1cc70a5a

Observation c7ce50ea-adca-467b-bfdf-d3b001847b5e · outbound

This paper cites A-IGZO FETs with High Current and Remarkable Stability for Vertical Channel Transistor(VCT) / 3D DRAM Applications,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories A-IGZO FETs with High Current and Remarkable Stability for Vertical Channel Transistor(VCT) / 3D DRAM Applications,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.239314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.239314Z digest=sha256:9a41f83e51303eddfd46e3bfde1e6f532c2af2c7cca67427bd43a3ba178cc8f3

Observation 99e2900d-b5fd-47d9-8030-c4632dcf7d21 · outbound

This paper cites Integration of 0.75V VDD Oxide-Semiconductor 1T1C Memory with Advanced Logic for An Ultra-Low-Power Low-Latency Cache Solution,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Integration of 0.75V VDD Oxide-Semiconductor 1T1C Memory with Advanced Logic for An Ultra-Low-Power Low-Latency Cache Solution,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T21:36:27.102419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T21:36:25.334829Z digest=sha256:1b08004c5e0da96d5f1ddefd51d6942c0e6bab9115ba97244fd042119a05fe16

Observation c1ae9546-5e4d-4e5b-8d9b-d6b9f42c9ea0 · outbound

This paper cites AMD Next-Generation “Zen 4.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories AMD Next-Generation “Zen 4

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T21:36:26.658354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-05T21:36:25.407922Z digest=sha256:3f25db471b946d186eb3d50dab3608598d44182853e4f7dcb27f634315974baf

Observation dabc142d-8cc0-4cbc-9102-09e6060fb135 · outbound

This paper cites Taming throughput-latency tradeoff in LLM inference with Sarathi‑Serve,.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Taming throughput-latency tradeoff in LLM inference with Sarathi‑Serve,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.462008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.462008Z digest=sha256:7dd933c049f2c459720df7fc88b6aa520e141ad7b8642b24562bab9d87b0ae68

Observation 53b1e519-fb93-4279-bff8-0706bb83a058 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.580836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.580836Z digest=sha256:11e7fb93bfa70585c101a812759d420f28f1b7216a8ef4b55fca7014455ab254

Observation c5f0cfb0-89f4-4d86-ae1b-1a65cf08a766 · outbound

This paper cites The Llama 3 Herd of Models.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.698784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.698784Z digest=sha256:4abd8187fad1551c487fd23a56d2608199cbdc4f5440663abcb9d6bf330d1f46

Observation 4610b7e7-14a7-40ff-a59a-8da2b5758fe4 · outbound

This paper cites GPT-4 Technical Report.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.814551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.814551Z digest=sha256:e97a34f6f03597312cf82b82e18707b3556414c2e708226594d0fee0afba4939

Observation b34dcaa5-f101-49a5-8375-06d237f6b7f2 · outbound

This paper cites an unresolved cited work.

Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T21:36:25.895941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:36:25.895941Z digest=sha256:67d1d43b88e6aa67977986b10b3ae441f5343e8169af3932e4d3e79ef2bac6a7

Pith citing papers

No inbound Pith citation observations are available.