Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:36:25.895941Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2508.08457.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:36:25.895941Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f47553a0-a5f2-4537-8721-603c27270229 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FLAT: An optimized dataflow for mitigating attention bottlenecks,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f131db9-5af6-470f-8c10-a70d592c3874 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a701258-f509-428e-8b57-655d2572f0a7 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089691a3-5ddb-47da-9315-a00db09fd2eb · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories CMOS+X: Stacking Persistent Embedded Memories based on Oxide Transistors upon GPGPU Platforms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d91127b3-f289-493d-af8e-e55852b8fb28 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Timeloop: A systematic approach to DNN accelerator evaluation,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7ce50ea-adca-467b-bfdf-d3b001847b5e · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories A-IGZO FETs with High Current and Remarkable Stability for Vertical Channel Transistor(VCT) / 3D DRAM Applications,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99e2900d-b5fd-47d9-8030-c4632dcf7d21 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Integration of 0.75V VDD Oxide-Semiconductor 1T1C Memory with Advanced Logic for An Ultra-Low-Power Low-Latency Cache Solution,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1ae9546-5e4d-4e5b-8d9b-d6b9f42c9ea0 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories AMD Next-Generation “Zen 4
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dabc142d-8cc0-4cbc-9102-09e6060fb135 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Taming throughput-latency tradeoff in LLM inference with Sarathi‑Serve,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b1e519-fb93-4279-bff8-0706bb83a058 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5f0cfb0-89f4-4d86-ae1b-1a65cf08a766 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4610b7e7-14a7-40ff-a59a-8da2b5758fe4 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories GPT-4 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34dcaa5-f101-49a5-8375-06d237f6b7f2 · outbound
Architecting Long-Context LLM Acceleration with Packing-Prefetch Scheduler and Ultra-Large Capacity On-Chip Memories Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.