Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T02:11:26.925234Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.19775.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T02:11:26.925234Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T01:55:09.658053Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T01:57:51.122520Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc66ec53-c92b-4c92-ac79-50a1fe7f11c5 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Vidur: A large-scale simulation frame- work for llm inference
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 38c70ef2-8da4-4000-9549-73029ad015d9 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve}
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2dd0a23f-066c-4141-9f14-2f9b8061721e · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f45074d-80ae-42d8-aad2-6c39e26f2718 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 854ab92f-9190-4b23-90fb-8d4efb86f63c · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm in a flash: Efficient large language model inference with limited memory
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0092719e-4193-4989-9453-d5ad9dacc9cc · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Exploiting cxl-based memory for distributed deep learning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb609ecb-212b-4d90-8ee7-f09461b3bce5 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Accelerating performance of gpu-based workloads using cxl
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2847a015-37f3-497a-93af-7dff61efb162 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Moe-lightning: High-throughput moe inference on memory-constrained gpus
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c290b306-54b3-42df-be41-c345b222bbce · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Lmcache: An efficient kv cache layer for enterprise-scale llm inference
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f109b5d0-54b4-4d5d-bbdc-be32baacb865 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm-inference- bench: Inference benchmarking of large language models on ai acceler- ators
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 589e859c-81ef-40aa-8b51-2161b16c942d · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles PagedEviction: Structured block-wise KV cache pruning for efficient large language model inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2a9a8dd0-c925-4770-b78a-d4aefb86c975 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Multi-Head Attention: Collaborate Instead of Concatenate
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f9307ee1-4d2d-4b79-8bff-1e65af4f311d · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Compute express link
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6d6d5136-e4d2-4ce1-941c-7d307fca72c2 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Corsair™. built for generative a
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation caa864e8-407d-4b4c-9f34-8013f04aa19c · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Available: https://www.d-matrix.ai/product/
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b9e4e9d7-a603-4429-9d15-7a292149cec6 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Why we decoupled execution to accelerate i/o
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 67452545-d618-444e-95e7-db09289c1fa3 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1bce6f5a-fda1-45f4-a32e-694187999c6a · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Accelerating LLM inference throughput via asynchronous KV cache prefetching
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4d9900ce-7436-4e07-95fb-61b603f96a7b · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ac2906b-e0cf-4bec-a6ad-817316c66a7f · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Nvidia dynamo
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3537f3d1-3d01-4b54-acdf-9fd19ef8cb21 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Gpipe: Efficient training of giant neural networks using pipeline parallelism
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f324fd4c-f6a8-48c0-913d-169da9adea20 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Calculon: a methodology and tool for high-level co-design of systems and large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c70d84ad-dc6d-4f3c-b67c-4eed754daada · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Efficient memory management for large language model serving with pagedattention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ed0f790-0cce-414c-85b2-f320cdaec115 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm inference serving: Survey of recent advances and opportunities
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 32936c06-218e-4973-a7ac-2cc489621af9 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles A Survey on Large Language Model Acceleration based on KV Cache Management
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d06ced6-091a-4814-acc9-9100f6319475 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Deepseek-v3 technical report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 23e374e1-46ba-46fc-a19a-a27bd1e04bdc · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Minicache: Kv cache compression in depth dimension for large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e2754aa6-4304-4c72-8865-4a43624110c5 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mlp-offload: Multi-level, multi-path offloading for llm pre-training to break the gpu memory wall
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eec82d19-bac2-4f3a-90a9-3fc0ccc63dce · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Openai o1 system card
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c14ce41-d42b-4b04-b685-b6e78eb71740 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles A survey on inference engines for large language models: Perspectives on optimization and efficiency
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 46755b82-94d7-4a76-b492-25849a04a9f3 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mooncake: A kvcache-centric disaggregated architecture for llm serving
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6843e24d-0796-440a-9bb1-dd6ec58d8003 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Prophet: An llm infer- ence engine optimized for head-of-line blocking
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 09a3c6a1-a70c-4365-8c11-c719c8b0b962 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Hbf: High bandwidth flash
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 391ab928-c06c-4c0c-ad09-c8546f47df1d · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ec55eb53-8863-414b-b09b-715fa2a859b9 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mechanistic interpretability of attention heads in reasoning llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation af623f19-d7cf-47db-b58b-dab28d57fc1c · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Chain-of-thought prompting elicits reasoning in large language mod- els
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1e349555-e289-4cd3-9c23-86a6dac0275f · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d943e6d4-c4c7-49da-b37d-7fac4718f538 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Characterizing the behavior and impact of kv caching on transformer inferences under concurrency
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4837cd02-1a1f-4be7-a555-c9a6b6b3e2ce · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2058e99c-1c98-459c-ab03-76f14c1e7216 · outbound
Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e16fca69-1acf-403f-88c4-af13a4444c3a · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4c2983f9-187a-4a27-8fc3-bc6bf4e8b52e · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.