Pith. sign in

Paper Citation Record · LEDGER

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2605.19775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19775 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T02:11:26.925234Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T01:55:09.658053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.122520Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact13
  • verified fuzzy27
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc66ec53-c92b-4c92-ac79-50a1fe7f11c5 · outbound

This paper cites Vidur: A large-scale simulation frame- work for llm inference.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Vidur: A large-scale simulation frame- work for llm inference

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.713366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:24f906e0d47bb814b8e756f1041e048c4511a2fb10fdd825527b928672ce638a

Observation 38c70ef2-8da4-4000-9549-73029ad015d9 · outbound

This paper cites Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve}.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve}

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.790233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:fdca1cd8f0e8e8408f360f7500061f42e2fa9fa91b3cc26faf5c89fc84552486

Observation 2dd0a23f-066c-4141-9f14-2f9b8061721e · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.271581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:ab72154bd917d4db89f02aa8b796452df0aff81a90664cf0e2e617ddf94efd2c

Observation 0f45074d-80ae-42d8-aad2-6c39e26f2718 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.293100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:b09b64728ffb1eed71e17b3fce6ac43e0de1fb21b279732ed18c04ac63121667

Observation 854ab92f-9190-4b23-90fb-8d4efb86f63c · outbound

This paper cites Llm in a flash: Efficient large language model inference with limited memory.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm in a flash: Efficient large language model inference with limited memory

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.717522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:7c573f205e6ad364cbbc44e2bbe1d023c0219c319ddb3485cfafc98650a340b7

Observation 0092719e-4193-4989-9453-d5ad9dacc9cc · outbound

This paper cites Exploiting cxl-based memory for distributed deep learning.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Exploiting cxl-based memory for distributed deep learning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:57.475972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:479d7e940451f61d30bfe25e9fa358d301e0bc5c567b004ee42d568696316a1b

Observation cb609ecb-212b-4d90-8ee7-f09461b3bce5 · outbound

This paper cites Accelerating performance of gpu-based workloads using cxl.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Accelerating performance of gpu-based workloads using cxl

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:57.489835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:a3bebc009b87368bb43bf829a7270c12a1afa6f7b2c855e0bb7b23e3b37424c6

Observation 2847a015-37f3-497a-93af-7dff61efb162 · outbound

This paper cites Moe-lightning: High-throughput moe inference on memory-constrained gpus.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Moe-lightning: High-throughput moe inference on memory-constrained gpus

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.780957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:0e71af014901d2960b00d97bebfa2b69593eb1ecbd543dfa7311fa8dca2e4f2f

Observation c290b306-54b3-42df-be41-c345b222bbce · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.246024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:f67524fef14cfde75543513f08273e51b3271ccba5ff9b954b2ae7d593613ae1

Observation f109b5d0-54b4-4d5d-bbdc-be32baacb865 · outbound

This paper cites Llm-inference- bench: Inference benchmarking of large language models on ai acceler- ators.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm-inference- bench: Inference benchmarking of large language models on ai acceler- ators

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.721706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:0d7e5f91d4297895c09c31c6b480396a1bcd2f9bfc415654610d410b6842ce54

Observation 589e859c-81ef-40aa-8b51-2161b16c942d · outbound

This paper cites PagedEviction: Structured block-wise KV cache pruning for efficient large language model inference.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles PagedEviction: Structured block-wise KV cache pruning for efficient large language model inference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.771929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:cba918ee7173d8dc65bff1d16945178c83cbced167086e8dfb3659869b2dbfac

Observation 2a9a8dd0-c925-4770-b78a-d4aefb86c975 · outbound

This paper cites Multi-Head Attention: Collaborate Instead of Concatenate.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Multi-Head Attention: Collaborate Instead of Concatenate

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.240864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:aabe901d45ce73e97b0e0c6c0e6baef32927e46a7467f36929b95d65ebda6783

Observation f9307ee1-4d2d-4b79-8bff-1e65af4f311d · outbound

This paper cites Compute express link.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Compute express link

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.725450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:b6b30dbddf56713aac8e8877c90c5c3940ea5f1f10ad0cf4aaaf594cd9f7e813

Observation 6d6d5136-e4d2-4ce1-941c-7d307fca72c2 · outbound

This paper cites Corsair™. built for generative a.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Corsair™. built for generative a

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.768204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:216f04b1de9c68d5cfdfedbc96b04c3768ad8771916b6db87ce75905443ae713

Observation caa864e8-407d-4b4c-9f34-8013f04aa19c · outbound

This paper cites Available: https://www.d-matrix.ai/product/.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Available: https://www.d-matrix.ai/product/

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.743392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:d8047b8033199bd73f5c9d8ff67a92ce2e74594268a2f421a7e18c3139a5530a

Observation b9e4e9d7-a603-4429-9d15-7a292149cec6 · outbound

This paper cites Why we decoupled execution to accelerate i/o.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Why we decoupled execution to accelerate i/o

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.755588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:6bf704ca675a1a08c2db120a64ebde1d40e94cd6b78402824c8655260fb81aa2

Observation 67452545-d618-444e-95e7-db09289c1fa3 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.251061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:4dc6414079b3ee3de579bacfd8030014ca3a0366cb9228ce790a1227fb9bd2dd

Observation 1bce6f5a-fda1-45f4-a32e-694187999c6a · outbound

This paper cites Accelerating LLM inference throughput via asynchronous KV cache prefetching.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Accelerating LLM inference throughput via asynchronous KV cache prefetching

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.776243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:531107b6cb07235b4d141007c610175eb5397c03bb02a05c394269f394c720f0

Observation 4d9900ce-7436-4e07-95fb-61b603f96a7b · outbound

This paper cites The Llama 3 Herd of Models.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles The Llama 3 Herd of Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.256516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:62f7a887437fed98150ba8122b6e005b0c88932bf09b340d4d5b2cad5e55a51c

Observation 8ac2906b-e0cf-4bec-a6ad-817316c66a7f · outbound

This paper cites Nvidia dynamo.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Nvidia dynamo

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.747390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:5ed7476800f82262a868c909dfcc3ad7a7aa2dcbe846df7332ec1559d8e7cbc8

Observation 3537f3d1-3d01-4b54-acdf-9fd19ef8cb21 · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.751430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:466dd5eae8d7b55691adfe4d73ec42ea9b42b0803748c2ea32fb21cdebbbda79

Observation f324fd4c-f6a8-48c0-913d-169da9adea20 · outbound

This paper cites Calculon: a methodology and tool for high-level co-design of systems and large language models.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Calculon: a methodology and tool for high-level co-design of systems and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.794436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:de1cdcdd1d48c410dcf64c37ff57b68551b98995f36943bc305eb85dba48591a

Observation c70d84ad-dc6d-4f3c-b67c-4eed754daada · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.759844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:648159d20c0c3fa72218a72dc68f2a0bc9f37a7bc90eef00f8abc4d7366e29d8

Observation 5ed0f790-0cce-414c-85b2-f320cdaec115 · outbound

This paper cites Llm inference serving: Survey of recent advances and opportunities.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Llm inference serving: Survey of recent advances and opportunities

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.798791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:05198d97f953dd1760efc5195a609c5c01b98f6a5efb4f6ee2e233c912d5d480

Observation 32936c06-218e-4973-a7ac-2cc489621af9 · outbound

This paper cites A Survey on Large Language Model Acceleration based on KV Cache Management.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles A Survey on Large Language Model Acceleration based on KV Cache Management

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.261881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:d17b757a9432434df32468c11f6048d6d476c8e97e5c852306fdb56cd3192b6f

Observation 3d06ced6-091a-4814-acc9-9100f6319475 · outbound

This paper cites Deepseek-v3 technical report.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Deepseek-v3 technical report

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.803051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:d11823164fa52f15b34515d81d3f2c795fb8d166aae76f103fa9113d7535ba86

Observation 23e374e1-46ba-46fc-a19a-a27bd1e04bdc · outbound

This paper cites Minicache: Kv cache compression in depth dimension for large language models.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Minicache: Kv cache compression in depth dimension for large language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.823757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:e8548274c7c6fdc8390704092584b2de24ceb2def473955937c8719861ac2728

Observation e2754aa6-4304-4c72-8865-4a43624110c5 · outbound

This paper cites Mlp-offload: Multi-level, multi-path offloading for llm pre-training to break the gpu memory wall.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mlp-offload: Multi-level, multi-path offloading for llm pre-training to break the gpu memory wall

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.811075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:9ff3068ff163651890f1e01ca01e24a38dfb570eece29c0476d52f14f440ed4a

Observation eec82d19-bac2-4f3a-90a9-3fc0ccc63dce · outbound

This paper cites Openai o1 system card.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Openai o1 system card

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.807002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:ddfb5b558a16850929ed55295411b65fe1224aaaa384a2575b870a752139a9ee

Observation 7c14ce41-d42b-4b04-b685-b6e78eb71740 · outbound

This paper cites A survey on inference engines for large language models: Perspectives on optimization and efficiency.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles A survey on inference engines for large language models: Perspectives on optimization and efficiency

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.267265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:b76e2bf07fc1bc2971869b6d84cd0a2b78d02af5953c03e46caaba846e6fc690

Observation 46755b82-94d7-4a76-b492-25849a04a9f3 · outbound

This paper cites Mooncake: A kvcache-centric disaggregated architecture for llm serving.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mooncake: A kvcache-centric disaggregated architecture for llm serving

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.785903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:37ecd4ee4c1888fc69832d3dcaa815c0a2c24951b568b1d069c8dbde93f8d0da

Observation 6843e24d-0796-440a-9bb1-dd6ec58d8003 · outbound

This paper cites Prophet: An llm infer- ence engine optimized for head-of-line blocking.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Prophet: An llm infer- ence engine optimized for head-of-line blocking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.819523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:ab39f7eef794c9e754b54868ffdeaee6238ac0f5d3cacbe9de0112f67d6d06e9

Observation 09a3c6a1-a70c-4365-8c11-c719c8b0b962 · outbound

This paper cites Hbf: High bandwidth flash.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Hbf: High bandwidth flash

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.815200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:5b8fdae3c7e3e9efbd27845f6478a50824ed4d9efa1bcd10ca3e61fc51495c8b

Observation 391ab928-c06c-4c0c-ad09-c8546f47df1d · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.276412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:eaef8dc8b77ed920c73e901dce2f9c618e271e7ef2a87562f7ff82052b5dd3f1

Observation ec55eb53-8863-414b-b09b-715fa2a859b9 · outbound

This paper cites Mechanistic interpretability of attention heads in reasoning llms.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Mechanistic interpretability of attention heads in reasoning llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.764292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:ae19664162d104d8c493dcacd41dff59bfc8cf6a29af6f926581fc23852401ee

Observation af623f19-d7cf-47db-b58b-dab28d57fc1c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language mod- els.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Chain-of-thought prompting elicits reasoning in large language mod- els

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.733884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:1362770a919394fa495935ac73508d9e499617361456c911813a77db5275bb63

Observation 1e349555-e289-4cd3-9c23-86a6dac0275f · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-20T02:12:58.287921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:671027ba91f86c3ff73b8b989a880a30655d86f4df96af73b02044222a419cd3

Observation d943e6d4-c4c7-49da-b37d-7fac4718f538 · outbound

This paper cites Characterizing the behavior and impact of kv caching on transformer inferences under concurrency.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Characterizing the behavior and impact of kv caching on transformer inferences under concurrency

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.739269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:246b11bed4a6c1f8e16fe1ea8714ba27b4efd7a22e5ee02155ae1083d8497aa7

Observation 4837cd02-1a1f-4be7-a555-c9a6b6b3e2ce · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.283403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:03e942c740d5fafe60de131b9491b2d36b001f9adfa39e5281c7a949de54d850

Observation 2058e99c-1c98-459c-ab03-76f14c1e7216 · outbound

This paper cites Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures.

Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles Insights into deepseek-v3: Scaling challenges and reflections on hardware for ai architectures

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T02:12:58.729560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T02:11:26.925234Z digest=sha256:5d1882d4e749cb6155412f555b04cbe88914f4818529666cfb6459fc67378dba

Pith citing papers

Observation e16fca69-1acf-403f-88c4-af13a4444c3a · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.093424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:19b6a44627791cebae03c04928d1c1cb118bb34055f96e77af78b08cf5f84599

Observation 4c2983f9-187a-4a27-8fc3-bc6bf4e8b52e · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.154963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:eb0c4eaeb73b82441f04d34350ee7ddfa308048a717cfd6dbf2b28318e3d47b8