Pith. sign in

Paper Citation Record · LEDGER

Efficiently Scaling Transformer Inference

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2211.05102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.05102 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 49 of 49 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:06.592617Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

57
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5cdcc096-9464-4c4c-8c12-3d87e1567d38 · inbound

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints cites this paper.

GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints Efficiently Scaling Transformer Inference

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:53:59.470349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T06:53:58.960357Z digest=sha256:5831aada9307f5bb4fb11fe312e929040c4ea8e53c19af6de76cdbd1a9247a73

Observation 7634cd19-3f13-4289-b6de-630640e8c514 · inbound

QLoRA: Efficient Finetuning of Quantized LLMs cites this paper.

QLoRA: Efficient Finetuning of Quantized LLMs Efficiently Scaling Transformer Inference

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:29:53.618992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T13:29:53.345251Z digest=sha256:488079a24eec10e844afb843968ff14f3f99002ff9621cb26daed49045a07676

Observation 56958681-84c9-4f80-a94e-4e5c691163df · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficiently Scaling Transformer Inference

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.145712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:3c186524f5fe9cceefd936f9c3b5041e6b027c7350db71aaf5460ccb1073316a

Observation 659423ce-ded9-4c05-970b-b38ca297a6d8 · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention Efficiently Scaling Transformer Inference

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.746609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:0b3030fdfa97e13baab18acea9ff18d50c3dfa8080735f753a2e7013e8515178

Observation 0f66068a-970b-4ac0-9265-9eb03bc23fb4 · inbound

Efficient Streaming Language Models with Attention Sinks cites this paper.

Efficient Streaming Language Models with Attention Sinks Efficiently Scaling Transformer Inference

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:40:13.346221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T00:40:12.924357Z digest=sha256:c245ee84a2697cb5d6bd10f7064102fa8f85f710a53d3e515a683d1bf5893aff

Observation c22765bb-c1de-4bf7-a67d-a0434dbcb6bb · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Efficiently Scaling Transformer Inference

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:36:17.998786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:2c5ac6e2ac820ffa9d722737d27ec3c64d4ad949be3d9372b7324e44bc95a2c8

Observation 905d6d42-0609-497b-9af6-656f9025c5a6 · inbound

AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution cites this paper.

AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution Efficiently Scaling Transformer Inference

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:35:56.995517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:35:56.995517Z digest=sha256:aed06a0be619b5061e5966ca243bae2f1837b621f02b9554bd69cffd6f0e62f0

Observation 3c13e79d-6ae7-4f19-9ac2-17fabad089b4 · inbound

MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series cites this paper.

MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series Efficiently Scaling Transformer Inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T13:02:19.444236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:02:19.444236Z digest=sha256:fb2b8cbef8f93889dc29c4f4139a7b58c5779fe724e7a3ce9e428a2537b81718

Observation 7e71a091-1a37-45d6-80f0-d925ea1e9ac6 · inbound

Scaling Deep Learning Training with MPMD Pipeline Parallelism cites this paper.

Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:21:38.863523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:21:38.863523Z digest=sha256:4bdd2aaf54b0f5b36e5edaed1586cf21e1a4e6bdd39d42cbfeca0dda020d7b0b

Observation 21a53ee0-b0f2-4a00-9062-ea2bb4d60a09 · inbound

TreeKV: Smooth Key-Value Cache Compression with Tree Structures cites this paper.

TreeKV: Smooth Key-Value Cache Compression with Tree Structures Efficiently Scaling Transformer Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:25:41.222232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:25:41.222232Z digest=sha256:1e18f1efbe2fb29742efae5c0f68ec83cae77220a57d7b0c9580d70ee8b8e934

Observation 6dfb7b8e-396e-40ef-bcd6-772041024548 · inbound

StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel cites this paper.

StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel Efficiently Scaling Transformer Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:07:54.856058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:07:54.856058Z digest=sha256:b782c0d7108ec75c2bb016e1e2dea6e96153b354080932c784bdc3c55f2995c1

Observation eb050189-b395-4a5f-945e-55fd897b7c6e · inbound

Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models cites this paper.

Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models Efficiently Scaling Transformer Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T18:41:32.496327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:41:32.496327Z digest=sha256:8b4392935fccb0fe7bcc7e80c0a35bdb4808650ba810085143e9a24bc4130cf2

Observation cec36e2b-a024-4f18-9252-28b3938dda3a · inbound

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project cites this paper.

Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Efficiently Scaling Transformer Inference

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:05:09.567263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:04:11.859563Z digest=sha256:3435200d1e95fbc805918f221ec8d3386ca913fc3b89b4b96b37eaa052d37049

Observation fa2ce1d6-8417-4827-8db7-a558c0aa6b3a · inbound

Cobra: Efficient Line Art COlorization with BRoAder References cites this paper.

Cobra: Efficient Line Art COlorization with BRoAder References Efficiently Scaling Transformer Inference

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:06.592617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:06.592617Z digest=sha256:11935ecabdf12ec6bdae1775cc5c7c01b62d6cccb4b64dadc42bd1cc7a587458

Observation ce55d1e5-9c2d-4744-b000-1daf1172ee2c · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Efficiently Scaling Transformer Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:56.481665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:56.481665Z digest=sha256:49cc2ced2df1ce7cee365d08164e447b7045e1101b437ecb24728e6febcec097

Observation 1702db5c-1acd-43d7-b486-f84714a7ed3b · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Efficiently Scaling Transformer Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:32:31.926378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:32:31.926378Z digest=sha256:ba7ac882957a8135e167ecc13f90add204bdb0de53910766e2e7b893e66d0d9c

Observation 8330db3e-096e-44b8-bfdf-25074bb9cfe1 · inbound

EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration cites this paper.

EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration Efficiently Scaling Transformer Inference

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:09:04.247876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:09:04.247876Z digest=sha256:e03c827e57c604ba7bf2547daf03da1272d0558c68994bbb6ca239ce1a664752

Observation 556c831d-9d4d-4ebb-a79a-5b862a0c33b8 · inbound

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models cites this paper.

Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficiently Scaling Transformer Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:23.701288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:23.701288Z digest=sha256:f305c2baad84a579df94f8567fb9704449bfe7c634abb7ceff56106605d6afc8

Observation e7098408-3cbe-4443-87ec-9a16f453c9c8 · inbound

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling cites this paper.

AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling Efficiently Scaling Transformer Inference

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.949143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:52.949143Z digest=sha256:575870aff0e22dfbdcfc0042c0bc3c60a307af2fcd4c2db70b7a0ef2e481c82d

Observation 48b0a8f9-3d01-4471-9ddb-e8c340b4c42d · inbound

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques cites this paper.

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques Efficiently Scaling Transformer Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:41.039701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:13:41.039701Z digest=sha256:fda4e5c19680c4f9d9d02ca1dcaf1483ffe799ffcc8a57ce0f1cfd45b314e12d

Observation 8b7f616e-de53-44e7-bf53-42aa1e695d13 · inbound

Characterizing Communication Patterns in Distributed Large Language Model Inference cites this paper.

Characterizing Communication Patterns in Distributed Large Language Model Inference Efficiently Scaling Transformer Inference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T16:14:22.520713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:14:22.520713Z digest=sha256:798764b2eb4f8508f87f0f1c3e146c2dbdad7dcea8b94a9d803a02e3b932dcf3

Observation 6cdc3b44-ecae-4a26-ae82-569ed53695b0 · inbound

Efficient Item ID Generation for Large-Scale LLM-based Recommendation cites this paper.

Efficient Item ID Generation for Large-Scale LLM-based Recommendation Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:46:49.751745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:46:49.751745Z digest=sha256:df1ab8b0bca4b2f502a6a06c2b63238ca565c688159f9e2a7d90d3e97bdd6767

Observation 4b11f7e9-a006-4f41-9ac0-d6e1264d6ff8 · inbound

From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill cites this paper.

From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill Efficiently Scaling Transformer Inference

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T08:51:08.938424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T08:47:29.759674Z digest=sha256:0f3753591e757c7757b8549330cc33c26c334c2350b4784427ec7a604c79a34c

Observation 5b5bebc3-06fd-4ab5-a9f6-29afd0d72509 · inbound

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse cites this paper.

Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Efficiently Scaling Transformer Inference

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:05:39.116323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T02:04:46.559881Z digest=sha256:35f2ae21b786f2150556079a810505452b21758d6d1c501f86905eb03c077154

Observation 3c24d81d-572b-4fff-9a4b-c0f397a30751 · inbound

Incremental Transformer Neural Processes cites this paper.

Incremental Transformer Neural Processes Efficiently Scaling Transformer Inference

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T21:52:58.923503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:52:58.923503Z digest=sha256:5795c1076553b69bd5e8c058d947db67778dc7328781b76a38ddd0ab6a944f7d

Observation 7b535baf-9b7d-4ea7-9668-a50fe86a4abc · inbound

Attention Residuals cites this paper.

Attention Residuals Efficiently Scaling Transformer Inference

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:04.533914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T06:39:04.312270Z digest=sha256:46ac312fd57cc3f67931430663a82f5022a192bc546985aa4a07cd767de6aae9

Observation d9e1594d-7aba-4c8d-947e-142a96c067e7 · inbound

Generating Counterfactual Patient Timelines from Real-World Data cites this paper.

Generating Counterfactual Patient Timelines from Real-World Data Efficiently Scaling Transformer Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:42:48.829434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T11:42:42.030779Z digest=sha256:f39e00d7ac19d4a9e911c5f55c171888702239c33cdd0b980682a7cba77f414c

Observation ce15ca23-97f6-48f2-a545-5d59ed51e8ee · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Efficiently Scaling Transformer Inference

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:52.018231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:3adc02f62c71f26dd25a1864a7240a587d266b43e0cbfd46453f3c19f6192d52

Observation 7ade546f-7fe1-4dad-ab0c-e03dd36c01ec · inbound

Benchmarking Compound AI Applications for Hardware-Software Co-Design cites this paper.

Benchmarking Compound AI Applications for Hardware-Software Co-Design Efficiently Scaling Transformer Inference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T16:17:08.045511Z digest=sha256:d886fd30ddcade24333493888a7d6e70d994c429798b66c69f86b883906805bb

Observation 74e644a0-27aa-498f-ba03-2f97857c51f8 · inbound

GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs cites this paper.

GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs Efficiently Scaling Transformer Inference

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:03.350227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:16:14.625291Z digest=sha256:b3a5c94c7b8b151bd5e8563e54e216feb3ffbd9768dc68b95896df43add95e6c

Observation e83c417c-21c8-4736-835a-9a50fb8a06fa · inbound

Continuous Semantic Caching for Low-Cost LLM Serving cites this paper.

Continuous Semantic Caching for Low-Cost LLM Serving Efficiently Scaling Transformer Inference

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:01:03.174067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T02:32:28.923367Z digest=sha256:2113d6b99596550f6fd304a6cba528b05d416ffba58482aa935c55cb003f262d

Observation 5e1e20ed-d45a-466c-92a1-cbd613639350 · inbound

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective cites this paper.

Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Efficiently Scaling Transformer Inference

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:23.906304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-07T16:41:23.234607Z digest=sha256:a51d697c10e3375775e2c23a9630259040141c8dd6fd632a4aaefa9ea164faf2

Observation 8848a48d-05dc-45ea-85a1-fe771068d49b · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Efficiently Scaling Transformer Inference

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.083007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-07T10:42:27.644514Z digest=sha256:09dddd86fb5408f381f040abde951e1b1c4e80a368cb72fe80fcface3ad0fdff

Observation 450f4a24-9764-49a5-b98e-82aebb3fc6f7 · inbound

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling cites this paper.

Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Efficiently Scaling Transformer Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T15:28:12.785926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:28:12.785926Z digest=sha256:d20a8ff2d22cdd6072815d79515f859a2f9e0218e313cd4b7eb5a45953132254

Observation 9e29908e-4592-4e72-a2f8-39160a6478fd · inbound

Speculative Decoding and the Curse of Multilinguality cites this paper.

Speculative Decoding and the Curse of Multilinguality Efficiently Scaling Transformer Inference

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:13.183398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T07:15:35.673216Z digest=sha256:b77d7d8f0fc8cad1bf0917208d7bfec170b37054a985cee127535f33a44ee462

Observation 6270690f-4f21-4548-85fb-8fd3762ffe9a · inbound

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models cites this paper.

HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models Efficiently Scaling Transformer Inference

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.495111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T08:46:01.500880Z digest=sha256:991bb46484dbde7524b7155095217bb991a12b975d58c1f857fb057008db95ba

Observation 4e7bce96-5d4b-473b-a00a-6879787917c9 · inbound

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation cites this paper.

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation Efficiently Scaling Transformer Inference

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:45.470405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T08:38:32.577228Z digest=sha256:dff1632e6f7eb97ed18953c573cea9b72150933fb1f1c50dbf0e5f1b9b7df3ba

Observation c43ee846-db3c-47c7-abef-8aa481831865 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization Efficiently Scaling Transformer Inference

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.375836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:94c96d9e101f8d22cbe957f5ae24f993c0531f1fbe5eb9194ab133ab945e48a1

Observation 617743ab-1c39-43d5-8a12-9e5686752c8e · inbound

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control cites this paper.

A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control Efficiently Scaling Transformer Inference

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:55:44.184963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T01:38:42.824342Z digest=sha256:8368ac15a5fdbd9b9aec2619f6752d7ace8efeed767b84b15d4f45a0d34f2094

Observation 65435f22-087c-4829-805e-b9e642318c20 · inbound

Design-CP: Context Parallelism for Design of Protein Nanoparticles cites this paper.

Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficiently Scaling Transformer Inference

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-12T02:29:52.764344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:29:52.764344Z digest=sha256:4dce30f76adbf46b9b97b9cf77178d1a073ad39a8b49f1f6440dc1136e0bf896

Observation 2e10be75-4514-4f73-910c-8c114af35cf2 · inbound

ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling cites this paper.

ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling Efficiently Scaling Transformer Inference

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T05:27:55.437107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T05:27:55.437107Z digest=sha256:5429310499edc2cdd82e42f353836fa2b28fbf3ff955e7f3b1315571a2d19981

Observation c36018a4-bdc7-44a3-98d5-ba5ac65e3bac · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Efficiently Scaling Transformer Inference

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.144012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:993a83a8e9608c00e159bb8935a76e25a91f63e1ddc05b2df920d96f33d4709d

Observation f169fdc9-3abf-4a0e-9e19-acaf3707c28b · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving Efficiently Scaling Transformer Inference

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.691294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:a9214af5bc32ff114fc6ca0e165b5f217083b0c461c93ed85ccf73ad70720adf

Observation 0dcb5b3c-aeb2-460f-9584-371ba5f6596f · inbound

The Cost and Network Limits of Space-Based AI Compute cites this paper.

The Cost and Network Limits of Space-Based AI Compute Efficiently Scaling Transformer Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T04:42:47.305650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:42:47.305650Z digest=sha256:f58e341e19baf9b950ca224514d6696de1924bef5b1f1c377b5b601b1e10ba51

Observation 91092c63-81f6-4338-b79e-094edccb645d · inbound

Efficient Clustering with Provable Guardrails for LLM Inference at Scale cites this paper.

Efficient Clustering with Provable Guardrails for LLM Inference at Scale Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T12:03:18.853036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:03:18.853036Z digest=sha256:57c95a2609a15af91541cea06eabf6b9c9ee8a55c007ad847d3a991705c5eb94

Observation 9bcae3a8-d910-4835-b8d0-a087730dd323 · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems Efficiently Scaling Transformer Inference

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:42.278984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:42.278984Z digest=sha256:552892fcae946229e0437d6b2b90290c8d55ad2b73550c6ed86491d0c795bd6e

Observation 52e0d750-d1b7-4592-ab19-68ddd67a108b · inbound

StrataCL: Fabric-Native Communication Library for Production Supernodes cites this paper.

StrataCL: Fabric-Native Communication Library for Production Supernodes Efficiently Scaling Transformer Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T15:42:57.155893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:42:57.155893Z digest=sha256:c3e6c9aa1b8a20604575106209f0461e1bf755d182773ec7b01a063fac290548

Observation 60fae710-b15f-47ad-aaab-3b0a667bc766 · inbound

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL cites this paper.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Efficiently Scaling Transformer Inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:57.766979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:57.766979Z digest=sha256:3ca52fc3b2ddfb0df654c774801e8bc7c2a7d6c38373cab2301ba491241b4cec

Observation 38d93a2f-d049-4842-a7ab-0c1bcbdb32c5 · inbound

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference cites this paper.

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference Efficiently Scaling Transformer Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:59.803368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:59.803368Z digest=sha256:906f6f286f4612231561e9559d7cbd19e227cdac9bd192393481ccaa80b96c9f