Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 49 inbound Pith citation observations for arXiv:2211.05102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:06.592617Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
57
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 5cdcc096-9464-4c4c-8c12-3d87e1567d38 · inbound
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints Efficiently Scaling Transformer Inference
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7634cd19-3f13-4289-b6de-630640e8c514 · inbound
QLoRA: Efficient Finetuning of Quantized LLMs Efficiently Scaling Transformer Inference
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 56958681-84c9-4f80-a94e-4e5c691163df · inbound
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Efficiently Scaling Transformer Inference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 659423ce-ded9-4c05-970b-b38ca297a6d8 · inbound
Efficient Memory Management for Large Language Model Serving with PagedAttention Efficiently Scaling Transformer Inference
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0f66068a-970b-4ac0-9265-9eb03bc23fb4 · inbound
Efficient Streaming Language Models with Attention Sinks Efficiently Scaling Transformer Inference
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c22765bb-c1de-4bf7-a67d-a0434dbcb6bb · inbound
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Efficiently Scaling Transformer Inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 905d6d42-0609-497b-9af6-656f9025c5a6 · inbound
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution Efficiently Scaling Transformer Inference
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c13e79d-6ae7-4f19-9ac2-17fabad089b4 · inbound
MarketGPT: Developing a Pre-trained transformer (GPT) for Modeling Financial Time Series Efficiently Scaling Transformer Inference
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e71a091-1a37-45d6-80f0-d925ea1e9ac6 · inbound
Scaling Deep Learning Training with MPMD Pipeline Parallelism Efficiently Scaling Transformer Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a53ee0-b0f2-4a00-9062-ea2bb4d60a09 · inbound
TreeKV: Smooth Key-Value Cache Compression with Tree Structures Efficiently Scaling Transformer Inference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dfb7b8e-396e-40ef-bcd6-772041024548 · inbound
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel Efficiently Scaling Transformer Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb050189-b395-4a5f-945e-55fd897b7c6e · inbound
Towards Sustainable NLP: Insights from Benchmarking Inference Energy in Large Language Models Efficiently Scaling Transformer Inference
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cec36e2b-a024-4f18-9252-28b3938dda3a · inbound
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project Efficiently Scaling Transformer Inference
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation fa2ce1d6-8417-4827-8db7-a558c0aa6b3a · inbound
Cobra: Efficient Line Art COlorization with BRoAder References Efficiently Scaling Transformer Inference
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce55d1e5-9c2d-4744-b000-1daf1172ee2c · inbound
CoDec: Prefix-Shared Decoding Kernel for LLMs Efficiently Scaling Transformer Inference
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1702db5c-1acd-43d7-b486-f84714a7ed3b · inbound
Hardware-Efficient Attention for Fast Decoding Efficiently Scaling Transformer Inference
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8330db3e-096e-44b8-bfdf-25074bb9cfe1 · inbound
EQuARX: Efficient Quantized AllReduce in XLA for Distributed Machine Learning Acceleration Efficiently Scaling Transformer Inference
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556c831d-9d4d-4ebb-a79a-5b862a0c33b8 · inbound
Infinite Sampling: Efficient and Stable Grouped RL Training for Large Language Models Efficiently Scaling Transformer Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7098408-3cbe-4443-87ec-9a16f453c9c8 · inbound
AbbIE: Autoregressive Block-Based Iterative Encoder for Efficient Sequence Modeling Efficiently Scaling Transformer Inference
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48b0a8f9-3d01-4471-9ddb-e8c340b4c42d · inbound
ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques Efficiently Scaling Transformer Inference
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b7f616e-de53-44e7-bf53-42aa1e695d13 · inbound
Characterizing Communication Patterns in Distributed Large Language Model Inference Efficiently Scaling Transformer Inference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdc3b44-ecae-4a26-ae82-569ed53695b0 · inbound
Efficient Item ID Generation for Large-Scale LLM-based Recommendation Efficiently Scaling Transformer Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b11f7e9-a006-4f41-9ac0-d6e1264d6ff8 · inbound
From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill Efficiently Scaling Transformer Inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5b5bebc3-06fd-4ab5-a9f6-29afd0d72509 · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse Efficiently Scaling Transformer Inference
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3c24d81d-572b-4fff-9a4b-c0f397a30751 · inbound
Incremental Transformer Neural Processes Efficiently Scaling Transformer Inference
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b535baf-9b7d-4ea7-9668-a50fe86a4abc · inbound
Attention Residuals Efficiently Scaling Transformer Inference
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d9e1594d-7aba-4c8d-947e-142a96c067e7 · inbound
Generating Counterfactual Patient Timelines from Real-World Data Efficiently Scaling Transformer Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ce15ca23-97f6-48f2-a545-5d59ed51e8ee · inbound
DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Efficiently Scaling Transformer Inference
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ade546f-7fe1-4dad-ab0c-e03dd36c01ec · inbound
Benchmarking Compound AI Applications for Hardware-Software Co-Design Efficiently Scaling Transformer Inference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 74e644a0-27aa-498f-ba03-2f97857c51f8 · inbound
GPU Acceleration of Sparse Fully Homomorphic Encrypted DNNs Efficiently Scaling Transformer Inference
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e83c417c-21c8-4736-835a-9a50fb8a06fa · inbound
Continuous Semantic Caching for Low-Cost LLM Serving Efficiently Scaling Transformer Inference
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5e1e20ed-d45a-466c-92a1-cbd613639350 · inbound
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective Efficiently Scaling Transformer Inference
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8848a48d-05dc-45ea-85a1-fe771068d49b · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Efficiently Scaling Transformer Inference
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 450f4a24-9764-49a5-b98e-82aebb3fc6f7 · inbound
Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling Efficiently Scaling Transformer Inference
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e29908e-4592-4e72-a2f8-39160a6478fd · inbound
Speculative Decoding and the Curse of Multilinguality Efficiently Scaling Transformer Inference
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6270690f-4f21-4548-85fb-8fd3762ffe9a · inbound
HyperQuant: A Rate-Distortion-Optimal Quantization Pipeline for Large Language and Diffusion Models Efficiently Scaling Transformer Inference
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4e7bce96-5d4b-473b-a00a-6879787917c9 · inbound
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation Efficiently Scaling Transformer Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c43ee846-db3c-47c7-abef-8aa481831865 · inbound
RoPE-Aware Bit Allocation for KV-Cache Quantization Efficiently Scaling Transformer Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 617743ab-1c39-43d5-8a12-9e5686752c8e · inbound
A Systematic Approach to Multi-Agent AI from Advanced Regulatory Control Theory: Safe and Auditable LLM Operator Agents for Process Control Efficiently Scaling Transformer Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 65435f22-087c-4829-805e-b9e642318c20 · inbound
Design-CP: Context Parallelism for Design of Protein Nanoparticles Efficiently Scaling Transformer Inference
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e10be75-4514-4f73-910c-8c114af35cf2 · inbound
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modeling Efficiently Scaling Transformer Inference
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36018a4-bdc7-44a3-98d5-ba5ac65e3bac · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Efficiently Scaling Transformer Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f169fdc9-3abf-4a0e-9e19-acaf3707c28b · inbound
Think Before You Grid-Search: Floor-First Triage for LLM Serving Efficiently Scaling Transformer Inference
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0dcb5b3c-aeb2-460f-9584-371ba5f6596f · inbound
The Cost and Network Limits of Space-Based AI Compute Efficiently Scaling Transformer Inference
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91092c63-81f6-4338-b79e-094edccb645d · inbound
Efficient Clustering with Provable Guardrails for LLM Inference at Scale Efficiently Scaling Transformer Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcae3a8-d910-4835-b8d0-a087730dd323 · inbound
KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems Efficiently Scaling Transformer Inference
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e0d750-d1b7-4592-ab19-68ddd67a108b · inbound
StrataCL: Fabric-Native Communication Library for Production Supernodes Efficiently Scaling Transformer Inference
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60fae710-b15f-47ad-aaab-3b0a667bc766 · inbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Efficiently Scaling Transformer Inference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d93a2f-d049-4842-a7ab-0c1bcbdb32c5 · inbound
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference Efficiently Scaling Transformer Inference
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.