Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2303.06865.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T09:50:33.189053Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
47
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation b6cb1e59-f963-43d7-9a41-e89dea8d448a · inbound
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad34b387-962f-4c5c-8a3c-889b798ae04d · inbound
H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b514fe48-8fb9-4ec3-9311-fb3b4b968682 · inbound
Efficient Memory Management for Large Language Model Serving with PagedAttention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f7f8442b-e6d9-4ea7-9d6c-ac4d4c74212a · inbound
AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f5c6d858-d87e-44cf-a489-0a3aa6738bcd · inbound
Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a64111e7-1b16-4e76-9299-57d4b5e2b293 · inbound
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7da1e05c-8151-459e-b767-4162b988f6b2 · inbound
Position: LLM Inference Should Be Evaluated as Energy-to-Token Production FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9338576d-07b3-400e-a5f0-267a21c1bbab · inbound
TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4f29e81-15f3-4070-8c88-5b061cbee79c · inbound
Motion-Compensated Weight Compression FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d6eb5ba2-eb03-402c-8623-afc058fd5251 · inbound
DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 17bb96a4-f528-44b7-a38b-8d275546cc8f · inbound
SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f44a03d9-0f6a-4f66-af40-cfea99326fcf · inbound
Do Transformers Need Three Projections? Systematic Study of QKV Variants FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e653dc00-faa0-4e9f-9722-0d681cb278c4 · inbound
From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fc2fa37-c415-4186-8e06-9dd863008722 · inbound
RoPE-Aware Bit Allocation for KV-Cache Quantization FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 00bd9ec3-3e97-43af-8c72-a27711523849 · inbound
GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c5aacbd5-6a12-4614-8d0d-3a86315c0e78 · inbound
What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 470deda4-87b7-40b5-a309-a9f786629989 · inbound
PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57070d58-0786-4889-b914-0d0d6ffff66f · inbound
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985de2a4-2f50-4365-b08d-a1b33ec5e371 · inbound
Transition-Aware Backend Dispatch for Edge LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.