Pith. sign in

Paper Citation Record · LEDGER

FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2303.06865.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.06865 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:50:33.189053Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

47
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b6cb1e59-f963-43d7-9a41-e89dea8d448a · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:29:11.378091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:b5194a27f6a5d70330722710b0bdbc2df8b756efd3b7b636ac5b87403dec7ab4

Observation ad34b387-962f-4c5c-8a3c-889b798ae04d · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T18:00:50.435347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:a1cd26d93aed48cd2191df09d9b211f536c78f8241273fa5f7905f4e7bcf89c1

Observation b514fe48-8fb9-4ec3-9311-fb3b4b968682 · inbound

Efficient Memory Management for Large Language Model Serving with PagedAttention cites this paper.

Efficient Memory Management for Large Language Model Serving with PagedAttention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:07.800919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T15:03:07.651839Z digest=sha256:403d34674d8810a4a9f58a32741ce451d57e3c53303c3ec61b83093383459237

Observation f7f8442b-e6d9-4ea7-9d6c-ac4d4c74212a · inbound

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention cites this paper.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:45:57.821474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:65421378a0e8d80647e3e24dc87aea783fe968597a0c1acd5d94c8de922941fb

Observation f5c6d858-d87e-44cf-a489-0a3aa6738bcd · inbound

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots cites this paper.

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:34.776645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T08:53:21.013613Z digest=sha256:9c3fb9bd081621da4e267050e1cb39cd4a4a8ec7a8bc93d6f40282240206aa57

Observation a64111e7-1b16-4e76-9299-57d4b5e2b293 · inbound

NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference cites this paper.

NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.425186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T14:14:33.939889Z digest=sha256:60ae749545830737772bc3966798e6de5683602e095fe89195b81c2b387020da

Observation 7da1e05c-8151-459e-b767-4162b988f6b2 · inbound

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production cites this paper.

Position: LLM Inference Should Be Evaluated as Energy-to-Token Production FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:19.393998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T05:17:24.147248Z digest=sha256:2a276748e3615e05a64ab41329db6db63158e52cd3bf7309faa3f18e58ac1448

Observation 9338576d-07b3-400e-a5f0-267a21c1bbab · inbound

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload cites this paper.

TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:03.728349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:08:07.318040Z digest=sha256:e3cd0622ef8042157ed959cb5975a8e880050d05440ba675fefa923dccccf1c3

Observation d4f29e81-15f3-4070-8c88-5b061cbee79c · inbound

Motion-Compensated Weight Compression cites this paper.

Motion-Compensated Weight Compression FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:04:40.443077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T12:58:27.637822Z digest=sha256:44792db217ae9ac4a728b726a843b87bfccdf213aaabe49665e0a45505064593

Observation d6eb5ba2-eb03-402c-8623-afc058fd5251 · inbound

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving cites this paper.

DisagFusion: Asynchronous Pipeline Parallelism and Elastic Scheduling for Disaggregated Diffusion Serving FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T20:53:57.615932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T20:46:51.896596Z digest=sha256:36e1c1dd2fa2c752a4df0c08540fe0b279faa656b1cd6e1692425eeb7083dd74

Observation 17bb96a4-f528-44b7-a38b-8d275546cc8f · inbound

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference cites this paper.

SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:13:17.841684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T10:07:06.141581Z digest=sha256:da6243ac935ed2f60ecae07e9000e853a7c11f235583154ad9a6cdcb77986942

Observation f44a03d9-0f6a-4f66-af40-cfea99326fcf · inbound

Do Transformers Need Three Projections? Systematic Study of QKV Variants cites this paper.

Do Transformers Need Three Projections? Systematic Study of QKV Variants FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.291462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:14:49.475561Z digest=sha256:83c206c366facfc9e1ca6d2dcd4753d53f2484b096e8c27e8843e7497f9d8f6e

Observation e653dc00-faa0-4e9f-9722-0d681cb278c4 · inbound

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs cites this paper.

From Rigid to Dynamic: Entropy-Guided Adaptive Inference for Long-Context LLMs FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.756499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T16:57:02.709967Z digest=sha256:33d46e746b266d5373b9b7eeecfa11d7b907904360c08b996ac96c9c3db7d170

Observation 9fc2fa37-c415-4186-8e06-9dd863008722 · inbound

RoPE-Aware Bit Allocation for KV-Cache Quantization cites this paper.

RoPE-Aware Bit Allocation for KV-Cache Quantization FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.345142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:05:13.790381Z digest=sha256:c6b7258bdb97e9349bedd1c800ae6afaf2b0899a2a494c6e59b8827144b257f4

Observation 00bd9ec3-3e97-43af-8c72-a27711523849 · inbound

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache cites this paper.

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:57:06.396714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-02T15:55:40.177742Z digest=sha256:3009548c512d9bdcc7a74fd9143524fb0a521b0b868647834709d0c00b5e8e24

Observation c5aacbd5-6a12-4614-8d0d-3a86315c0e78 · inbound

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents cites this paper.

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 105

Resolution
verified exact
local_arxiv, observed 2026-07-10T01:36:44.214493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-10T01:26:59.421158Z digest=sha256:e711942379c82672479d8d2e5305f49585c6b412c5d11853f07b2ff1047dc96f

Observation 470deda4-87b7-40b5-a309-a9f786629989 · inbound

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization cites this paper.

PagedWeight: Efficient MoE LLM Serving with Dynamic Quality-Aware Weight Quantization FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:10.165093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:10.165093Z digest=sha256:9283ed3c81474a52e41d0b7e0fd32bdc32edff2f7bc68d73cabce33ed7750204

Observation 57070d58-0786-4889-b914-0d0d6ffff66f · inbound

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration cites this paper.

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T09:50:33.189053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:50:33.189053Z digest=sha256:c90ffd0cd2c93cb8e08ef08fcd21feb25c9e66b0baa8edfa9bccc6ecc87aebd5

Observation 985de2a4-2f50-4365-b08d-a1b33ec5e371 · inbound

Transition-Aware Backend Dispatch for Edge LLM Inference cites this paper.

Transition-Aware Backend Dispatch for Edge LLM Inference FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:58.812521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:58.812521Z digest=sha256:d643bc254580551a1ae4e9545895b9548154e844c42323da16a91e95d3a5f8a4