Pith. sign in

Paper Citation Record · LEDGER

HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.01164.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.01164 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:24:02.495472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T15:30:17.903404Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bc4ff586-5d51-4d6c-bab2-8fbf670db8c4 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:44.160888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:44.160888Z digest=sha256:7bc499d98f9038305886f4285cc10eb8f679eba4f07af71e4dd8e36b2a6d6b67

Observation 78edecc1-f912-4d40-81f1-8d8254a6173a · inbound

Deploying Foundation Model Powered Agent Services: A Survey cites this paper.

Deploying Foundation Model Powered Agent Services: A Survey HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T13:09:45.100032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:09:45.100032Z digest=sha256:cbd7568dc79003c255c924bc6097919953ba0e2b87c427c4c29f8855da44d751

Observation d679780a-a501-4a46-9f38-09526025853d · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.863029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.863029Z digest=sha256:3b03dbd5ffac7acd7aed2d1e10e9d6da8ea26f55a3c6f8d8c246a229e84e8c37

Observation 9990155a-14b6-4bca-b2cc-a8aa0aa8f9ea · inbound

WindVE: Collaborative CPU-NPU Vector Embedding cites this paper.

WindVE: Collaborative CPU-NPU Vector Embedding HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:42:31.684392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:42:31.684392Z digest=sha256:44cb68fa38e445fe758e864fbec627b9241764161c9ab7756bffd18ba1e05564

Observation 750fd8ea-9f65-4f17-a014-ee7cfe6b1cf2 · inbound

RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU cites this paper.

RAGDoll: Efficient Offloading-based Online RAG System on a Single GPU HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T12:24:02.495472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:24:02.495472Z digest=sha256:6cef20f368baa8de1998b1ba34e0a00af65af0dc29ebcd9e72d49d44eaead214

Observation 2932364a-1366-44ef-bafe-feca4697decb · inbound

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips cites this paper.

SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:30:17.913894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T15:26:01.283448Z digest=sha256:b04f205c3f92a25ff2472c6e3fa3192a21a018e0219687e5296fdc93d78fcd86