Pith. sign in

Paper Citation Record · LEDGER

Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2406.01566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.01566 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:45:03.061729Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T13:36:59.461933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 309083a9-c73c-461e-ba36-00dd5810ebce · inbound

HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware cites this paper.

HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.592114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T21:21:03.598124Z digest=sha256:7c4db37cb93f14803c772502a5e4ef67b3e0a3d24565a2adfd937e0ae0336c76

Observation b8f0a6c8-6de0-4e39-81f8-0a2c3a94eedf · inbound

AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding cites this paper.

AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T17:39:10.360919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:39:10.360919Z digest=sha256:cda1ea46bf95389b501579b307e13b47649ba5093f2a32223872537088d1d818

Observation 7765e1a6-ff32-4733-8bf1-631b1b924de7 · inbound

DeServe: Towards Affordable Offline LLM Inference via Decentralization cites this paper.

DeServe: Towards Affordable Offline LLM Inference via Decentralization Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:16:53.070722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:16:53.070722Z digest=sha256:ddc14b196a1e2b7e4938443cfb8058dae3260fac8856afc89245956b77536120

Observation e4f76555-430e-415b-bb32-0659880987cb · inbound

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs cites this paper.

Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T18:02:12.424564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:02:12.424564Z digest=sha256:4dc0416cd1702d6b1104eb6fb6717bc7a0f0507020cb9b4cec80fabaa3f036b5

Observation e90f9fba-9644-48d7-95b4-e75fe7430ea5 · inbound

EcoServe: Designing Carbon-Aware AI Inference Systems cites this paper.

EcoServe: Designing Carbon-Aware AI Inference Systems Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T20:31:35.738204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:31:35.738204Z digest=sha256:f26bc4dac4a52027a3c05932d526f11a7ab1f4ebb79cce1bd68cdd530b412b8d

Observation 3200e3ad-3931-4619-a58f-92ab2c9c5492 · inbound

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment cites this paper.

HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T11:32:21.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:32:21.524538Z digest=sha256:62bccc89bfc6defb562689c53f8729fe62f96d0c3dff328d87f5ee62f7f48d51

Observation 75d28a88-1da7-4443-97bf-0a4f4d4a3f55 · inbound

gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling cites this paper.

gLLM: Global Balanced Pipeline Parallelism System for Distributed LLM Serving with Token Throttling Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:45:03.061729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:45:03.061729Z digest=sha256:f795675679eab7aafd89c4354951400ef95140e7423d5b4655a6a88fc577753e

Observation d1b1fac3-c08e-4b7d-83c9-3ae853b6603b · inbound

Taming the Titans: A Survey of Efficient LLM Inference Serving cites this paper.

Taming the Titans: A Survey of Efficient LLM Inference Serving Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:10.000897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:49:10.000897Z digest=sha256:4c0a054453562a756b051866e76186be7dd3ebe70ffce73c60af036f35be139a

Observation bee64b80-c58e-449c-91a5-cdef69b47ff3 · inbound

Harmonia: End-to-End RAG Serving Optimization cites this paper.

Harmonia: End-to-End RAG Serving Optimization Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:04.566083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:04.566083Z digest=sha256:4a8df5f296cccf6426873ca87514c407d3b66a0d0b1ca806265a7f31f600d9b6

Observation 273bddff-6fb4-44da-a491-669bb0f35e2a · inbound

HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training cites this paper.

HALoS: Hierarchical Asynchronous Local SGD over Slow Networks for Geo-Distributed Large Language Model Training Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:46:06.563160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:46:06.563160Z digest=sha256:396bc78827d1a13bae8f887e0a325b705a5e14aff8f37982245702b2420a06d8

Observation a8e76d3e-e5dd-453b-8500-9c5428386ce6 · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.463491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:c1acc9a1d0f3949311b34f71da6721a78b1f53371cb335da5f288fad8a203db3