Pith. sign in

Paper Citation Record · LEDGER

FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2406.06858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06858 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:17:26.613570Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T16:37:23.032697Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b78cbcf5-abcd-4ff4-88d0-d514868036a1 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.905725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b9e3f73b19795a7b080affc5da9dd01eb269f8aed73b018cbb2650e28ce1a1cc

Observation 6234bd18-0851-4f19-bf58-ca78097e27f9 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.074100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:c2a2f0c687f12d2d76d400864ae81f5678c86ebd20e3807aa02753de16c80d9b

Observation 36560a66-9cab-4c92-b8b5-e8ead0d488de · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.613570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.613570Z digest=sha256:8143b433fcc0713d5090f2eccfeb03c8094537247527a0e8359ec1a74bfa590c

Observation 5a55e24b-f663-41e9-bce7-61a6d0ac6096 · inbound

Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training cites this paper.

Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:23:02.055646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:23:02.055646Z digest=sha256:4e41d311e96dea155e3b8e86b3b3667db127ca3183a287017971fdac6f6ea259

Observation 66a69f32-dddc-49f4-a323-e604a0241173 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.192888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:06:41.150046Z digest=sha256:099395947044e0419b3f0905cc1d6b8cc1b6970f66a4b97398d3433f70fd48a9

Observation 75a4fb2f-d353-4a89-af7c-d4aa47922422 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T07:21:23.050371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:21:23.050371Z digest=sha256:66e2f63fe9d707c42ab7155b0adc64018692531d113b994fe25d18b95f5c3966

Observation ded769fd-618c-4cfd-8e1c-05085d49689c · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:52.155399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:de559a058d99e00f2e44fc91d7c4490653792ddf0107d2c9a8fca95ed3cbc34d

Observation 149779f0-7572-4a28-bc9a-3fe110071f81 · inbound

UCCL-Zip: Lossless Compression Supercharged GPU Communication cites this paper.

UCCL-Zip: Lossless Compression Supercharged GPU Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:41:37.154204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T06:36:43.101810Z digest=sha256:12afacb501c6724f899aa80a97636830ffab92068e8705de77585a6e2d37e492

Observation 07f0e3c7-1f59-4311-8ea7-3f5818b42a26 · inbound

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training cites this paper.

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:22:20.756427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T02:20:00.625923Z digest=sha256:2cc2b75969c61f065004538a79121a5e47c024e1404e07f0c4ff6312aa089fc5

Observation b65b1355-0268-416d-a3a2-5ce181ef3430 · inbound

CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training cites this paper.

CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:41.075961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T04:25:11.075964Z digest=sha256:01f4fd73fbfcef175c12de65989457707ba054c9cdb32c69c6d1feea14fb27aa

Observation c38437b1-1815-4e53-a38c-390ddc3e9605 · inbound

Eliminating Hidden Serialization in Multi-Node Megakernel Communication cites this paper.

Eliminating Hidden Serialization in Multi-Node Megakernel Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:09.187848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T18:35:12.493334Z digest=sha256:cc3647a95ef116ae5efa82c9f027b0328cf2cdb363e325cb3e16fcb9bad14f40

Observation 5cde0bf3-bdb3-46fb-b1be-c9d9a26652be · inbound

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs cites this paper.

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:00:34.257503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T15:57:35.983246Z digest=sha256:63e4ef46dcc5adc8c2452055337f023623b6a6f36376270f5eb7b6b6a2e4f6d6

Observation 77edfa72-05a8-4614-b0ab-497e1a131a5b · inbound

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning cites this paper.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.822291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:2d6c513c4cad0adca86a39c26940121d4197ef841f308b3e069a36f891542f69

Observation ec067641-9638-4f11-b05a-fc3b1175138c · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.929283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:e0490ab03978f38686f2ffc53ef3efbd490284adeeefb09fd137b037a2634ec0

Observation 4c6317c2-539f-45bb-a7a9-2486fe295f97 · inbound

Accelerating Compound LLM Training Workloads with Maestro cites this paper.

Accelerating Compound LLM Training Workloads with Maestro FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.582680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T05:11:46.845356Z digest=sha256:b052cf2d0f65158d1df7de6feed14d839c2411f996da0e324010ff86450d966a

Observation d9d49c1d-c31e-41fa-b91b-1395b8b888dc · inbound

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments cites this paper.

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:56:21.770249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:55:15.475044Z digest=sha256:d54c681192e034fee7af195827e2cdeb2f92063b6f4b1a694a04d50a26efbfde

Observation 7c39fc60-040f-49e6-a4bd-16362d558f1c · inbound

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments cites this paper.

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:09:44.740788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T05:08:56.760052Z digest=sha256:e3bb5765bd33fc1bf5b47fa586b48c85473a5e88d05f71595b4579856651f148

Observation 33873138-bff7-4437-aa4c-9cef24c30730 · inbound

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM cites this paper.

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:57:44.426109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T19:56:25.334846Z digest=sha256:e76599234d759d040bf393e122795aa88781ce290e94c97fe0799984e2502240

Observation ae85bc83-67ca-4a7f-9913-bee03d9430b2 · inbound

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training cites this paper.

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:22:55.884632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T01:19:28.044984Z digest=sha256:7d45781f2dd3bd25e2cd675f384bdfc4cc308ee480b89ef68544f24908f7bc69

Observation cd568f4f-06ea-427e-9e82-52318b011536 · inbound

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling cites this paper.

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:36:17.107636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:34:47.923752Z digest=sha256:d49e12806726d542a973513d9c728e7166091acb710536985164459efe948820

Observation 960b7ab6-2eda-4316-ae6a-e842017e4303 · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.772900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T02:49:09.109990Z digest=sha256:9fedee39bc4692c65e9ebde88e0295d510499567cbe90697335f0d160727e46e

Observation c781b13e-235b-44cf-a188-0b1c603b0626 · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:14:46.959202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T15:10:25.071087Z digest=sha256:cd1fc9db692f6169d13858ec09907a5aaf96c13fe2bba7addaceb0f8622193bb

Observation e069f1e0-dcf9-42e0-810f-69d4f84114f1 · inbound

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads cites this paper.

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.978887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T15:00:07.823857Z digest=sha256:0bfe28a4e763fdef4818947b4ae7e7f6239c94f89385d9dbdd0f4723a00908b5

Observation f3af15ee-3cd5-4858-9954-930c5a3fe13b · inbound

Piper: A Programmable Distributed Training System cites this paper.

Piper: A Programmable Distributed Training System FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.661496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T11:34:02.562929Z digest=sha256:0b891fecdd83f646dd6c9773c40b9b67ca6ccc8f86527eb30042ff7d97c4ae83

Observation 60a77571-6f7c-4581-b4ad-e2f1058e3c8b · inbound

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems cites this paper.

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:37:23.034143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-10T16:30:09.561233Z digest=sha256:5f98f4f30ee41e6a93f46deb4ca84056716c42abda940bf17121eca12390473b

Observation bbed7484-4fc1-4b72-ac3d-6708d4c2539b · inbound

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference cites this paper.

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T01:18:48.061500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:18:48.061500Z digest=sha256:4b6fcc97fd7c567a7373b2c97fec2b689039f8c16520250934c5b425a2fd2081

Observation f31e81af-a01c-4acb-b7da-553543687284 · inbound

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference cites this paper.

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:00:11.581450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:00:11.581450Z digest=sha256:a18749960c27f75da864e7039a5036dc4f4fb04106c261c0e7252b4eb5e2ccff