Pith. sign in

Paper Citation Record · LEDGER

FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2406.06858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06858 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:26:47.368308Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dfef0897-770b-4fd4-a0a1-ed55b8188b05 · inbound

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference cites this paper.

Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:11:39.060378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:11:39.060378Z digest=sha256:e700ba11418ca54d9a8cc9b94d6f18c7a853a40027cbecf46ae2d2c6b2297ac4

Observation d09eaa3c-6e16-49b1-a691-5e8124ab3bba · inbound

Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking cites this paper.

Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:02:00.256179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:02:00.256179Z digest=sha256:174d5c3b934dd7a2973ece3e5b526580f8e2261a72e0305bf6a995bfdfba45a2

Observation ceb8792e-e274-40fb-bafe-20a3cce01657 · inbound

PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving cites this paper.

PRESERVE: Prefetching Model Weights and KV-Cache in Distributed LLM Serving FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:34:07.768670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:34:07.768670Z digest=sha256:abb5b0be33c8c56cde7e461ddb9b121252fb1d0427c89edc5e5234bd1226ba60

Observation 0dfe1503-fe16-4ecd-9e8c-9c6b07ef8dd9 · inbound

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models cites this paper.

FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:05:55.126143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:05:55.126143Z digest=sha256:44b710509aaa7983ce80ca0628773077f6d1494eb7b8230c4cb928fa0815830d

Observation b78cbcf5-abcd-4ff4-88d0-d514868036a1 · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.905725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7820086b06a433cb0134308f5ec90e8ef014364586ccbc43b64c9652b0c0e076

Observation 5a8e9774-b5e6-4664-8e58-07f53de3b53d · inbound

H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips cites this paper.

H2:Towards Efficient Large-Scale LLM Training on Hyper-Heterogeneous Cluster over 1,000 Chips FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:17.541894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:17.541894Z digest=sha256:419082083ee5a16643f6a0cfa7b974dcd02561ffe62466b9808757f5504a3ace

Observation 12e87fc3-18af-4c55-805c-fde39d1251fd · inbound

KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads cites this paper.

KPerfIR: Towards an Open and Compiler-centric Ecosystem for GPU Kernel Performance Tooling on Modern AI Workloads FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.095644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.095644Z digest=sha256:0f8b83af3fec7b4e17ac6324d9c989045458e414b41644d82be12a24197b107c

Observation 169b84dd-eb35-4d65-8765-b7364f8470bc · inbound

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure cites this paper.

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:50.463847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:50.463847Z digest=sha256:9f7309783584f6a9416e55bb8b2d6731c1198fe2b3ec0bfc64e1a968c1327b80

Observation 6234bd18-0851-4f19-bf58-ca78097e27f9 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.074100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:85c7340954aa2bbde085a1cf26a19fe783d0250ace4887f817f3a3fbcbc9c731

Observation 36560a66-9cab-4c92-b8b5-e8ead0d488de · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:26.613570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:26.613570Z digest=sha256:ebffeafd4eafd927e4b2951fcc35f8f96c79429ae221fb23097a3ab297ac686e

Observation 5a55e24b-f663-41e9-bce7-61a6d0ac6096 · inbound

Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training cites this paper.

Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T08:23:02.055646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:23:02.055646Z digest=sha256:c4af60f1ca8f18bb0b84f8f3c6685d63b78751aa3cfa4984c51e28ad1378e47a

Observation 66a69f32-dddc-49f4-a323-e604a0241173 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.192888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T10:06:41.150046Z digest=sha256:bf27cd21b3f0a65875255ab3b7827dbda67bfd984d4b53849621ddec7061c748

Observation 75a4fb2f-d353-4a89-af7c-d4aa47922422 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T07:21:23.050371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:21:23.050371Z digest=sha256:66a77b6c2b245401ee6bc5010d49322fef02bb14b571b3d4cba751c3aa0f5f7f

Observation ded769fd-618c-4cfd-8e1c-05085d49689c · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:52.155399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:9ce05c59f0637391f8112127b5c4af6171acfeb518be705f819fd84cca6dcb66

Observation 149779f0-7572-4a28-bc9a-3fe110071f81 · inbound

UCCL-Zip: Lossless Compression Supercharged GPU Communication cites this paper.

UCCL-Zip: Lossless Compression Supercharged GPU Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:41:37.154204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T06:36:43.101810Z digest=sha256:e77462ed8f9ce92c990699c1992f21e5d15537ac665cf9375853e2709a04f978

Observation 07f0e3c7-1f59-4311-8ea7-3f5818b42a26 · inbound

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training cites this paper.

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:22:20.756427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T02:20:00.625923Z digest=sha256:310572de290b48bad3217b7d44bd0e944e6c31b7db1c5bc19fb2f7cfefadb892

Observation b65b1355-0268-416d-a3a2-5ce181ef3430 · inbound

CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training cites this paper.

CommFuse: Hiding Tail Latency via Communication Decomposition and Fusion for Distributed LLM Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:46:41.075961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T04:25:11.075964Z digest=sha256:4650b1e2297842bfd323261929e72fb91070648c592cc8a9d9f671eecf760b0c

Observation c38437b1-1815-4e53-a38c-390ddc3e9605 · inbound

Eliminating Hidden Serialization in Multi-Node Megakernel Communication cites this paper.

Eliminating Hidden Serialization in Multi-Node Megakernel Communication FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:11:09.187848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-09T18:35:12.493334Z digest=sha256:9996b8469a26b1e1bec72d76b0daad6592a9bbfb72ac8c7017653316e7eba454

Observation 5cde0bf3-bdb3-46fb-b1be-c9d9a26652be · inbound

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs cites this paper.

DITRON: Distributed Multi-level Tiling Compiler for Parallel Tensor Programs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:00:34.257503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T15:57:35.983246Z digest=sha256:24260edbf0e373d3f7d39a4ac63c251599f1cb59102bc0e69a7395b3753d8530

Observation 77edfa72-05a8-4614-b0ab-497e1a131a5b · inbound

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning cites this paper.

ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:46:14.822291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T01:38:03.841465Z digest=sha256:be92b33fa3359104bd42ade3ac10326799a091a92be788f8d777bd55223c6e66

Observation ec067641-9638-4f11-b05a-fc3b1175138c · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.929283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:0b2713e96a582feb2ec76dd50f1202f37fc0b376d53d1f9052ccf157e4366798

Observation 4c6317c2-539f-45bb-a7a9-2486fe295f97 · inbound

Accelerating Compound LLM Training Workloads with Maestro cites this paper.

Accelerating Compound LLM Training Workloads with Maestro FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:25.582680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T05:11:46.845356Z digest=sha256:c670be10b3618a9c91b05fccf0d37732b399969458171df7ef2f27ba5db28e96

Observation d9d49c1d-c31e-41fa-b91b-1395b8b888dc · inbound

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments cites this paper.

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:56:21.770249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:55:15.475044Z digest=sha256:7c2c2ddcd90250c6da7a9f74c8f84ddbac6553853f5969448b3df4611833a846

Observation 7c39fc60-040f-49e6-a4bd-16362d558f1c · inbound

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments cites this paper.

TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:09:44.740788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T05:08:56.760052Z digest=sha256:23caa03115e831ef2ca179716d0b43087fe8b41cb62cb5db0bd2316dbd57d197

Observation 33873138-bff7-4437-aa4c-9cef24c30730 · inbound

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM cites this paper.

A Few GPUs, A Whole Lotta Scale: Faithful LLM Training Emulation with PrismLLM FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:57:44.426109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-19T19:56:25.334846Z digest=sha256:22ceedd9ce9630731ec6d4ee75e445fefbf735fdeb61a9938f7e069a06dca2cf

Observation ae85bc83-67ca-4a7f-9913-bee03d9430b2 · inbound

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training cites this paper.

Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:22:55.884632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T01:19:28.044984Z digest=sha256:46cf066f44ac558dd7abfe549fef5d3864aa71eb2adcda27dd114c0d33576dca

Observation cd568f4f-06ea-427e-9e82-52318b011536 · inbound

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling cites this paper.

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:36:17.107636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T08:34:47.923752Z digest=sha256:858ddba389cb6b584e8ba0357695bbffe3e1e2197b6d700fcedc4f6751a39c4b

Observation 960b7ab6-2eda-4316-ae6a-e842017e4303 · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.772900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T02:49:09.109990Z digest=sha256:449076bf781492a576d1ebc6fc4a150bde0137497769f104141455971407fcbd

Observation c781b13e-235b-44cf-a188-0b1c603b0626 · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:14:46.959202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T15:10:25.071087Z digest=sha256:fc2fe86ea7861dd25178734b936c3b854dd4e9e34614067d0da132a32f6a7ed0

Observation e069f1e0-dcf9-42e0-810f-69d4f84114f1 · inbound

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads cites this paper.

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.978887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T15:00:07.823857Z digest=sha256:0eff914c5fc56e724640f5a8dc78cbfeafa0f5a6504d05a76f369b6756da7a56

Observation f3af15ee-3cd5-4858-9954-930c5a3fe13b · inbound

Piper: A Programmable Distributed Training System cites this paper.

Piper: A Programmable Distributed Training System FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.661496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T11:34:02.562929Z digest=sha256:ebcb5bf279408389d3229e42fe23e1ec367a99c43670adec3f8ff9a4314cdc90

Observation 60a77571-6f7c-4581-b4ad-e2f1058e3c8b · inbound

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems cites this paper.

CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-10T16:37:23.034143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-10T16:30:09.561233Z digest=sha256:66f17e7026ecb7458dd3656650628657244318422aebc94d85aac1d7977176c8

Observation bbed7484-4fc1-4b72-ac3d-6708d4c2539b · inbound

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference cites this paper.

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T01:18:48.061500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:18:48.061500Z digest=sha256:665594377f70d00e4f2836eae5892915122a006c4e92abca6c3bee2ac39099ae

Observation f31e81af-a01c-4acb-b7da-553543687284 · inbound

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference cites this paper.

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T00:00:11.581450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:00:11.581450Z digest=sha256:f026e6ac81756a318562af65271fb7c3c9e2b67e08b1ea17c75d0d6b85edce4b

Observation f74c1e6c-f28a-4de2-8e5c-1af97a64e03a · inbound

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization cites this paper.

SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:26:47.368308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:26:47.368308Z digest=sha256:2859c0e48aaaefc359749f8f67ac5d6fe2026a1ff6361d6147fa44289f2e3a8d