Pith. sign in

Paper Citation Record · LEDGER

Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2504.19442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19442 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T17:17:35.012041Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:37:44.599691Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b82ecf43-30b1-4223-95f0-bf96799fa64f · inbound

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference cites this paper.

TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:26:40.426398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T14:25:14.150581Z digest=sha256:5e0f176a188f337d604e08c42d402850b14da61a45c0f4ccd39c7af1646a92ae

Observation 6022ff37-35f4-49d2-8ed7-d9d7a4d24545 · inbound

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication cites this paper.

DMA-Latte: Expanding the Reach of DMA Offloads to Latency-bound ML Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:30:33.083551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T00:27:43.845408Z digest=sha256:be7539e265146fde697d05de28a8aabb717d3d9596e2514c04bcb544244a39d0

Observation 5f4c9bf8-c59f-4d2c-819f-3ee88f462063 · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:35.012041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:35.012041Z digest=sha256:8d64b701d37b12c1b9c7c5a13fa9d273e5acb969c86ffbcad3eda99f8866f833

Observation aad259e2-f4d6-4b9d-a755-929a04214a70 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:07:43.209953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T10:06:41.150046Z digest=sha256:cf2a5ffd3d8e73bd29ed80b8346ac4c65ac75a6fb6de170dc90fbfab36a12f8a

Observation 12a7ce50-5cc3-4baa-9d83-7e4967d02403 · inbound

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap cites this paper.

Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication Overlap Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T07:21:23.181366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:21:23.181366Z digest=sha256:042619e59ae135908fc9ea531f4933ad3663982392bdf5a77dd9b1df779ff94d

Observation bb172f6c-e390-4cab-ad04-4f2e76452a46 · inbound

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators cites this paper.

DeepStack: Scalable and Accurate Design Space Exploration for Distributed 3D-Stacked AI Accelerators Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:51.981196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T19:04:17.725111Z digest=sha256:38395c9e90ff540c55d33542d7ab93298cf1e767d9efd5aec9798947b848cbf9

Observation f4df9561-beba-4fdf-938e-bef5ded9aed7 · inbound

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel cites this paper.

Event Tensor: A Unified Abstraction for Compiling Dynamic Megakernel Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:36:02.757348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T13:45:53.346024Z digest=sha256:f01a54a0aa7467cfe53741d858f4c5fec77b73031b11e4b440ae34df9ecdb2f0

Observation 3ad8e02c-cd63-41e6-a347-8849af033675 · inbound

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training cites this paper.

UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.137902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T02:20:00.625923Z digest=sha256:55e92c36830e5e8b5381ea96f33762e7536752bafb65a3b0abed537c907dd8fa

Observation 504f7409-3dba-4dc3-89b8-4e17ce7042d4 · inbound

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training cites this paper.

FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:05.339124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T01:14:18.241481Z digest=sha256:dec9ae16c2f2a15c2ba803bcfb3fcbd901c88996c2323c20ada4d6660565769c

Observation 2014761c-7057-48e6-950b-90133d74b0ac · inbound

Eliminating Hidden Serialization in Multi-Node Megakernel Communication cites this paper.

Eliminating Hidden Serialization in Multi-Node Megakernel Communication Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:11:09.153963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T18:35:12.493334Z digest=sha256:0b9d42289c6867733774f2a7b0f63f121881054087159c043f42dfe29102a22c

Observation 3cf7948d-804f-48db-afe3-8846d520fcb2 · inbound

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding cites this paper.

NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:53:54.557117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T01:51:38.671365Z digest=sha256:01016d631c24de4c4336727147d89c13d5aa2f93e6c544a41c7ecfd04143aa88

Observation d631532b-9e71-4ecb-b590-2f9b335e9624 · inbound

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling cites this paper.

DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:36:17.082595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T08:34:47.923752Z digest=sha256:f7d6c9cda37f6e671802363b1f68dee15354e1cd5c9a448476315f91f2f3faf7

Observation f7e1ebf8-4eba-42e4-8548-ac37829603ee · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.750739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T02:49:09.109990Z digest=sha256:584d6bae1263910d032a6010089b375e61a22b91a260032ffcf5f18fae637a07

Observation 9b15cc1c-a649-46f9-a1de-332a5b2edf7b · inbound

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs cites this paper.

HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUs Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:14:46.917726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T15:10:25.071087Z digest=sha256:8697978145a15c9f4fad6e9d9bcd706803ed6975cc63837f08a2ef980377a546

Observation 6fe09191-b2fe-407e-b7ca-9fd37eccf6a8 · inbound

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters cites this paper.

HetCCL: Enabling Collective Communication For Mixed-Vendor Heterogeneous Clusters Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:22:37.227181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T20:16:40.966177Z digest=sha256:ce6435c647f3c0fd93e057f1234be737cf9007a9929c992f8d29d551955c0891

Observation f7a96141-20b8-4940-b17d-16c900ec5a1e · inbound

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Finite Elements in Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:37:44.601363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-03T07:37:40.688370Z digest=sha256:5cadbe909527f14c6501be4ecb81f2ccceca76fc9c33213f5c48fb452e7c5c17

Observation a3b2296a-e9f8-4a83-89c6-0fc88688d802 · inbound

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage cites this paper.

Coalesced Matrix-Free Geometric Multigrid on Persistent Cell-Wise Storage Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T02:40:50.976808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:40:50.976808Z digest=sha256:1949ff4fabb380b6dfa1cf25cd30a1a4da07a0630993df8fe0d2da7fdb3f33c0

Observation 448dfd36-bda8-4f71-b366-2421e9937c5a · inbound

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference cites this paper.

SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-13T01:18:48.061500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T01:18:48.061500Z digest=sha256:fa1a674c641d0b48f35b9aa7a2c57c423373dded8fb50a371865920d6ede2355