Pith. sign in

Paper Citation Record · LEDGER

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2607.16100.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16100 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:25:20.987873Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved68
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a9ad259-59ce-4d54-81d1-e51c4da5b342 · outbound

This paper cites Deepseek-v3 technical report,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3 technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:13.775104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:13.775104Z digest=sha256:7853ccb0d82a484ac3121299041c1b86dee93542fd13b672933f685015d8376b

Observation 0b3788ff-7317-43c7-938d-9aa9caffd13f · outbound

This paper cites Deepseek-v3/r1 671b deployment guide: Gpu require- ments,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3/r1 671b deployment guide: Gpu require- ments,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:13.912073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:13.912073Z digest=sha256:62a2df9fc20ec5265f4cf636ce2f3788dbed71ee9271ae89301bc03776045d7b

Observation 6e35e9c3-5090-4f82-99a9-a745e6e9b133 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Efficient memory management for large language model serving with pagedattention,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.061863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.061863Z digest=sha256:971b3725d84d676b966de8b057f77a847346ee2582b3db710d65668e26da1c2b

Observation afbebda9-16f1-445a-8503-05689076a4b6 · outbound

This paper cites State of ai: An empirical 100 trillion token study with openrouter,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives State of ai: An empirical 100 trillion token study with openrouter,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.205913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.205913Z digest=sha256:f15b741f3484a6875cd9e87143fe5ec634bd3be8da570c61f19fc6ce87efe0bf

Observation c8f2e045-189d-482d-a104-95b85da321dd · outbound

This paper cites Sglang: efficient execution of structured language model programs,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Sglang: efficient execution of structured language model programs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.336839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.336839Z digest=sha256:cafa4f20d62d547662074d84b517ca16737dc93d5356639d866048ad983bd607

Observation f3903cc2-f50f-4753-8ce8-39a453417ea1 · outbound

This paper cites Tensorrt-llm: A library for optimizing large language model inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Tensorrt-llm: A library for optimizing large language model inference,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.520829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.520829Z digest=sha256:72853d5f04ebb05fa23c1bde29173fa1ef4f4dde248889034a042d3cc82c4224

Observation 1038d302-deab-4ff1-8a98-511433a776a5 · outbound

This paper cites CoreWeave Pricing: Instance Pricing,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives CoreWeave Pricing: Instance Pricing,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.834402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.834402Z digest=sha256:fbc664aee86679c53ac2ef30d3ee23736e16c1571c9009ce1f7fd970c7cb0c60

Observation 09056ba9-bbfb-4762-973b-7222605e7748 · outbound

This paper cites Scaling-up pytorch inference: Serving billions of daily nlp inferences with onnx runtime,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Scaling-up pytorch inference: Serving billions of daily nlp inferences with onnx runtime,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.004162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.004162Z digest=sha256:bb39e914f7aaab78d7b1f29e19f88d7e6bb0d391b145cd6094cd2f5c89af4ae5

Observation e789aa8c-d279-4c85-84b0-cf6a4576cf83 · outbound

This paper cites Understanding data movement in tightly coupled heterogeneous systems: A case study with the grace hopper superchip,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Understanding data movement in tightly coupled heterogeneous systems: A case study with the grace hopper superchip,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.133947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.133947Z digest=sha256:9a15ba686fd51751dcefde4df5f004e8710135f5f9a550e39e040c99fa4fb6e6

Observation ba6e4d6b-f423-46cb-a0eb-4d6c06e43fee · outbound

This paper cites Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms ,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms ,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.320821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.320821Z digest=sha256:097d00e1d6d974999f4f3b6564d90421db60bd9bd26160e7c75a1f11a2f24539

Observation e5220af4-d77e-4356-afdb-85ce704bc8c2 · outbound

This paper cites NVSHMEM communication library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVSHMEM communication library,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.476182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.476182Z digest=sha256:43e3c1a9aa8248a8aaf4de371a5efd1765c42cbc84a50e092c750a9402188adc

Observation f7731326-971c-4750-b79f-2deac8501581 · outbound

This paper cites ROCm Communication Collectives Library (RCCL) Documentation,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives ROCm Communication Collectives Library (RCCL) Documentation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.616175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.616175Z digest=sha256:832920c3c3c3dcecdb5d88249a1e9041237076be10621f672e79b4f9beea7988

Observation af79158d-7236-4e53-a8a2-d0bf6213a478 · outbound

This paper cites AMD ROCm Documentation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMD ROCm Documentation, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.792579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.792579Z digest=sha256:23209bfe4286d5fb2b930e5cc7fd27d17465224a2dc0310de44a890bc1a4a907

Observation 7d36d779-647d-4b94-b44a-7a884e24b51c · outbound

This paper cites Accessed: February 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: February 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.907979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.907979Z digest=sha256:e595e9800b61c24e2149f9280b8400d54ce841701b2cd3de7e9440da5bac11ea

Observation f4b5d931-844d-4bb2-b6c4-bc1ca39339b5 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.056139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.056139Z digest=sha256:acd86a047937c5be540cb3e89ae5fbe4c53dbefbaade258a8c13d7ad9c14fccc

Observation f1ca6079-e9d0-4b9a-9fd3-08e37a72c1ad · outbound

This paper cites An introduction to cuda-aware mpi,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives An introduction to cuda-aware mpi,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.165890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.165890Z digest=sha256:21d4af6c042be56d689d08ff423d2a062f734116994e43d8c3aafb70bccc6aad

Observation ae8a94c5-8d78-4b20-abc1-adb893da5acd · outbound

This paper cites Enabling fast inference and resilient training with nccl 2.27,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enabling fast inference and resilient training with nccl 2.27,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.359083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.359083Z digest=sha256:530301699ac0ab85f7dd928288195ec34dd656cce1a4072429eb4d2f24a12aca

Observation a34d0e29-9dd8-496f-aa53-0b6ec32830d7 · outbound

This paper cites Gpu-initiated networking for nccl,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpu-initiated networking for nccl,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.494336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.494336Z digest=sha256:b2ada653898d9738ec5b34003c99acdab14c3c1801d81ed6f2483888adb42d39

Observation 39a171ef-35a5-41c5-aca5-eb4f25696c3f · outbound

This paper cites NVIDIA, 2025.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.661046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.661046Z digest=sha256:edc0ac43dd5fa7a7163acefe0d51719c1d5d5776c3302fd2b6294ccd8f3da3dc

Observation eb7c6918-a749-4e52-9d5b-c23d4e978426 · outbound

This paper cites Fusing communication and compute with new device api and copy engine collectives in nvidia nccl 2.28,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Fusing communication and compute with new device api and copy engine collectives in nvidia nccl 2.28,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.784410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.784410Z digest=sha256:e4eb6d9d2f7a394d749bd139bb31f77458c1b470add3495d8c5067ba43b46ccc

Observation d9f3656d-35c1-48a9-bc7b-29758dbb8a80 · outbound

This paper cites PyTorch Foundation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch Foundation, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.857736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.857736Z digest=sha256:e78f38ab30143a48444dde488da15822cb417d4b17f2be09ead0370cf25a2f38

Observation 6553156a-7e56-4d0e-a740-6027b92cbe45 · outbound

This paper cites NHR@FAU, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NHR@FAU, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.919904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.919904Z digest=sha256:42f41ac8c76c42fdbec5cac0b777843c1ef48aaa22415e9e27b8c58f6cb702fa

Observation 472f42f1-784e-4951-96e8-a644b2f1a4ec · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.982682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.982682Z digest=sha256:1399f05e993a22af2b48582c5039442d0a4a9bec0764609e95e4fd6213a89178

Observation 1b5388cb-f193-45bd-85b0-93ad50f1a373 · outbound

This paper cites TensorFlow: Large-scale machine learning on heterogeneous systems,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives TensorFlow: Large-scale machine learning on heterogeneous systems,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.085329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.085329Z digest=sha256:afbc2f676a5899f248b48bd4de32130a90123e909293f00093a19d64b26c87bd

Observation 88c2afe9-aa80-4fc4-a87d-596cccc7a504 · outbound

This paper cites NVIDIA Corporation,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.160785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.160785Z digest=sha256:c690aeb117474afd3c84b9fd6a399aadd9a66de796c07240283fb1b8617a11a3

Observation b3a6a6a7-957e-44d1-ae52-321de4b59c6c · outbound

This paper cites NVIDIA Corporation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.331377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.331377Z digest=sha256:5f22e9d016e22ea22e7f976520d9b9a752d52328968acd614cd443ef10c7ee63

Observation 8c77e832-ca44-4a66-9721-44529684ea50 · outbound

This paper cites Llm inference beyond a single node: From bottlenecks to mitigations with fast all-reduce communication,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Llm inference beyond a single node: From bottlenecks to mitigations with fast all-reduce communication,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.416705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.416705Z digest=sha256:2e6569215733f2a6285e4edb736364e0dbe0e6981877acfa655565375016a8bf

Observation 5e9e99dc-cd28-492c-9afa-81c4ee1ec49f · outbound

This paper cites Msccl++: Rethinking gpu communication abstractions for ai inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Msccl++: Rethinking gpu communication abstractions for ai inference,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.512162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.512162Z digest=sha256:2935f4fc121e46bed822779fc93ba9fdbd1d10d5a7c8bc316c611d79990bf859

Observation 555e25db-a80a-457d-a863-9cd2a6899968 · outbound

This paper cites Nvidia gtc 2025 - built for reasoning, vera rubin, kyber, cpo, dynamo inference, jensen math, feynman,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nvidia gtc 2025 - built for reasoning, vera rubin, kyber, cpo, dynamo inference, jensen math, feynman,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.600194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.600194Z digest=sha256:9784ad637a6c9b75cad9d4e5d3749e58ec5091a0f59476d2680ef3be2812f5da

Observation 07f573ca-22b3-4bc6-8073-ee66872fefe8 · outbound

This paper cites Colossus: xai’s supercomputer for grok,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Colossus: xai’s supercomputer for grok,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.688668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.688668Z digest=sha256:a2b9108d1638c8b645c7a3e2794728039aacba6def2e60cb120f4b597a8b86f1

Observation 84ec75fb-c0dc-46ea-83d4-410030bb12c2 · outbound

This paper cites Recent improvement to open mpi allreduce and the impact to application performance,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Recent improvement to open mpi allreduce and the impact to application performance,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.756420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.756420Z digest=sha256:1c64e463d15345128e3ea49ad70321d7d46f7a841301f77f2d0041f82589a2c7

Observation 749fc5d6-ab5c-4ce0-b39b-d74a13e3be6f · outbound

This paper cites Revisiting the time cost model of allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Revisiting the time cost model of allreduce,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.859961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.859961Z digest=sha256:18254003869e2f982e69b48ba0de27979192a73d9e1150acc716ddb58055e3e6

Observation fafed51a-9d45-45a3-a82e-2b0eddf758a9 · outbound

This paper cites xccl: A survey of industry-led collective communication libraries for deep learning,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives xccl: A survey of industry-led collective communication libraries for deep learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.951226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.951226Z digest=sha256:abce0cac84475094e9d9051d09a37657f62487dea32c130f19357f2f84ca0e81

Observation f7d1e527-d487-4679-9914-4503d0286d0f · outbound

This paper cites Hear: Homomorphically encrypted allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hear: Homomorphically encrypted allreduce,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.026695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.026695Z digest=sha256:94cc07a9d36fd4606471d643ac2b41ab724448d6a4ee6d7ad417add3f7139f5f

Observation fac11bdf-0895-49b3-8972-47ca79be7a87 · outbound

This paper cites A survey of mpi usage in the us exascale computing project,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A survey of mpi usage in the us exascale computing project,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.218838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.218838Z digest=sha256:f95c8ccd6af0a423a6add9932ab6ff46ca71341f35f54bcff9708d2546a9f94d

Observation 7a162cce-669a-4284-ad8d-842c61632797 · outbound

This paper cites A large-scale study of mpi usage in open-source hpc applications,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A large-scale study of mpi usage in open-source hpc applications,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.463178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.463178Z digest=sha256:afdd8d69a237f62faaf92682686291ed2cf060c2ebb278bdb10a04005e8e387d

Observation a0ae434e-702d-4a55-9279-4ed928885529 · outbound

This paper cites Short-circuiting rings for low-latency allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Short-circuiting rings for low-latency allreduce,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.590758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.590758Z digest=sha256:79cd704fbd6a64d743a0f12bb383b1fd2448da12c8393cf5190dec8019c00247

Observation f41d774e-b592-4c46-a878-9c84b31b3ad8 · outbound

This paper cites Cuda c++ programming guide.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Cuda c++ programming guide

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.684480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.684480Z digest=sha256:e77a80a1c0af7555becce0839832cc3c4f4fd696d48f67f700a6e49aa1b5c593

Observation 0252082f-c101-40bb-ae82-dc161755edab · outbound

This paper cites Parallel thread execution isa version 8.0.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Parallel thread execution isa version 8.0

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.772103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.772103Z digest=sha256:c0cb35cdcbd8a98233a5a2c67a8c432c78a6a54212f445a465269aeda0fb6c25

Observation afb4469f-2c07-4271-a037-cdab3a06269b · outbound

This paper cites vllm container (version 26.02-py3),.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm container (version 26.02-py3),

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.845547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.845547Z digest=sha256:93eec01b07805475d38a2ace40944597732953d6e4c15bbba3b0b6f41bb13f59

Observation 784214d7-9123-4747-8d89-936f6bdeeace · outbound

This paper cites Network-offloaded bandwidth-optimal broadcast and allgather for distributed ai,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Network-offloaded bandwidth-optimal broadcast and allgather for distributed ai,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.896433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.896433Z digest=sha256:12c230552c5e9697d778d3d47275c7e6bcfec0c3e8c57cb2e57dc5f8258c76be

Observation 2f37d146-a83d-4999-ac3a-30f47a9a516a · outbound

This paper cites Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.083488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.083488Z digest=sha256:a625fd1a54a913fd29fd0762eacc8686832dc30c5d1fe5ab7c2a694ad4eda987

Observation 8701c828-c2be-4774-8eba-d71b78d59efd · outbound

This paper cites vllm – technology radar entry,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm – technology radar entry,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.181539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.181539Z digest=sha256:747399fe36014da2ebb66b7718913198b7825e3d195cefccbb8d106824af835a

Observation 323375ca-c4ab-487e-ba67-eee2cd45c0fe · outbound

This paper cites The huge potential implications of long-context inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The huge potential implications of long-context inference,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.273592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.273592Z digest=sha256:fa5095e3469bb97377b960cf6e735b867f44bdefcd251eb2bd222e974af2eb67

Observation 0ffa4cbf-30c7-47a0-80ab-8881c503fa2b · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.006976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.006976Z digest=sha256:7c0c4e9111d576285b2a51f40954a9dddefa96b5cda3979ca66c36fd85fb8cc4

Observation c1127a6f-4641-49b7-aa98-7e64e2d365ea · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Chain of agents: Large language models collaborating on long-context tasks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.440969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.440969Z digest=sha256:82ae990434fcc1d7063ba20b3507f0d0efdcdd84b2ecc71814948577c78830a2

Observation 9748d69e-2b66-4ed0-b867-61dc58a6b85b · outbound

This paper cites Breaking the boundaries of long- context llm inference: Adaptive kv management on a single commodity gpu,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Breaking the boundaries of long- context llm inference: Adaptive kv management on a single commodity gpu,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.528704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.528704Z digest=sha256:46ec9c429a714b83b78c90a2b40d96736400e2e5e632afa004e8193d1855f783

Observation 1d5f83da-1bb3-438b-abbe-41a0e30dcfb5 · outbound

This paper cites Prefill-decode disaggregation.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Prefill-decode disaggregation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.629418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.629418Z digest=sha256:c463ae42b82a63e308fdc93bf6a596d6093db65f46ff6798290a8acc6d796070

Observation 8aafc676-e597-4ac0-ac20-bcf8c7c6341d · outbound

This paper cites Is long context all you need? leveraging llm’s extended context for nl2sql,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Is long context all you need? leveraging llm’s extended context for nl2sql,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.382129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.382129Z digest=sha256:e14dc2c27e4e5e1daefb17b0451c4ef2ac8c8f4912579dfb72cc4d60323f1b96

Observation 7e090cab-d299-4f45-92fe-6193df7732e8 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.766032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.766032Z digest=sha256:db21d91e264c514c0d8e1d379834635df7bdc13a214c9f956badc6d3af0ca1fa

Observation 874e5964-549e-489c-a52b-7356e2e0dc95 · outbound

This paper cites Accessed: March 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.916482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.916482Z digest=sha256:7ddb0f6782f5fa96f32ff3c11c08efdb81f5e7bde8bfb11ccc5fea09399b9cae

Observation 4f1b716c-61c9-489b-b789-8d1fbb58b7ae · outbound

This paper cites Amrex and pyamrex: Looking beyond the exas- cale computing project,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Amrex and pyamrex: Looking beyond the exas- cale computing project,

Reference 54

Resolution
verified exact
doi, observed 2026-08-01T21:28:32.655939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-01T21:25:19.998763Z digest=sha256:46593ffcf792cab5995a4d6911329b2e83ce400eb7deed057d146ab1b4e9d9be

Observation 91ae64af-c590-47a8-b86f-c6d071eaf952 · outbound

This paper cites Slo-aware compute resource allocation for prefill-decode disaggregated llm inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Slo-aware compute resource allocation for prefill-decode disaggregated llm inference,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.670034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.670034Z digest=sha256:5c8b47e045fd0ad61a19e1535efc067f79e0cce3ee6d218a0f5467c28a5e34d1

Observation 2a329e9e-4041-45be-90bd-4bc80afa968c · outbound

This paper cites Qmcpack: an open source ab initio quantum monte carlo package for the electronic structure of atoms, molecules and solids,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack: an open source ab initio quantum monte carlo package for the electronic structure of atoms, molecules and solids,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.252829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.252829Z digest=sha256:bdffec835be93fa7e435b636a2f45df537f43a563ba43d80f377cd06116a19b2

Observation d3c997e9-79f3-4624-b9f7-6a224d926070 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 57

Resolution
parse uncertain
no resolver link, observed 2026-08-01T21:25:19.821885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.821885Z digest=sha256:77263eafb9750cc9055dd49ef351867cd3c2a8226b3619ced26a4f006a5c8f33

Observation 6005d775-7c55-4096-9985-7542cd1c6bc9 · outbound

This paper cites NVIDIA HPCG Benchmark,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA HPCG Benchmark,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.385515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.385515Z digest=sha256:d1be8af930ec25eb6548f47c136653bbbcb7a877858abe99363a94d9e90996e6

Observation 7a2f491e-e59a-479f-8ee1-c644e10c9fc3 · outbound

This paper cites Accessed: March 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.464831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.464831Z digest=sha256:fdda0240c62520a299a76899b3b0e1d1ce7b42c967de5d9022e64f3654c2ffa6

Observation be738fcb-82d0-463c-96d2-c14d401588b6 · outbound

This paper cites AMReX issue #4821: Device-initiated collectives in NCCL/RCCL.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMReX issue #4821: Device-initiated collectives in NCCL/RCCL

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.077961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.077961Z digest=sha256:e6fd644fef955bdd964eba1bb977ef350d2ac86edd09729757f3ccf047c921c8

Observation 4702eab3-4a03-4345-9f35-ca16d8197c20 · outbound

This paper cites New research infrastructure: ’alps’ supercomputer inaugurated,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives New research infrastructure: ’alps’ supercomputer inaugurated,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.626672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.626672Z digest=sha256:26f41153938aa1ed4ae6c45459e890c8223ca5dfbc465bd9835ff711214d59d3

Observation 1f8ab145-110b-45b5-b5cf-879508324856 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Flashinfer: Efficient and customizable attention engine for llm inference serving,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.731027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.731027Z digest=sha256:23d8e305f2191b4be49ce25e7b48a887b83017bddca55e1b1a28a2bc26dbb92d

Observation cd35d561-5800-452b-ba99-0071dc64e305 · outbound

This paper cites Qmcpack issue #4654,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack issue #4654,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.332056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.332056Z digest=sha256:b54c36a7ca02c8a63b7ee3526fbc14d3329f3c9ad69eb958f9ce622d7ffea386

Observation 90e64a8e-5e91-4298-a3e9-ab2d51c5fe27 · outbound

This paper cites Nccl ep: Towards a unified expert parallel communication api for nccl,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nccl ep: Towards a unified expert parallel communication api for nccl,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.853881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.853881Z digest=sha256:d694d7e2e1c4f39446e0b456ddca4fd94d3ef8ad131abc8990e3112f73527325

Observation 02d42690-5e33-42c5-80e1-106ff10d1444 · outbound

This paper cites Enhancing dis- tributed inference performance with the nvidia inference transfer library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enhancing dis- tributed inference performance with the nvidia inference transfer library,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.934999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.934999Z digest=sha256:408ff4251064ac2d53f27ebd3be8afdab27c1bbb48f9333db1ced731bdfabde7

Observation 29b8628c-a227-432b-84c2-a7050299050e · outbound

This paper cites The elpa library: scal- able parallel eigenvalue solutions for electronic structure theory and computational science,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The elpa library: scal- able parallel eigenvalue solutions for electronic structure theory and computational science,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.549073Z digest=sha256:e3e5c44b9d94478f8d3b8fb310790200671f001c669a7778e1ae66bdbb62f46d

Observation 0836cd97-20e2-42c9-abfe-54893ca7b307 · outbound

This paper cites Deepep: an efficient expert-parallel communication library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepep: an efficient expert-parallel communication library,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.804322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.804322Z digest=sha256:dcfb141dbab5d1914bad5f37a141ab61b76782131f43b9cda5b43b186634960d

Observation aa173ed3-ed6f-40fd-bc13-383abd1e9ad4 · outbound

This paper cites Gpt-5.4 chatbot,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpt-5.4 chatbot,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.987873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.987873Z digest=sha256:02fc39b62263802790409509f67bdf321b1c94932f02cf7120d45a6375007314

Observation ae2cf742-ceec-41a6-b499-f3cdccdb1871 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.335994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.335994Z digest=sha256:eda1f08605ad3318602bcf7db900ee15c38e135035d16b0f4758edbf24e40cf7

Observation 6947ed3d-468a-4bd4-b977-0ba2f5b4af63 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.096674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.096674Z digest=sha256:a610cbe74ca00ead373761cc8cc737ff4cd0c8b3e40ba98a5f97d4da126ba046

Observation f0c5898d-7da4-438e-8d51-52d3a6726d18 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.236864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.236864Z digest=sha256:8cccf9a3aaefc11effde6aba49f93371267763a22681a4c42d85a2e0b8197a7c

Observation 78451cc0-945c-465e-902a-fe41b4d093f9 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.176773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.176773Z digest=sha256:8c2e1f5ac49cc2d77d3920ef316eca82c4f16ae166199286f8afee7d9c06dde4

Pith citing papers

No inbound Pith citation observations are available.