Pith. sign in

Paper Citation Record · LEDGER

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2607.16100.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16100 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:25:20.987873Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved68
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4a9ad259-59ce-4d54-81d1-e51c4da5b342 · outbound

This paper cites Deepseek-v3 technical report,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3 technical report,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:13.775104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:13.775104Z digest=sha256:5cb737e517ee809a583c4548e8dd593aa4a90571607b5545b0e103677ab6d97a

Observation 0b3788ff-7317-43c7-938d-9aa9caffd13f · outbound

This paper cites Deepseek-v3/r1 671b deployment guide: Gpu require- ments,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3/r1 671b deployment guide: Gpu require- ments,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:13.912073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:13.912073Z digest=sha256:5aab262db4e0c97e5bee4824ab00000ff5a7e1ade29ea4466bc56b6a85b96097

Observation 6e35e9c3-5090-4f82-99a9-a745e6e9b133 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Efficient memory management for large language model serving with pagedattention,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.061863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.061863Z digest=sha256:1c2315855dd4bf18fd45f8b225996eb4963b8ae4e06c6b40458cbcaf0e8c3fff

Observation afbebda9-16f1-445a-8503-05689076a4b6 · outbound

This paper cites State of ai: An empirical 100 trillion token study with openrouter,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives State of ai: An empirical 100 trillion token study with openrouter,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.205913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.205913Z digest=sha256:219b3658ace673b69f841fafa62331ede243b38ac2ee5638cd380722f88c8417

Observation c8f2e045-189d-482d-a104-95b85da321dd · outbound

This paper cites Sglang: efficient execution of structured language model programs,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Sglang: efficient execution of structured language model programs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.336839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.336839Z digest=sha256:6ed56e37db6374fc9d6c5504c42ee61b4fee297045b9cd063f445eee88238454

Observation f3903cc2-f50f-4753-8ce8-39a453417ea1 · outbound

This paper cites Tensorrt-llm: A library for optimizing large language model inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Tensorrt-llm: A library for optimizing large language model inference,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.520829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.520829Z digest=sha256:7aafc890f615102a53a151a9bfd8a03d81ade636f4f1f0e302b3a604d11f4efb

Observation 1038d302-deab-4ff1-8a98-511433a776a5 · outbound

This paper cites CoreWeave Pricing: Instance Pricing,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives CoreWeave Pricing: Instance Pricing,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:14.834402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:14.834402Z digest=sha256:853a3d73f37e292f7dd912f5f7912c05133a6c86d827df458dde5e28d36ec23f

Observation 09056ba9-bbfb-4762-973b-7222605e7748 · outbound

This paper cites Scaling-up pytorch inference: Serving billions of daily nlp inferences with onnx runtime,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Scaling-up pytorch inference: Serving billions of daily nlp inferences with onnx runtime,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.004162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.004162Z digest=sha256:b023dc80f937fdcf62994941e6569c51cd00428e46da82f58d5b3ecb447c3c1b

Observation e789aa8c-d279-4c85-84b0-cf6a4576cf83 · outbound

This paper cites Understanding data movement in tightly coupled heterogeneous systems: A case study with the grace hopper superchip,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Understanding data movement in tightly coupled heterogeneous systems: A case study with the grace hopper superchip,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.133947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.133947Z digest=sha256:649ddc697e8ded361431fabb51647727fdebe37cbbbc15364f41a7d029c644da

Observation ba6e4d6b-f423-46cb-a0eb-4d6c06e43fee · outbound

This paper cites Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms ,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms ,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.320821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.320821Z digest=sha256:5ec92823012256422b49294ccb1a7f9c4e1e42f2244368d9192a84ba3066772e

Observation e5220af4-d77e-4356-afdb-85ce704bc8c2 · outbound

This paper cites NVSHMEM communication library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVSHMEM communication library,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.476182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.476182Z digest=sha256:0b9f21dad3fb2c3fd58a893436001ba7667ee6ec1851e0ada67a12fd676ae72e

Observation f7731326-971c-4750-b79f-2deac8501581 · outbound

This paper cites ROCm Communication Collectives Library (RCCL) Documentation,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives ROCm Communication Collectives Library (RCCL) Documentation,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.616175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.616175Z digest=sha256:0d027010718950cd2f1ce49bbe57a0de52fb685d80c86523437d2d0172c7e29a

Observation af79158d-7236-4e53-a8a2-d0bf6213a478 · outbound

This paper cites AMD ROCm Documentation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMD ROCm Documentation, 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.792579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.792579Z digest=sha256:d69778561fead900cc162a5e356b42dd88f085aa6cd0caf6bb85fc158e9d9710

Observation 7d36d779-647d-4b94-b44a-7a884e24b51c · outbound

This paper cites Accessed: February 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: February 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:15.907979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:15.907979Z digest=sha256:4d2bf7a2cabce19e18d2c431d4a7825073b6336325b5f537dbe98306c3baf36e

Observation f4b5d931-844d-4bb2-b6c4-bc1ca39339b5 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.056139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.056139Z digest=sha256:38e2ee1a032142c31c31b48ec9131ce9b9ec55bfa272f4c6e0131ca2c189483b

Observation f1ca6079-e9d0-4b9a-9fd3-08e37a72c1ad · outbound

This paper cites An introduction to cuda-aware mpi,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives An introduction to cuda-aware mpi,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.165890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.165890Z digest=sha256:4bfd994445af851caa0a39e2e33befb69b1583e4425e11be1803eac03e7a3013

Observation ae8a94c5-8d78-4b20-abc1-adb893da5acd · outbound

This paper cites Enabling fast inference and resilient training with nccl 2.27,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enabling fast inference and resilient training with nccl 2.27,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.359083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.359083Z digest=sha256:a7fe014379edddc9250d5fc233417365f7f5fd274289187ff415495153f1caa5

Observation a34d0e29-9dd8-496f-aa53-0b6ec32830d7 · outbound

This paper cites Gpu-initiated networking for nccl,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpu-initiated networking for nccl,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.494336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.494336Z digest=sha256:a4c58507a284a77c98701ed09ccd8bc3850076f3fa9d0de478af15f0f6005a34

Observation 39a171ef-35a5-41c5-aca5-eb4f25696c3f · outbound

This paper cites NVIDIA, 2025.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.661046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.661046Z digest=sha256:79d4126e792976852e97c5685a50988d30a3b15a4194bb4d35babe0cc5a93fda

Observation eb7c6918-a749-4e52-9d5b-c23d4e978426 · outbound

This paper cites Fusing communication and compute with new device api and copy engine collectives in nvidia nccl 2.28,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Fusing communication and compute with new device api and copy engine collectives in nvidia nccl 2.28,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.784410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.784410Z digest=sha256:ec5d5c8748704fa2bfe31551434b4a26483cc7112d3493304d02dacdc6380587

Observation d9f3656d-35c1-48a9-bc7b-29758dbb8a80 · outbound

This paper cites PyTorch Foundation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch Foundation, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.857736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.857736Z digest=sha256:d3590cd28b16f7364af4d2e8f99af0c556eba14095ad83aba6fca8f3d64bd1ac

Observation 6553156a-7e56-4d0e-a740-6027b92cbe45 · outbound

This paper cites NHR@FAU, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NHR@FAU, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.919904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.919904Z digest=sha256:6518a09061d4e5e4115f17518c452120e3c580d7fa93fded4a9f814e48bf7894

Observation 472f42f1-784e-4951-96e8-a644b2f1a4ec · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:16.982682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:16.982682Z digest=sha256:7d8c90054cb30ca680d83f0ac4398125e8358afcae605b65df28bc5bdd20c906

Observation 1b5388cb-f193-45bd-85b0-93ad50f1a373 · outbound

This paper cites TensorFlow: Large-scale machine learning on heterogeneous systems,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives TensorFlow: Large-scale machine learning on heterogeneous systems,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.085329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.085329Z digest=sha256:82055d1e6d4a9936bc4ff717910d054f672815da013ff7f350ab5917968ceddc

Observation 88c2afe9-aa80-4fc4-a87d-596cccc7a504 · outbound

This paper cites NVIDIA Corporation,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.160785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.160785Z digest=sha256:3a182b898a2cf5359ffb6a91fcf674389380c313aef3ae04c727dcde6bd14edc

Observation b3a6a6a7-957e-44d1-ae52-321de4b59c6c · outbound

This paper cites NVIDIA Corporation, 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation, 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.331377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.331377Z digest=sha256:a4923dabc7b96fb8db6b03d83e75ca41cbfa63271e1cf816f613cd957b4687a9

Observation 8c77e832-ca44-4a66-9721-44529684ea50 · outbound

This paper cites Llm inference beyond a single node: From bottlenecks to mitigations with fast all-reduce communication,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Llm inference beyond a single node: From bottlenecks to mitigations with fast all-reduce communication,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.416705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.416705Z digest=sha256:b1214cc0a5cfbf32bfae74d04ffc9ec62ef6acf5288e9d34a3950d44117cbafd

Observation 5e9e99dc-cd28-492c-9afa-81c4ee1ec49f · outbound

This paper cites Msccl++: Rethinking gpu communication abstractions for ai inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Msccl++: Rethinking gpu communication abstractions for ai inference,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.512162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.512162Z digest=sha256:007cbbc0301ce485082616e6f5a65a061821421ce20ea925c7e1c5d82b0da202

Observation 555e25db-a80a-457d-a863-9cd2a6899968 · outbound

This paper cites Nvidia gtc 2025 - built for reasoning, vera rubin, kyber, cpo, dynamo inference, jensen math, feynman,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nvidia gtc 2025 - built for reasoning, vera rubin, kyber, cpo, dynamo inference, jensen math, feynman,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.600194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.600194Z digest=sha256:6c90f17e212391a693d89d58e1dc749da03a10fb06bd7f36fd05763bd48e4366

Observation 07f573ca-22b3-4bc6-8073-ee66872fefe8 · outbound

This paper cites Colossus: xai’s supercomputer for grok,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Colossus: xai’s supercomputer for grok,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.688668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.688668Z digest=sha256:92b1c9192ef3c823ad73495950ff45218fc9feb4b931ffeb2d5dfe688f9febc7

Observation 84ec75fb-c0dc-46ea-83d4-410030bb12c2 · outbound

This paper cites Recent improvement to open mpi allreduce and the impact to application performance,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Recent improvement to open mpi allreduce and the impact to application performance,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.756420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.756420Z digest=sha256:5a0736cdc0f102e78ef97933c6d26f6ce7d1d109ee1b12f584c4d4d9f89659de

Observation 749fc5d6-ab5c-4ce0-b39b-d74a13e3be6f · outbound

This paper cites Revisiting the time cost model of allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Revisiting the time cost model of allreduce,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.859961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.859961Z digest=sha256:c50628addf938dcece48870e6de819765e2c9726a1ec309428e9a22a01b1b7e8

Observation fafed51a-9d45-45a3-a82e-2b0eddf758a9 · outbound

This paper cites xccl: A survey of industry-led collective communication libraries for deep learning,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives xccl: A survey of industry-led collective communication libraries for deep learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.951226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.951226Z digest=sha256:c1310bbeb5e3cb52aba28cfdb527b2d2df76d32424a317c6e374708bc5659bae

Observation f7d1e527-d487-4679-9914-4503d0286d0f · outbound

This paper cites Hear: Homomorphically encrypted allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hear: Homomorphically encrypted allreduce,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.026695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.026695Z digest=sha256:d9e4d8962ca890c3776f729443617295a70af572930e056de972b9c5b4c744e7

Observation fac11bdf-0895-49b3-8972-47ca79be7a87 · outbound

This paper cites A survey of mpi usage in the us exascale computing project,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A survey of mpi usage in the us exascale computing project,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.218838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.218838Z digest=sha256:34adb8598640ceb0b11e720dc2ee14d508049491f80e9b929c1c3ee025e61eb5

Observation 7a162cce-669a-4284-ad8d-842c61632797 · outbound

This paper cites A large-scale study of mpi usage in open-source hpc applications,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A large-scale study of mpi usage in open-source hpc applications,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.463178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.463178Z digest=sha256:01884cad292946ec1dea349419fd9c6e6a8c09c7d2c970a9032d46343f32255d

Observation a0ae434e-702d-4a55-9279-4ed928885529 · outbound

This paper cites Short-circuiting rings for low-latency allreduce,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Short-circuiting rings for low-latency allreduce,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.590758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.590758Z digest=sha256:5c2d81b837b538f5eaff1244fd5a526de7e1e5a15fd22a5fdc7053f0c7b38ccb

Observation f41d774e-b592-4c46-a878-9c84b31b3ad8 · outbound

This paper cites Cuda c++ programming guide.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Cuda c++ programming guide

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.684480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.684480Z digest=sha256:9145a082f9663ffd603a0c71d0b618aa979b1cbb7caa238a2e749a096595b103

Observation 0252082f-c101-40bb-ae82-dc161755edab · outbound

This paper cites Parallel thread execution isa version 8.0.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Parallel thread execution isa version 8.0

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.772103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.772103Z digest=sha256:3355aad7a38cd9890aaa774b4ade560e0db09c99dd33a89c782d86c8bb61aca6

Observation afb4469f-2c07-4271-a037-cdab3a06269b · outbound

This paper cites vllm container (version 26.02-py3),.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm container (version 26.02-py3),

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.845547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.845547Z digest=sha256:92b04479ebb4e556631cdd2c9b7e88a7e9d7aad793781efbc8e06ee7a879551f

Observation 784214d7-9123-4747-8d89-936f6bdeeace · outbound

This paper cites Network-offloaded bandwidth-optimal broadcast and allgather for distributed ai,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Network-offloaded bandwidth-optimal broadcast and allgather for distributed ai,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.896433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.896433Z digest=sha256:7998d3a81fb85a1ef3cfd39508642c3ec8fe02307be98eb5045e2ea6135ed6ce

Observation 2f37d146-a83d-4999-ac3a-30f47a9a516a · outbound

This paper cites Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.083488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.083488Z digest=sha256:3c90445a59fd39cb8d13093c7f7321cc45da412d114614e537cd0fd23ed010cb

Observation 8701c828-c2be-4774-8eba-d71b78d59efd · outbound

This paper cites vllm – technology radar entry,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm – technology radar entry,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.181539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.181539Z digest=sha256:ac28fb6f2326f1153e4c7bb327ed6b82af8cc08935611695e8d372fcba34b820

Observation 323375ca-c4ab-487e-ba67-eee2cd45c0fe · outbound

This paper cites The huge potential implications of long-context inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The huge potential implications of long-context inference,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.273592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.273592Z digest=sha256:ef2297c16e63941e1c61f61b436898de819a7c2aaac9c9337880122b162426dd

Observation 0ffa4cbf-30c7-47a0-80ab-8881c503fa2b · outbound

This paper cites Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.006976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.006976Z digest=sha256:1acd5accefe8e954f6ecec23c2b5162a3b00487450a02f7e488ef0f397245e29

Observation c1127a6f-4641-49b7-aa98-7e64e2d365ea · outbound

This paper cites Chain of agents: Large language models collaborating on long-context tasks,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Chain of agents: Large language models collaborating on long-context tasks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.440969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.440969Z digest=sha256:1405d8e4429ad189f95ea95746fb7f644a4ef87d3e2ad7de5170eb17c0d0b612

Observation 9748d69e-2b66-4ed0-b867-61dc58a6b85b · outbound

This paper cites Breaking the boundaries of long- context llm inference: Adaptive kv management on a single commodity gpu,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Breaking the boundaries of long- context llm inference: Adaptive kv management on a single commodity gpu,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.528704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.528704Z digest=sha256:63334e2dd9f99719a41cdba3e79f065724e39b7b75676e8d36be0ec4ffac73bf

Observation 1d5f83da-1bb3-438b-abbe-41a0e30dcfb5 · outbound

This paper cites Prefill-decode disaggregation.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Prefill-decode disaggregation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.629418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.629418Z digest=sha256:0af60f616f5887a2eebedd9f58f1ceed95fe5bbbcd73114a6cfdc9e7cb66d5e1

Observation 8aafc676-e597-4ac0-ac20-bcf8c7c6341d · outbound

This paper cites Is long context all you need? leveraging llm’s extended context for nl2sql,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Is long context all you need? leveraging llm’s extended context for nl2sql,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.382129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.382129Z digest=sha256:666ab460f8f7ae0bed1c730c141bf6004a6030e512092b0021be6ebda852af1c

Observation 7e090cab-d299-4f45-92fe-6193df7732e8 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.766032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.766032Z digest=sha256:70c78f1e4393cc8dff6943d9cb93330672d4a81f6701ab9a9deb2c0eae775d79

Observation 874e5964-549e-489c-a52b-7356e2e0dc95 · outbound

This paper cites Accessed: March 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.916482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.916482Z digest=sha256:016195792e8a20e77cfd9ebc478718319e8bc23918902023fc4a59b278d21466

Observation 4f1b716c-61c9-489b-b789-8d1fbb58b7ae · outbound

This paper cites Amrex and pyamrex: Looking beyond the exas- cale computing project,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Amrex and pyamrex: Looking beyond the exas- cale computing project,

Reference 54

Resolution
verified exact
doi, observed 2026-08-01T21:28:32.655939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-01T21:25:19.998763Z digest=sha256:07935476ff6ef6d363079def03fef32ca77e6b8b2e4484e6ca2791caddd0d98a

Observation 91ae64af-c590-47a8-b86f-c6d071eaf952 · outbound

This paper cites Slo-aware compute resource allocation for prefill-decode disaggregated llm inference,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Slo-aware compute resource allocation for prefill-decode disaggregated llm inference,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:19.670034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.670034Z digest=sha256:c52e2874706387ee6e12dc8078734f07c8cfbf73eceb88437fa0582bd0b7154c

Observation 2a329e9e-4041-45be-90bd-4bc80afa968c · outbound

This paper cites Qmcpack: an open source ab initio quantum monte carlo package for the electronic structure of atoms, molecules and solids,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack: an open source ab initio quantum monte carlo package for the electronic structure of atoms, molecules and solids,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.252829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.252829Z digest=sha256:6a10052c49441aeac5075c93c2eb7788e6884bd5b932e97a6ee5aa07017368fe

Observation d3c997e9-79f3-4624-b9f7-6a224d926070 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 57

Resolution
parse uncertain
no resolver link, observed 2026-08-01T21:25:19.821885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:19.821885Z digest=sha256:78e8e6f43bb8cde5015fdaa60e514301292b006b6ef41b74750e113b54f2e051

Observation 6005d775-7c55-4096-9985-7542cd1c6bc9 · outbound

This paper cites NVIDIA HPCG Benchmark,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA HPCG Benchmark,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.385515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.385515Z digest=sha256:28fa6469715b589644b3406eb53aa25c95e1e47dc1baec01dff1594ea196edc4

Observation 7a2f491e-e59a-479f-8ee1-c644e10c9fc3 · outbound

This paper cites Accessed: March 2026.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.464831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.464831Z digest=sha256:18025c90f299865bce5c3cb8b44cf993db38c97019cb560d752e119adf372355

Observation be738fcb-82d0-463c-96d2-c14d401588b6 · outbound

This paper cites AMReX issue #4821: Device-initiated collectives in NCCL/RCCL.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMReX issue #4821: Device-initiated collectives in NCCL/RCCL

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.077961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.077961Z digest=sha256:4cd65f011a7e63e94884e5f53ac069d7df201d0a7adef1a0f85ba069f8d5c0b7

Observation 4702eab3-4a03-4345-9f35-ca16d8197c20 · outbound

This paper cites New research infrastructure: ’alps’ supercomputer inaugurated,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives New research infrastructure: ’alps’ supercomputer inaugurated,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.626672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.626672Z digest=sha256:df034a2a05cffc381cda0c94d8a821ccdccdb69708d69db009fa2f1a09f7d63b

Observation 1f8ab145-110b-45b5-b5cf-879508324856 · outbound

This paper cites Flashinfer: Efficient and customizable attention engine for llm inference serving,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Flashinfer: Efficient and customizable attention engine for llm inference serving,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.731027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.731027Z digest=sha256:02add0087a047b784016510a16f2c233260d09b0d27531fe4a1fa04f89ea397c

Observation cd35d561-5800-452b-ba99-0071dc64e305 · outbound

This paper cites Qmcpack issue #4654,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack issue #4654,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.332056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.332056Z digest=sha256:82bc80b75b3800b316fe82f94299b6c668e6707ef68fa98f0666e93f2b8dfa6a

Observation 90e64a8e-5e91-4298-a3e9-ab2d51c5fe27 · outbound

This paper cites Nccl ep: Towards a unified expert parallel communication api for nccl,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nccl ep: Towards a unified expert parallel communication api for nccl,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.853881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.853881Z digest=sha256:200d9834751af3899c5a2b42d17418bc4945c7bf928f4a94d8b7efbf88c11a62

Observation 02d42690-5e33-42c5-80e1-106ff10d1444 · outbound

This paper cites Enhancing dis- tributed inference performance with the nvidia inference transfer library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enhancing dis- tributed inference performance with the nvidia inference transfer library,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.934999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.934999Z digest=sha256:44351ebef6f6dee565559f29b8777a1fe8c22b408283445c3146060f056ceae4

Observation 29b8628c-a227-432b-84c2-a7050299050e · outbound

This paper cites The elpa library: scal- able parallel eigenvalue solutions for electronic structure theory and computational science,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The elpa library: scal- able parallel eigenvalue solutions for electronic structure theory and computational science,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.549073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.549073Z digest=sha256:c212afa3265718676ab8b9a06d9217cb1820be356fa8f6fafea5b10c5bfeb452

Observation 0836cd97-20e2-42c9-abfe-54893ca7b307 · outbound

This paper cites Deepep: an efficient expert-parallel communication library,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepep: an efficient expert-parallel communication library,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.804322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.804322Z digest=sha256:e91c5e547322939d6697b1c5850d0911a3a21d3583c6a2ae2abcbda8186a5f0c

Observation aa173ed3-ed6f-40fd-bc13-383abd1e9ad4 · outbound

This paper cites Gpt-5.4 chatbot,.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpt-5.4 chatbot,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.987873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.987873Z digest=sha256:01e9c1146959de75a64c9d5d3430fa499c6425cd54ca35143721b55934a1989a

Observation ae2cf742-ceec-41a6-b499-f3cdccdb1871 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.335994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.335994Z digest=sha256:12c65d9c1b83a93babf2acad0b8a7349022f7fb9b2cd32fdb0ca935812e0d285

Observation 6947ed3d-468a-4bd4-b977-0ba2f5b4af63 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:18.096674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:18.096674Z digest=sha256:5317086b812b3263a6917d20465d3c100862b86ca0b56add327515e2be4d5c52

Observation f0c5898d-7da4-438e-8d51-52d3a6726d18 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:17.236864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:17.236864Z digest=sha256:78973ce38a452c9fa5306d18dce1eb650e5dbee594f79d4f00b172031afef5c8

Observation 78451cc0-945c-465e-902a-fe41b4d093f9 · outbound

This paper cites an unresolved cited work.

Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T21:25:20.176773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:25:20.176773Z digest=sha256:0d22a6ff5b376ec565b24b09b1ff9087a63a25b52a1a0e7a1fd0a49777b146c2

Pith citing papers

No inbound Pith citation observations are available.