Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:25:20.987873Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2607.16100.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:25:20.987873Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4a9ad259-59ce-4d54-81d1-e51c4da5b342 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3 technical report,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3788ff-7317-43c7-938d-9aa9caffd13f · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepseek-v3/r1 671b deployment guide: Gpu require- ments,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e35e9c3-5090-4f82-99a9-a745e6e9b133 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Efficient memory management for large language model serving with pagedattention,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afbebda9-16f1-445a-8503-05689076a4b6 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives State of ai: An empirical 100 trillion token study with openrouter,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f2e045-189d-482d-a104-95b85da321dd · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Sglang: efficient execution of structured language model programs,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3903cc2-f50f-4753-8ce8-39a453417ea1 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Tensorrt-llm: A library for optimizing large language model inference,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1038d302-deab-4ff1-8a98-511433a776a5 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives CoreWeave Pricing: Instance Pricing,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09056ba9-bbfb-4762-973b-7222605e7748 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Scaling-up pytorch inference: Serving billions of daily nlp inferences with onnx runtime,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e789aa8c-d279-4c85-84b0-cf6a4576cf83 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Understanding data movement in tightly coupled heterogeneous systems: A case study with the grace hopper superchip,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba6e4d6b-f423-46cb-a0eb-4d6c06e43fee · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Demystifying NCCL: An In-Depth Analysis of GPU Communication Protocols and Algorithms ,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5220af4-d77e-4356-afdb-85ce704bc8c2 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVSHMEM communication library,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7731326-971c-4750-b79f-2deac8501581 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives ROCm Communication Collectives Library (RCCL) Documentation,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af79158d-7236-4e53-a8a2-d0bf6213a478 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMD ROCm Documentation, 2026
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d36d779-647d-4b94-b44a-7a884e24b51c · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: February 2026
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4b5d931-844d-4bb2-b6c4-bc1ca39339b5 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ca6079-e9d0-4b9a-9fd3-08e37a72c1ad · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives An introduction to cuda-aware mpi,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae8a94c5-8d78-4b20-abc1-adb893da5acd · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enabling fast inference and resilient training with nccl 2.27,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34d0e29-9dd8-496f-aa53-0b6ec32830d7 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpu-initiated networking for nccl,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39a171ef-35a5-41c5-aca5-eb4f25696c3f · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7c6918-a749-4e52-9d5b-c23d4e978426 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Fusing communication and compute with new device api and copy engine collectives in nvidia nccl 2.28,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f3656d-35c1-48a9-bc7b-29758dbb8a80 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch Foundation, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6553156a-7e56-4d0e-a740-6027b92cbe45 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NHR@FAU, 2026
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 472f42f1-784e-4951-96e8-a644b2f1a4ec · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives PyTorch: An Imperative Style, High-Performance Deep Learning Library
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b5388cb-f193-45bd-85b0-93ad50f1a373 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives TensorFlow: Large-scale machine learning on heterogeneous systems,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c2afe9-aa80-4fc4-a87d-596cccc7a504 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a6a6a7-957e-44d1-ae52-321de4b59c6c · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA Corporation, 2026
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c77e832-ca44-4a66-9721-44529684ea50 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Llm inference beyond a single node: From bottlenecks to mitigations with fast all-reduce communication,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9e99dc-cd28-492c-9afa-81c4ee1ec49f · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Msccl++: Rethinking gpu communication abstractions for ai inference,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555e25db-a80a-457d-a863-9cd2a6899968 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nvidia gtc 2025 - built for reasoning, vera rubin, kyber, cpo, dynamo inference, jensen math, feynman,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f573ca-22b3-4bc6-8073-ee66872fefe8 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Colossus: xai’s supercomputer for grok,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ec75fb-c0dc-46ea-83d4-410030bb12c2 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Recent improvement to open mpi allreduce and the impact to application performance,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 749fc5d6-ab5c-4ce0-b39b-d74a13e3be6f · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Revisiting the time cost model of allreduce,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fafed51a-9d45-45a3-a82e-2b0eddf758a9 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives xccl: A survey of industry-led collective communication libraries for deep learning,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d1e527-d487-4679-9914-4503d0286d0f · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hear: Homomorphically encrypted allreduce,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac11bdf-0895-49b3-8972-47ca79be7a87 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A survey of mpi usage in the us exascale computing project,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a162cce-669a-4284-ad8d-842c61632797 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives A large-scale study of mpi usage in open-source hpc applications,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ae434e-702d-4a55-9279-4ed928885529 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Short-circuiting rings for low-latency allreduce,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41d774e-b592-4c46-a878-9c84b31b3ad8 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Cuda c++ programming guide
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0252082f-c101-40bb-ae82-dc161755edab · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Parallel thread execution isa version 8.0
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb4469f-2c07-4271-a037-cdab3a06269b · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm container (version 26.02-py3),
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784214d7-9123-4747-8d89-936f6bdeeace · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Network-offloaded bandwidth-optimal broadcast and allgather for distributed ai,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f37d146-a83d-4999-ac3a-30f47a9a516a · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Comparative analysis of large language model inference serving systems: A performance study of vllm and huggingface tgi,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8701c828-c2be-4774-8eba-d71b78d59efd · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives vllm – technology radar entry,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 323375ca-c4ab-487e-ba67-eee2cd45c0fe · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The huge potential implications of long-context inference,
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ffa4cbf-30c7-47a0-80ab-8881c503fa2b · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Hilfer fractional advection-diffusion equations with power-law initial condition; a Numerical study using variational iteration method
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1127a6f-4641-49b7-aa98-7e64e2d365ea · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Chain of agents: Large language models collaborating on long-context tasks,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9748d69e-2b66-4ed0-b867-61dc58a6b85b · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Breaking the boundaries of long- context llm inference: Adaptive kv management on a single commodity gpu,
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d5f83da-1bb3-438b-abbe-41a0e30dcfb5 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Prefill-decode disaggregation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aafc676-e597-4ac0-ac20-bcf8c7c6341d · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Is long context all you need? leveraging llm’s extended context for nl2sql,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e090cab-d299-4f45-92fe-6193df7732e8 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874e5964-549e-489c-a52b-7356e2e0dc95 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f1b716c-61c9-489b-b789-8d1fbb58b7ae · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Amrex and pyamrex: Looking beyond the exas- cale computing project,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91ae64af-c590-47a8-b86f-c6d071eaf952 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Slo-aware compute resource allocation for prefill-decode disaggregated llm inference,
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a329e9e-4041-45be-90bd-4bc80afa968c · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack: an open source ab initio quantum monte carlo package for the electronic structure of atoms, molecules and solids,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3c997e9-79f3-4624-b9f7-6a224d926070 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6005d775-7c55-4096-9985-7542cd1c6bc9 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives NVIDIA HPCG Benchmark,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a2f491e-e59a-479f-8ee1-c644e10c9fc3 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Accessed: March 2026
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be738fcb-82d0-463c-96d2-c14d401588b6 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives AMReX issue #4821: Device-initiated collectives in NCCL/RCCL
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4702eab3-4a03-4345-9f35-ca16d8197c20 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives New research infrastructure: ’alps’ supercomputer inaugurated,
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f8ab145-110b-45b5-b5cf-879508324856 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Flashinfer: Efficient and customizable attention engine for llm inference serving,
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd35d561-5800-452b-ba99-0071dc64e305 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Qmcpack issue #4654,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e64a8e-5e91-4298-a3e9-ab2d51c5fe27 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Nccl ep: Towards a unified expert parallel communication api for nccl,
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d42690-5e33-42c5-80e1-106ff10d1444 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Enhancing dis- tributed inference performance with the nvidia inference transfer library,
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b8628c-a227-432b-84c2-a7050299050e · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives The elpa library: scal- able parallel eigenvalue solutions for electronic structure theory and computational science,
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0836cd97-20e2-42c9-abfe-54893ca7b307 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Deepep: an efficient expert-parallel communication library,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa173ed3-ed6f-40fd-bc13-383abd1e9ad4 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Gpt-5.4 chatbot,
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2cf742-ceec-41a6-b499-f3cdccdb1871 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6947ed3d-468a-4bd4-b977-0ba2f5b4af63 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c5898d-7da4-438e-8d51-52d3a6726d18 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78451cc0-945c-465e-902a-fe41b4d093f9 · outbound
Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.