REVIEW 4 major objections 5 minor 115 references
NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Non-uniform network access inside large multi-die GPUs is a first-order cost in GPU-to-GPU communication, and routing plus placement that respect it make latency-sensitive collectives up to 1.8x faster.
desk verdict A useful new framing of intra-socket I/O locality for scale-up GPUs, with real motivating measurements, but headline speedup claims are internally inconsistent and the evaluation is simulation-only. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the NAR domain: a statically partitioned slice of physical address space that constrains a CU's off-chip flits to a subset of I/O ports physically near that CU, while still load-balancing across ports within the subset. NAP's algorithm takes a collective plan (threadblock access patterns over fixed-size chunks), first groups threadblocks and chunks into logical groups, then greedily maps each group to a NAR domain and, within it, to the CU and HBM stack closest to the I/O ports subject to resource budgets such as outstanding-request capacity and bandwidth-delay product. The paper also exposes NAR granularity, the number of domains, as the tuning knob that trades low latency against port load balancing. These mechanisms are implemented with existing threadblock-to-CU affinity masks and page-to-stack placement, plus a modification to address-hashing logic, so they are backward-compatible and disable-able.
What would settle it
Run a 1 MB All-Gather on a next-generation multi-die GPU with a rail-optimized single-level Clos, and compare round-robin placement with full-port hashing against the paper's NAP+NAR scheme. If the measured speedup comes nowhere near 1.8x -- for example, because scale-up link latency or switch queueing so dominates the remote write that the choice of CU, HBM stack, and I/O port no longer changes the total time -- the central claim fails. A simpler check is to measure remote-store latency from the fastest versus slowest compute unit to the same remote address: the paper predicts that gap grows with socket size.
Extended reading notes
Core claim
The paper's central claim is that the remote-store path between GPUs has two comparable legs: the on-chip trip from a compute unit to an edge I/O port and through the receiving socket to its memory stack, and the scale-up link between sockets; for the modeled next-generation system the on-chip leg reaches 0.9 $\mu$s while the scale-up leg is 1 $\mu$s. Because these legs are comparable, where a threadblock sits, where its buffers live, and which I/O port hashes the traffic determine the latency of a small collective. The paper argues that NAR plus NAP addresses all three choices together: NAR routes flits to a physically close subset of ports, and NAP places communicating threadblocks on nearby CUs and their chunks on nearby HBM stacks, with a greedy two-phase algorithm that resolves resource contention. In simulation, the combined scheme reduces All-Gather, All-Reduce, and All-to-All execution time by up to 80% (up to 1.8x speedup) for small collectives, with a 1.32x geomean communication speedup in prefill and 1.56x in decode across the LLM workloads.
Load-bearing premise
The load-bearing premise is that moving data across the GPU's interior to the edge I/O port (up to 0.9 $\mu$s) costs about as much time as sending it between GPUs over the scale-up link (1 $\mu$s), so the exact placement of threads, data, and ports inside the socket is a first-order factor; if inter-GPU link latency grows faster than on-chip latency, or if compute hides the on-chip transfer, the speedups shrink.
Editorial extensions
If this is right
- Small, latency-bound collectives (up to roughly 1 MB) benefit most, with 1.5-1.9x speedups that decay to about 1.1x by 64 MB as bandwidth saturation hides the on-chip spatial effect.
- The scheme cuts decode time-per-output-token by 7% on average and 28% at best, and prefill time-to-first-token by 6% on average and 11% at best, across the 12 evaluated LLM architectures and pod sizes from 2 to 64 GPUs.
- Placement without routing is not enough: NAP alone can slow down 10 MB collectives because hashed traffic traverses the whole socket and congests the on-chip network; routing and placement must be designed jointly.
- The best NAR granularity depends on threadblock count, so the allocation algorithm should choose the number of domains per workload rather than fixing it in hardware.
- Concurrent compute kernels degrade the NUNA-aware collective less than the baseline, with degradation appearing only at roughly twice the GEMM size, because localized I/O traffic avoids mid-die congestion.
Reading between the lines
- The 1.8x figure is tied to the modeled parameter regime: on sockets where the scale-up link dominates the on-chip leg the same policies would yield less, while on still-larger dies the gains could exceed 1.8x.
- The same NAR/NAP machinery generalizes to any latency-sensitive off-chip traffic with a known access plan, such as remote atomics or sparse attention token routing, not just the three collectives evaluated.
- Because the paper serializes compute and communication, its end-to-end numbers are optimistic about exposure; if future systems overlap compute and collectives well, per-collective speedups would remain but the 7% TPOT gain would shrink.
- A natural testable extension is to derive the topology input to NAP from measured per-CU remote-store latencies rather than geometric layout, so the method stays calibrated if real meshes or packaging differ from the modeled 24x6 mesh.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces the term non-uniform network access (NUNA) to describe how intra-socket wire distance between compute units, HBM stacks, and I/O ports makes inter-GPU scale-up communication latency depend on spatial placement. It proposes two optimizations: NUNA-aware routing (NAR), which restricts latency-sensitive flits to I/O ports physically closer to the requesting CU via statically partitioned address ranges, and NUNA-aware placement (NAP), which places collective threadblocks and memory pages near those I/O ports. The paper profiles three real AMD GPU systems to show that remote-access latency varies with CU and address location, then evaluates NAR and NAP in the ASTRA-sim 3.0 simulator. The headline claims are up to 1.8x collective speedups for NAP+NAR over the locality-unaware baseline and 7% mean (28% max) end-to-end time-per-output-token improvement for LLM inference.
Significance. If the claimed speedups are validated, the paper makes a timely contribution: it identifies a new locality dimension for scale-up GPU systems and provides a coherent, backward-compatible software/hardware mechanism for exploiting it. Strengths include the real-system profiling of MI210, MI355X, and MI300X, which documents that remote-access latency variation grows with socket size; the isolation study showing that NAR is necessary for NAP to help; and the sensitivity studies for collective size and NAR granularity. The paper does not, however, ship code or validation data, and its central quantitative claims depend on a self-built simulator and on two strong assumptions: the NoC/scale-up latency ratio and a serial compute-communication model. Those assumptions are load-bearing for the 1.8x and 7% headline numbers, so the result is conditional until the assumptions are tested and the internal speedup accounting is reconciled.
major comments (4)
- [§1 and Abstract] The manuscript states two incompatible headline speedups. The abstract claims up to 1.8x faster collectives for NAP+NAR, while the introduction says 'NAP and NAR applied together can reduce collective execution time by up to 80%' (a 5x speedup). Both cannot describe the same comparison against the same baseline; the 80% reduction is hard to reconcile with the 1.8x upper bound reported later. This needs to be corrected and the exact comparison (which configuration, collective, message size, and pod size) stated for each number.
- [§3.1.3 and §1] The motivating 1.8x 'uncontended GPU-to-GPU transfer slowdown' is not supported by the paper's own latency numbers. Section 3.1.3 gives a NoC transfer of up to 0.9 us and a scale-up latency of 1 us; with a 3x NoC spread (0.3 to 0.9 us) plus a fixed 1 us scale-up term, the worst/best total transfer ratio is (1.9)/(1.3) = 1.46x, not 1.8x. Because NAP and NAR act only on the intra-socket segment, any fixed scale-up latency dilutes the benefit, and the claimed 1.8x collective speedup therefore depends on the assumed latency ratio. The paper provides no sensitivity analysis over this ratio. I ask the authors to (i) show exactly how the 1.8x transfer-slowdown and 1.8x collective-speedup numbers are derived, and (ii) add a sensitivity study varying the NoC/scale-up latency ratio (e.g., 0.3x to 3x of the assumed 0.9 us) and report collective speedup as a function of that ratio.
- [§5.2] The end-to-end inference results are computed under a serial compute/communication model, as stated in Section 5.2: 'we model compute and communication serially since intra-batch overlap is ongoing research.' This assumption removes any overlap between collective communication and compute, which exposes the full communication time and therefore inflates the measured benefit of faster collectives. Since the paper itself cites prior overlap work (e.g., T3 [80] and Splitwise [78]), a serial model is a worst-case assumption, not a neutral one. The 7% mean TPOT claim is thus conditional on zero overlap. The authors should quantify how the end-to-end speedup changes under partial or full overlap (e.g., by applying an overlap factor to the collective time in their analytical model).
- [§5.1 and §3.2] The central collective-speedup results come from ASTRA-sim 3.0, a simulator co-authored by this team (reference [98]), and the paper does not ship the simulator configuration, trace-generation tool, or validation data. Section 3.2's real-system profiling validates the existence of latency variation on current hardware, but it does not validate that the simulated 24x6 mesh with the chosen hop latencies reproduces those measurements. Without a calibration/validation experiment comparing simulated remote-store latency distributions against the measured MI300X values, the quantitative claims (1.8x for NAP+NAR, up to 28% TPOT improvement) cannot be independently checked. I would like the paper to include a validation section, or at minimum release the exact configuration and a comparison of simulated versus measured latency distributions.
minor comments (5)
- [§1, §4.1.2, §6.3] There are several typos: 'Specifcially' in the introduction, 'to to limit' in Section 4.1.2, and 'traffics' in Section 6.3.
- [§2.2, Figure 2] Figure 2 reports collective sizes derived from an in-house trace-generation tool, but the methodology is deferred to Section 5.2 and no validation is given for the trace shapes or collective sizes. A short validation sentence (e.g., comparison with known model configurations) would help the reader trust the prefill/decode CDFs.
- [§5.3, Table 2] The simulated system table lists '144 I/O ports (72 per die)' for a 4-compute-die socket, while Section 4.1.1 and Figure 5 describe a socket with four I/O ports. The figure is clearly illustrative, but the relationship between the illustrative four-port model and the simulated 144-port model should be stated explicitly to avoid confusion.
- [§6.2] The description of the UA latency histogram as 'the Gaussian distribution of individual path latencies in a mesh' is informal; the reported 920 ns average and 1,320 ns maximum are more informative. Reporting percentiles (e.g., p50/p99/max) directly in the text would make the comparison clearer.
- [§6.7] The phrase 'exposed communication ratio is approximately 4:1 for prefill and 5:1 for decode' is ambiguous: it is not clear whether this is compute-to-communication or communication-to-compute, and how the ratio is computed. Clarifying the definition would help the Amdahl's law argument.
Circularity Check
No significant circularity: the speedup claims are simulated outcomes from stated models, not predictions that reduce to their inputs by construction.
full rationale
The paper's central claims are empirical and simulation results conditional on stated parameters, not derivations that reduce to their inputs. Section 3.2 profiles real AMD Instinct GPUs (MI210, MI355X, MI300X) and measures latency variation, providing an independent anchor for the NUNA phenomenon. The 1.5x/1.8x collective speedups and 7%/28% TPOT gains are outputs of ASTRA-sim 3.0 simulations using explicit NoC mesh, hop latencies, I/O port counts, and scale-up switch parameters (Table 2, Section 5.3), rather than relabeled fits. NAR and NAP are evaluated against a round-robin/hash baseline; their benefit is computed by the simulator's cycle-level model, not assumed by construction. The use of ASTRA-sim 3.0 [98] is a self-citation, but it is used as a simulation tool with independently stated inputs, and the paper also presents real-system profiling; no load-bearing claim is justified solely by an unverified self-cited theorem. The paper's '1.8x/1.45x transfer slowdowns' in Section 1 are not arithmetically consistent with the 0.9 us NoC / 1 us scale-up numbers in Section 3.1.3, and the serial compute/communication model in Section 5.2 may inflate end-to-end gains, but these are correctness and robustness concerns, not circularity. No equation in the paper equates an output to an input by construction.
Assumptions & free parameters
free parameters (5)
- NAR granularity (narmax) =
2-12 domains in sensitivity experiments; compiler-selected in practice
- CU outstanding request budget (Rmax CU) =
700 requests per CU
- HBM stack in-flight budget (bandwidth-delay product) =
not specified numerically
- Number of threadblocks for collectives ('best-performing') =
not reported
- NAR usage threshold (collective size) =
unstated; 'statically determined' per architecture
assumptions (5)
- domain assumption Intra-socket NoC transfer delay will be comparable to scale-up link latency (around 1 us) on next-generation GPUs, making intra-socket distance a first-order factor for inter-GPU communication.
- domain assumption The simulated 24x6 mesh NoC with X-Y routing, 4 compute dies, and single-level Clos scale-up is representative of future multi-die GPU systems.
- domain assumption Collectives up to roughly 100 MB are latency-bound rather than bandwidth-bound, so point-to-point latency reductions translate into collective speedups.
- domain assumption Compute and communication do not overlap in inference execution, so reducing communication time directly reduces end-to-end time.
- domain assumption The real-system latency variation in Figure 4 is caused by physical distance rather than by contention or hashing artifacts.
Cite this review
Pith. "Pith review of NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems." pith.science (2026). https://pith.science/paper/XXVHHPR7
@misc{pith2026260800867,
author = {Pith},
title = {Pith review of: NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems},
year = {2026},
howpublished = {\url{https://pith.science/paper/XXVHHPR7}},
note = {Machine review of arXiv:2608.00867}
}
read the original abstract
Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly. While previous research has optimized for compute and memory locality within a socket, the spatial impact on inter-GPU communication has not been well-studied. We introduce the term non-uniform network access (NUNA) to describe this emerging optimization dimension in multi-GPU systems. We specifically focus on latency-sensitive collective communication, common in machine learning inference. First, we highlight the need for NUNA-aware routing (NAR), choosing optimized, spatially-aware inter-GPU paths in large scale-up network topologies. Second, we introduce NUNA-aware placement (NAP), placing threadblocks and data near I/O to optimize the inter-GPU traffic. We demonstrate that the NAP optimizations alone offer up to 1.5x collective speedups over a locality-unaware baseline. Combining NAP with NAR yields up to 1.8x faster collectives over the locality-unaware baseline. This leads to 7% mean (28% max) time per output token speedup in machine learning inference.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[80]
Suchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena, and Matthew D. Sinclair. 2024. T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives. InProceedings of the 29th ACM International Confer- ence on Architectural Support for Programming Languages and Operating Systems, Volume 2(La Jolla, CA, USA)(ASPLOS ’24). Asso...
arXiv 2024
-
[78]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient generative llm inference using phase splitting. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 118–132
2024
-
[98]
William Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, and Bradford M. Beckmann. 2026. ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fi...
work page Pith review arXiv 2026
-
[1]
[n. d.]. GitHub - ROCm/rccl: ROCm Communication Collectives Library (RCCL) — github.com. https://github.com/ROCm/rccl. [Accessed 31-07-2025]
2025
-
[2]
[n. d.]. NVLink & NVSwitch for Advanced Multi-GPU Communication — nvidia.com. https://www.nvidia.com/en-us/data-center/nvlink/. [Accessed 31-07-2025]
2025
-
[3]
[n. d.]. SPCL - ATLAHS — spcl.inf.ethz.ch. https://spcl.inf.ethz.ch/Research/ Scalable_Networking/ATLAHS/. [Accessed 10-06-2026]
2026
-
[4]
NVIDIA GB200 Interconnect Architecture Analysis: NVLink, InfiniBand, and Future Trends - NADDOD Blog — naddod.com
2024. NVIDIA GB200 Interconnect Architecture Analysis: NVLink, InfiniBand, and Future Trends - NADDOD Blog — naddod.com. https://www.naddod.com/blog/nvidia-gb200-interconnect-architecture- analysis-nvlink-infiniband-and-future-trends. [Accessed 31-07-2025]
2024
-
[5]
NVIDIA Collective Communications Library (NCCL)
2025. NVIDIA Collective Communications Library (NCCL). https://developer. nvidia.com/nccl. [Accessed 31-07-2025]
2025
Show all 115 references
-
[6]
Advanced Micro Devices, Inc. [n. d.].AMD Instinct TM MI250 microarchitecture — ROCm Documentation. Online documentation page
-
[7]
2021.Introducing AMD CDNA TM 2 Architecture
Advanced Micro Devices, Inc. 2021.Introducing AMD CDNA TM 2 Architecture. White Paper. AMD. https://www.amd.com/content/dam/amd/en/documents/ instinct-business-docs/white-papers/amd-cdna2-white-paper.pdf Copyright notice shows 2021; document includes performance notes as of Ja...
2021
-
[8]
2025.HIP documentation; HIP 6.4.43484 Documen- tation
Advanced Micro Devices Inc. 2025.HIP documentation; HIP 6.4.43484 Documen- tation. [Accessed 01-08-2025]
2025
-
[9]
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementat...
2024
-
[10]
AMD. [n. d.]. AMD InstinctTM MI350X GPUs. https://www.amd.com/en/ products/accelerators/instinct/mi350/mi350x.html. [Accessed 31-07-2025]
2025
-
[11]
AMD. 2025. AMD Advancing AI 2025. https://www.amd.com/en/corporate/ events/advancing-ai.html. [Accessed 31-07-2025]
2025
-
[12]
Michael Andersch, Greg Palmer, Ronny Krashinsky, Nick Stam, Vishal Mehta, Gonzalo Brito, and Sridhar Ramaswamy. 2022. NVIDIA Hopper Architecture In-Depth. https://developer.nvidia.com/blog/nvidia-hopper-architecture-in- depth/. [Accessed 31-07-2025]
2022
-
[13]
Akhil Arunkumar, Evgeny Bolotin, Benjamin Cho, Ugljesa Milic, Eiman Ebrahimi, Oreste Villa, Aamer Jaleel, Carole-Jean Wu, and David Nellans. 2017. MCM-GPU: Multi-chip-module GPUs for continued performance scalability. ACM SIGARCH Computer Architecture News45, 2 (2017), 320–332
2017
-
[14]
Mohammad Banikazemi, Vijay Moorthy, and Dhabaleswar K Panda. 1998. Effi- cient collective communication on heterogeneous networks of workstations. InProceedings. 1998 International Conference on Parallel Processing (Cat. No. 98EX205). IEEE, 460–467
1998
-
[15]
Jeff Barr. 2019. Amazon ec2 update–inf1 instances with AWS inferentia chips for high performance cost-effective inferencing.A WS News Blog(2019)
2019
-
[16]
Bradford M Beckmann and David A Wood. 2004. Managing wire delay in large chip-multiprocessor caches. In37th International Symposium on Microarchitec- ture (MICRO-37’04). IEEE, 319–330
2004
-
[17]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901
2020
-
[18]
Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B Jablin, Nacho Navarro, and Wen-mei W Hwu. 2015. Automatic parallelization of kernels in shared- memory multi-GPU nodes. InProceedings of the 29th ACM on International Conference on Supercomputing. 3–13
2015
-
[19]
Li-Jhan Chen, Hsiang-Yun Cheng, Po-Han Wang, and Chia-Lin Yang. 2017. Improving GPGPU Performance via Cache Locality Aware Thread Block Sched- uling.IEEE Computer Architecture Letters16, 2 (2017), 127–131. https: //doi.org/10.1109/LCA.2017.2693371
2017
-
[20]
Jaehong Cho, Hyunmin Choi, and Jongse Park. 2025. LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure .IEEE Computer Architecture Letters24, 02 (July 2025), 361–364. https://doi.org/10.1109/LCA.2025.3628325
2025
-
[21]
Jack Choquette. 2022. Nvidia hopper gpu: Scaling performance. In2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1–46
2022
-
[22]
Charles Clos. 1953. A study of non-blocking switching networks.Bell System Technical Journal32, 2 (1953), 406–424
1953
-
[23]
Patrick H Coppock, Brian Zhang, Eliot H Solomon, Vasilis Kypriotis, Leon Yang, Bikash Sharma, Dan Schatzberg, Todd C Mowry, and Dimitrios Skarlatos
-
[24]
David Correa. 2025. Artificial Intelligence Chip Market Expected to Reach $460.9 Billion by 2034. https://www.einpresswire.com/article/866225614/artificial- intelligence-chip-market-expected-to-reach-460-9-billion-by-2034. [Accessed Nov. 11, 2025]
2025
-
[25]
Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, and Yifan Xiong. 2023. Mscclang: Microsoft collective communication language. InPro- ceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volu...
2023
-
[26]
2004.Principles and practices of interconnection networks
William James Dally and Brian Patrick Towles. 2004.Principles and practices of interconnection networks. Elsevier
2004
-
[27]
Preyesh Dalmia, Rajesh Shashi Kumar, and Matthew D Sinclair. 2024. CPElide: Efficient multi-chiplet GPU implicit synchronization. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 700–717
2024
-
[28]
2026.DeepSeek-V4: Towards highly efficient million-token context intelligence
DeepSeek-AI. 2026.DeepSeek-V4: Towards highly efficient million-token context intelligence. Technical Report. DeepSeek AI. https://huggingface.co/deepseek- ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf
2026
-
[29]
Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...
2025 arXiv
-
[30]
2018.Aluminum: An asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems
Nikoli Dryden, Naoya Maruyama, Tim Moon, Tom Benson, Andy Yoo, Marc Snir, and Brian Van Essen. 2018.Aluminum: An asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems. Technical Report. Lawrence Livermore Nation...
2018
-
[31]
Ege Erdil. 2025. Inference economics of language models.arXiv preprint arXiv:2506.04645(2025)
2025 arXiv
-
[32]
Amel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun, and Adwait Jog. 2025. NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU Systems. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 1064–1078
2025
-
[33]
Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, et al. 2024. Rdma over ethernet for distributed training at meta scale. InProceedings of the ACM SIGCOMM 2024 Conference. 57–70
2024
-
[34]
Krishnan Geeyarpuram. [n. d.]. Building Custom AI Infrastructure with NVLink Fusion. Presentation slides. https://hoti.org/assets/slides/2025_08_21_day2_ Invited_talk_NVIDIA_NVLink_Fusion.pdf
-
[35]
Raja Gond, Nipun Kwatra, and Ramachandran Ramjee. 2025. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference. arXiv preprint arXiv:2505.11329(2025)
2025 arXiv
-
[36]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)
2024 arXiv
-
[37]
Fei Gui, Kaihui Gao, Li Chen, Dan Li, Vincent Liu, Ran Zhang, Hongbing Yang, and Dian Xiong. 2025. Accelerating design space exploration for LLM training systems with multi-experiment parallel simulation. InProceedings of the 22nd USENIX Symposium on Networked Systems Design a...
2025
-
[38]
Nikos Hardavellas, Michael Ferdman, Babak Falsafi, and Anastasia Ailamaki
-
[39]
Ben Hawks, Gregor von Laszewski, Matthew D Sinclair, Marco Colombo, Shivaram Venkataraman, Rutwik Jain, Yiwei Jiang, Nhan Tran, and Geoffrey Fox. 2025. An MLCommons Scientific Benchmarks Ontology.arXiv preprint arXiv:2511.05614(2025)
2025
-
[40]
Kevin Hsieh, Eiman Ebrahimi, Gwangsun Kim, Niladrish Chatterjee, Mike O’Connor, Nandita Vijaykumar, Onur Mutlu, and Stephen W Keckler. 2016. Transparent offloading and mapping (TOM) enabling programmer-transparent near-data processing in GPU systems.ACM SIGARCH Computer Archit...
2016
-
[41]
Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong. 2023. Tutel: Adaptive mixture-of-experts at scale.arXiv:2206.0338(2023). https://doi.org/10.4855...
2023 doi
-
[42]
Ziyang Jia, Laxmi N Bhuyan, and Daniel Wong. 2024. Pccl: Energy-efficient llm training with power-aware collective communication. In2024 IEEE 42nd International Conference on Computer Design (ICCD). IEEE, 84–91
2024
-
[43]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou 12 Hanna, Florian Bressand, et al . 2024. Mixtral of experts.arXiv preprint arXiv:2401.04088(2024)
2024 arXiv
-
[44]
Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M Aamodt, and John Kim. 2024. Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 885–898
2024
-
[45]
Myoungsoo Jung. 2025. Compute Can’t Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure.arXiv preprint arXiv:2507.07223(2025)
2025 arXiv
-
[46]
2025.Introducing UALink 200G 1.0 Specifica- tion
Nathan Kalyanasundharam. 2025.Introducing UALink 200G 1.0 Specifica- tion. https://ualinkconsortium.org/wp-content/uploads/2025/04/UALink-1.0- White_Paper_FINAL.pdf
2025
-
[47]
Mahmoud Khairy, Vadim Nikiforov, David Nellans, and Timothy G Rogers. 2020. Locality-centric data and threadblock management for massive GPUs. In2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1022–1036
2020
-
[48]
Changkyu Kim, Doug Burger, and Stephen W Keckler. 2002. An adaptive, non- uniform cache structure for wire-delay dominated on-chip caches. InProceedings of the 10th international conference on Architectural support for programming languages and operating systems. 211–222
2002
-
[49]
Hyojong Kim, Ramyad Hadidi, Lifeng Nai, Hyesoon Kim, Nuwan Jayasena, Yasuko Eckert, Onur Kayiran, and Gabriel H Loh. 2017. CODA: Enabling Co- location of Computation and Data for Near-Data Processing.arXiv preprint arXiv:1710.09517(2017)
2017 arXiv
-
[50]
Hyeonjin Kim and William J Song. 2023. LAS: Locality-aware scheduling for GEMM-accelerated convolutions in GPUs.IEEE Transactions on Parallel and Distributed Systems34, 5 (2023), 1479–1494
2023
-
[51]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principl...
2023
-
[52]
Christoph Lameter. 2013. NUMA (Non-uniform memory access): An overview: NUMA becomes more common because memory controllers get close to execu- tion units on microprocessors.Queue11, 7 (2013), 40–51
2013
-
[53]
Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R Tallent, and Kevin J Barker. 2019. Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect.IEEE Transactions on Parallel and Distributed Systems31, 1 (2019), 94–110
2019
-
[54]
Ang Li, Shuaiwen Leon Song, Weifeng Liu, Xu Liu, Akash Kumar, and Henk Corporaal. 2017. Locality-aware CTA clustering for modern GPUs.ACM SIGARCH Computer Architecture News45, 1 (2017), 297–311
2017
-
[55]
Zonghang Li, Wenjiao Feng, Mohsen Guizani, and Hongfang Yu. 2024. TPI- LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices. arXiv:2410.00531 [cs.DC] https://arxiv.org/abs/2410.00531
2024 arXiv
-
[56]
Tim Luhnen, Tobias Marschner, and Sohan Lal. 2024. Benchmarking Thread Block Cluster. In2024 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1–7
2024
-
[57]
Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, and Xiaowen Chu
-
[58]
Clemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl, and Volker Markl
-
[59]
Xiaoyu Ma and David Patterson. 2026. Challenges and Research Directions for Large Language Model Inference Hardware.arXiv preprint arXiv:2601.05047 (2026)
2026
-
[60]
Brandi Martina and Liz Stine. [n. d.].AMD Unveils Strategy to Lead the $1 Trillion Compute Market and Accelerate Next Phase of Growth. Advanced Micro Devices, Inc. https://ir.amd.com/news-events/press-releases/detail/1266/amd- unveils-strategy-to-lead-the-1-trillion-compute-ma...
-
[61]
Philip K McKinley, Yih-jia Tsai, and David F Robinson. 1995. Collective com- munication in wormhole-routed massively parallel computers.Computer28, 12 (1995), 39–50
1995
-
[62]
AI Meta. 2025. The llama 4 herd: The beginning of a new era of natively multi- modal ai innovation.https://ai. meta. com/blog/llama-4-multimodal-intelligence/, checked on4, 7 (2025), 2025
2025
-
[63]
Ugljesa Milic, Oreste Villa, Evgeny Bolotin, Akhil Arunkumar, Eiman Ebrahimi, Aamer Jaleel, Alex Ramirez, and David Nellans. 2017. Beyond the socket: NUMA-aware GPUs. InProceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture(Cambridge, Massachusett...
2017
- [64]
-
[65]
NVIDIA. [n. d.].NVIDIA Fabric Manager NVIDIA Fabric Manager — docs.nvidia.com. [Accessed 31-07-2025]
2025
-
[66]
2016.NVIDIA Tesla P100: The Most Advanced Datacen- ter Accelerator Ever Built — Featuring Pascal GP100, the World’s Fastest GPU
NVIDIA Corporation. 2016.NVIDIA Tesla P100: The Most Advanced Datacen- ter Accelerator Ever Built — Featuring Pascal GP100, the World’s Fastest GPU. Whitepaper WP-08019-001 v01.1. NVIDIA. https://images.nvidia.com/content/ pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf
2016
-
[67]
NVIDIA Corporation. 2022. The NVLink Network Switch. Presentation slides, Hot Chips 34 (HC34). https://hc34.hotchips.org/assets/program/conference/ day2/Network%20and%20Switches/NVSwitch%20HotChips%202022%20r5.pdf
2022
-
[68]
NVIDIA Corporation
NVIDIA Corporation 2023.NVIDIA NVLink SGXLS10 Switch Systems User Man- ual: Introduction. NVIDIA Corporation. https://docs.nvidia.com/networking/ display/sgxh100/introduction Last updated Dec. 13, 2023
2023
-
[69]
2026.NVIDIA CUDA C++ Programming Guide
NVIDIA Corporation. 2026.NVIDIA CUDA C++ Programming Guide. NVIDIA Corporation. Version 13.3
2026
-
[70]
NVIDIA Corporation
NVIDIA Corporation 2026.NVIDIA Multi-Instance GPU User Guide. NVIDIA Corporation. https://docs.nvidia.com/datacenter/tesla/mig-user-guide/latestn
2026
-
[71]
OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, Mark Chen,...
2025 arXiv
-
[72]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[73]
Muhammad Osama, Ryan Swann, Karthik Sangaiah, Sonali Singh, and Ganesh Dasika. [n. d.]. Deep dive into the MI300 compute and memory partition modes; ROCm Blogs — rocm.blogs.amd.com. https://rocm.blogs.amd.com/software- tools-optimization/compute-memory-modes/README.html. [Acce...
2025
-
[74]
Myle Ott, Sam Shleifer, Min Xu, Priya Goyal, Quentin Duval, and Vittorio Caggiano. [n. d.].Fully Sharded Data Parallel: faster AI training with fewer GPUs. https://engineering.fb.com/2021/07/15/open-source/fsdp/
2021
-
[75]
Saptadeep Pal, Daniel Petrisko, Matthew Tomei, Puneet Gupta, Subramanian S Iyer, and Rakesh Kumar. 2019. Architecting waferscale processors-a gpu case study. In2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 250–263
2019
-
[76]
Dylan Patel, Xie Myron, Daniel Nishball, Ivan Chiam, Patrick Zhou, Doug OLaughlin, and Wega Chu. [n. d.]. NVIDIA GTC 2025 – Built For Reason- ing, Vera Rubin, Kyber, CPO, Dynamo Inference, Jensen Math, Feynman. https://semianalysis.com/2025/03/19/nvidia-gtc-2025-built-for-reas...
2025
-
[77]
Dylan Patel, Myron Xie, and Gerald Wong. 2023. AI Capacity Con- straints—CoWoS and HBM Supply Chain.SemiAnalysis.[Online]. A vailable: https://www. semianalysis. com/p/ai-capacity-constraints-cowos-and(2023)
2023
-
[79]
2024.Cross-Stack Optimizations for Sequence-Based Models on GPUs
Suchita Pati. 2024.Cross-Stack Optimizations for Sequence-Based Models on GPUs. The University of Wisconsin-Madison
2024
-
[81]
David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis- Miquel Munguia, Daniel Rothchild, David R So, Maud Texier, and Jeff Dean
-
[82]
Le Qin, Junwei Cui, Weilin Cai, Meng Niu, Yan Yang, and Jiayi Huang. 2025. Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture. 659–674
2025
-
[83]
Gabin Schieffer, Ruimin Shi, Stefano Markidis, Andreas Herten, Jennifer Faj, and Ivy Peng. 2024. Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric. InSC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage a...
2024
-
[84]
David Schor. 2021. 5th Gen CoWoS-S Extends 3 Reticle Size. https://fuse. wikichip.org/news/6031/5th-gen-cowos-s-extends-3-reticle-size/. [Accessed 31-07-2025]
2021
-
[85]
Aashaka Shah, Abhinav Jangda, Binyang Li, Caio Rocha, Changho Hwang, Jithin Jose, Madan Musuvathi, Olli Saarikivi, Peng Cheng, Qinghua Zhou, et al. 2025. MSCCL++: Rethinking GPU Communication Abstractions for Cutting-edge AI Applications.arXiv preprint arXiv:2504.09014(2025)
2025
-
[86]
2025.Optimizing for Low-Latency Communication in Inference Workloads with JAX and XLA
Jaya Shankar, TJ Xu, and Tejash Shah. 2025.Optimizing for Low-Latency Communication in Inference Workloads with JAX and XLA. NVIDIA Tech- nical Blog. https://developer.nvidia.com/blog/optimizing-for-low-latency- communication-in-inference-workloads-with-jax-and-xla/ Blog post
2025
-
[87]
Siyuan Shen, Tommaso Bonato, Zhiyi Hu, Pasquale Jordan, Tiancheng Chen, and Torsten Hoefler. 2025. ATLAHS: An Application-centric Network Sim- ulator Toolchain for AI, HPC, and Distributed Storage. InProceedings of the International Conference for High Performance Computing, N...
2025
-
[88]
Siddharth Singh, Mahua Singh, and Abhinav Bhatele. 2025. The Big Send-off: High Performance Collectives on GPU-based Supercomputers.arXiv preprint arXiv:2504.18658(2025)
2025
-
[89]
Alan Smith, Eric Chapman, Chintan Patel, Raja Swaminathan, John Wuu, Tyrone Huang, Wonjun Jung, Alexander Kaganov, Hugh McIntyre, and Ramon Man- gaser. 2024. 11.1 AMD InstinctTM MI300 series modular chiplet package–HPC and AI accelerator for exa-class systems. In2024 IEEE Inte...
2024
-
[90]
Alan Smith, Gabriel H Loh, John Wuu, Samuel Naffziger, Tyrone Huang, Hugh McIntyre, Ramon Mangaser, Wonjun Jung, and Raja Swaminathan. 2024. AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization. In2024 IEEE Symposium on VLSI Technology and Circuits (VLSI...
2024
-
[91]
Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, et al. 2023. Chakra: Advancing performance benchmarking and co-design using standardized execution traces.arXiv preprint arXiv:23...
2023 arXiv
-
[92]
Andrew Tee, Nicholas Curtis, Noah Wolfe, and Daniel Wong. 2025. The MALL is Open: Exploring Shared Caches and Latency in AMD CDNA™3 GPUs. In Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1110–1116
2025
-
[93]
Ajay Tirumala and Raymond Wong. 2024. Nvidia blackwell platform: Advancing generative ai and accelerated computing. In2024 IEEE Hot Chips 36 Symposium (HCS). IEEE Computer Society, 1–33
2024
-
[94]
Walker J Turner, John W Poulton, John M Wilson, Xi Chen, Stephen G Tell, Matthew Fojtik, Thomas H Greer, Brian Zimmer, Sanquan Song, Nikola Nedovic, et al. 2018. Ground-referenced signaling for intra-chip and short-reach chip-to- chip interconnects. In2018 IEEE Custom Integrat...
2018
-
[95]
Ben Verghese, Scott Devine, Anoop Gupta, and Mendel Rosenblum. 1996. Oper- ating system support for improving data locality on CC-NUMA compute servers. SIGPLAN Not.31, 9 (Sept. 1996), 279–289. https://doi.org/10.1145/248209.237205
1996
-
[96]
Nandita Vijaykumar, Eiman Ebrahimi, Kevin Hsieh, Phillip B Gibbons, and Onur Mutlu. 2018. The locality descriptor: A holistic cross-layer abstraction to express data locality in GPUs. In2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEEE, 829–842
2018
-
[97]
Xizheng Wang, Qingxu Li, Yichi Xu, Gang Lu, Dan Li, Li Chen, Heyang Zhou, Linkang Zheng, Sen Zhang, Yikai Zhu, Yang Liu, Pengcheng Zhang, Kun Qian, Kunling He, Jiaqi Gao, Ennan Zhai, Dennis Cai, and Binzhang Fu. 2025. SimAI: unifying architecture design and performance tuning ...
2025
-
[99]
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. 2022. Sustainable ai: Environmental implications, challenges and opportunities.Pro- ceedings of machine learning and systems4 (20...
2022
-
[100]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025). 14
2025 arXiv
-
[101]
Jie Amy Yang, Jongsoo Park, Srinivas Sridharan, and Ping Tak Peter Tang
-
[102]
Jialiang Zhang, Michael Swift, and Jing Li. 2022. Software-defined address mapping: a case on 3d memory. InProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 70–83
2022
-
[103]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068 (2022)
2022 arXiv
-
[104]
Xia Zhao, Magnus Jahre, Yuhua Tang, Guangda Zhang, and Lieven Eeckhout
-
[105]
Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. 2024. HeteGen: Efficient Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices. InProceedings of Ma- chine Learning and Systems, P. Gibbons, G. Pekhimenko, and C...
2024
-
[106]
Gonza- lez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonza- lez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. InAdvances in Neural Infor...
2024 doi
-
[107]
InConference on Knowledge Discovery and Data Mining (KDD), Vol
Training deep learning recommendation model with quantized collective communications. InConference on Knowledge Discovery and Data Mining (KDD), Vol. 95
-
[108]
Shizhuo Zhu, Illia Shkirko, Jacob Levinson, Zhengrong Wang, and Tony Nowatzki. 2024. SPGPU: Spatially Programmed GPU.IEEE Computer Ar- chitecture Letters(2024). 15
2024
-
[114]
Tianhao Zheng, David Nellans, Arslan Zulfiqar, Mark Stephenson, and Stephen W Keckler. 2016. Towards high performance paged memory for GPUs. In2016 IEEE International Symposium on High Performance Computer Architec- ture (HPCA). IEEE, 345–357
2016
-
[2009]
InProceedings of the 36th annual international symposium on Computer architecture
Reactive NUCA: near-optimal block placement and replication in dis- tributed caches. InProceedings of the 36th annual international symposium on Computer architecture. 184–195
-
[2020]
InProceedings of the 2020 ACM SIGMOD International Con- ference on Management of Data(Portland, OR, USA)(SIGMOD ’20)
Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects. InProceedings of the 2020 ACM SIGMOD International Con- ference on Management of Data(Portland, OR, USA)(SIGMOD ’20). Associ- ation for Computing Machinery, New York, NY, USA, 1633–1649. https: //doi.or...
2020
-
[2022]
The carbon footprint of machine learning training will plateau, then shrink.Computer55, 7 (2022), 18–28
2022
-
[2023]
InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2
NUBA: Non-uniform bandwidth GPUs. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 544–559
-
[2024]
In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
Benchmarking and dissecting the nvidia hopper gpu architecture. In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 656–667
-
[2025]
In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles
LithOS: An operating system for efficient machine learning on GPUs. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 1–17
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.