Pith. sign in

REVIEW 4 major objections 5 minor 115 references

NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Non-uniform network access inside large multi-die GPUs is a first-order cost in GPU-to-GPU communication, and routing plus placement that respect it make latency-sensitive collectives up to 1.8x faster.

desk verdict A useful new framing of intra-socket I/O locality for scale-up GPUs, with real motivating measurements, but headline speedup claims are internally inconsistent and the evaluation is simulation-only. read the letter →

arxiv 2608.00867 v1 pith:XXVHHPR7 submitted 2026-08-01 cs.DC cs.AR

classification cs.DCcs.AR
keywords non-uniformnetworkaccess(NUNA)multi-dieGPUscale-upcollectivecommunicationlatency-sensitivecollectivesLLMinferencethreadblockplacementmemorypage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is trying to establish that, in large multi-die GPUs, the physical location of a threadblock, its data, and the I/O port its traffic leaves from changes remote-memory latency by almost 2x, and that this non-uniform network access (NUNA) is a first-order cost in GPU-to-GPU communication. It claims that two software-controllable policies -- NUNA-aware routing (NAR), which sends latency-sensitive flits only to nearby I/O ports, and NUNA-aware placement (NAP), which puts threadblocks and memory pages near those ports -- jointly make latency-sensitive collectives up to 1.8x faster than a locality-unaware baseline. On an end-to-end evaluation across 12 LLM architectures on 2-64 GPUs, the combined scheme cuts time per output token by 7% on average and up to 28%, with prefill time-to-first-token down 6% on average. The practical stake is that LLM inference decode is dominated by small, latency-bound collectives, which is exactly the regime where NUNA gains are largest.

What carries the argument

The load-bearing object is the NAR domain: a statically partitioned slice of physical address space that constrains a CU's off-chip flits to a subset of I/O ports physically near that CU, while still load-balancing across ports within the subset. NAP's algorithm takes a collective plan (threadblock access patterns over fixed-size chunks), first groups threadblocks and chunks into logical groups, then greedily maps each group to a NAR domain and, within it, to the CU and HBM stack closest to the I/O ports subject to resource budgets such as outstanding-request capacity and bandwidth-delay product. The paper also exposes NAR granularity, the number of domains, as the tuning knob that trades low latency against port load balancing. These mechanisms are implemented with existing threadblock-to-CU affinity masks and page-to-stack placement, plus a modification to address-hashing logic, so they are backward-compatible and disable-able.

What would settle it

Run a 1 MB All-Gather on a next-generation multi-die GPU with a rail-optimized single-level Clos, and compare round-robin placement with full-port hashing against the paper's NAP+NAR scheme. If the measured speedup comes nowhere near 1.8x -- for example, because scale-up link latency or switch queueing so dominates the remote write that the choice of CU, HBM stack, and I/O port no longer changes the total time -- the central claim fails. A simpler check is to measure remote-store latency from the fastest versus slowest compute unit to the same remote address: the paper predicts that gap grows with socket size.

Watch

Extended reading notes

Core claim

The paper's central claim is that the remote-store path between GPUs has two comparable legs: the on-chip trip from a compute unit to an edge I/O port and through the receiving socket to its memory stack, and the scale-up link between sockets; for the modeled next-generation system the on-chip leg reaches 0.9 $\mu$s while the scale-up leg is 1 $\mu$s. Because these legs are comparable, where a threadblock sits, where its buffers live, and which I/O port hashes the traffic determine the latency of a small collective. The paper argues that NAR plus NAP addresses all three choices together: NAR routes flits to a physically close subset of ports, and NAP places communicating threadblocks on nearby CUs and their chunks on nearby HBM stacks, with a greedy two-phase algorithm that resolves resource contention. In simulation, the combined scheme reduces All-Gather, All-Reduce, and All-to-All execution time by up to 80% (up to 1.8x speedup) for small collectives, with a 1.32x geomean communication speedup in prefill and 1.56x in decode across the LLM workloads.

Load-bearing premise

The load-bearing premise is that moving data across the GPU's interior to the edge I/O port (up to 0.9 $\mu$s) costs about as much time as sending it between GPUs over the scale-up link (1 $\mu$s), so the exact placement of threads, data, and ports inside the socket is a first-order factor; if inter-GPU link latency grows faster than on-chip latency, or if compute hides the on-chip transfer, the speedups shrink.

Editorial extensions

If this is right

  • Small, latency-bound collectives (up to roughly 1 MB) benefit most, with 1.5-1.9x speedups that decay to about 1.1x by 64 MB as bandwidth saturation hides the on-chip spatial effect.
  • The scheme cuts decode time-per-output-token by 7% on average and 28% at best, and prefill time-to-first-token by 6% on average and 11% at best, across the 12 evaluated LLM architectures and pod sizes from 2 to 64 GPUs.
  • Placement without routing is not enough: NAP alone can slow down 10 MB collectives because hashed traffic traverses the whole socket and congests the on-chip network; routing and placement must be designed jointly.
  • The best NAR granularity depends on threadblock count, so the allocation algorithm should choose the number of domains per workload rather than fixing it in hardware.
  • Concurrent compute kernels degrade the NUNA-aware collective less than the baseline, with degradation appearing only at roughly twice the GEMM size, because localized I/O traffic avoids mid-die congestion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 1.8x figure is tied to the modeled parameter regime: on sockets where the scale-up link dominates the on-chip leg the same policies would yield less, while on still-larger dies the gains could exceed 1.8x.
  • The same NAR/NAP machinery generalizes to any latency-sensitive off-chip traffic with a known access plan, such as remote atomics or sparse attention token routing, not just the three collectives evaluated.
  • Because the paper serializes compute and communication, its end-to-end numbers are optimistic about exposure; if future systems overlap compute and collectives well, per-collective speedups would remain but the 7% TPOT gain would shrink.
  • A natural testable extension is to derive the topology input to NAP from measured per-CU remote-store latencies rather than geometric layout, so the method stays calibrated if real meshes or packaging differ from the modeled 24x6 mesh.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces the term non-uniform network access (NUNA) to describe how intra-socket wire distance between compute units, HBM stacks, and I/O ports makes inter-GPU scale-up communication latency depend on spatial placement. It proposes two optimizations: NUNA-aware routing (NAR), which restricts latency-sensitive flits to I/O ports physically closer to the requesting CU via statically partitioned address ranges, and NUNA-aware placement (NAP), which places collective threadblocks and memory pages near those I/O ports. The paper profiles three real AMD GPU systems to show that remote-access latency varies with CU and address location, then evaluates NAR and NAP in the ASTRA-sim 3.0 simulator. The headline claims are up to 1.8x collective speedups for NAP+NAR over the locality-unaware baseline and 7% mean (28% max) end-to-end time-per-output-token improvement for LLM inference.

Significance. If the claimed speedups are validated, the paper makes a timely contribution: it identifies a new locality dimension for scale-up GPU systems and provides a coherent, backward-compatible software/hardware mechanism for exploiting it. Strengths include the real-system profiling of MI210, MI355X, and MI300X, which documents that remote-access latency variation grows with socket size; the isolation study showing that NAR is necessary for NAP to help; and the sensitivity studies for collective size and NAR granularity. The paper does not, however, ship code or validation data, and its central quantitative claims depend on a self-built simulator and on two strong assumptions: the NoC/scale-up latency ratio and a serial compute-communication model. Those assumptions are load-bearing for the 1.8x and 7% headline numbers, so the result is conditional until the assumptions are tested and the internal speedup accounting is reconciled.

major comments (4)
  1. [§1 and Abstract] The manuscript states two incompatible headline speedups. The abstract claims up to 1.8x faster collectives for NAP+NAR, while the introduction says 'NAP and NAR applied together can reduce collective execution time by up to 80%' (a 5x speedup). Both cannot describe the same comparison against the same baseline; the 80% reduction is hard to reconcile with the 1.8x upper bound reported later. This needs to be corrected and the exact comparison (which configuration, collective, message size, and pod size) stated for each number.
  2. [§3.1.3 and §1] The motivating 1.8x 'uncontended GPU-to-GPU transfer slowdown' is not supported by the paper's own latency numbers. Section 3.1.3 gives a NoC transfer of up to 0.9 us and a scale-up latency of 1 us; with a 3x NoC spread (0.3 to 0.9 us) plus a fixed 1 us scale-up term, the worst/best total transfer ratio is (1.9)/(1.3) = 1.46x, not 1.8x. Because NAP and NAR act only on the intra-socket segment, any fixed scale-up latency dilutes the benefit, and the claimed 1.8x collective speedup therefore depends on the assumed latency ratio. The paper provides no sensitivity analysis over this ratio. I ask the authors to (i) show exactly how the 1.8x transfer-slowdown and 1.8x collective-speedup numbers are derived, and (ii) add a sensitivity study varying the NoC/scale-up latency ratio (e.g., 0.3x to 3x of the assumed 0.9 us) and report collective speedup as a function of that ratio.
  3. [§5.2] The end-to-end inference results are computed under a serial compute/communication model, as stated in Section 5.2: 'we model compute and communication serially since intra-batch overlap is ongoing research.' This assumption removes any overlap between collective communication and compute, which exposes the full communication time and therefore inflates the measured benefit of faster collectives. Since the paper itself cites prior overlap work (e.g., T3 [80] and Splitwise [78]), a serial model is a worst-case assumption, not a neutral one. The 7% mean TPOT claim is thus conditional on zero overlap. The authors should quantify how the end-to-end speedup changes under partial or full overlap (e.g., by applying an overlap factor to the collective time in their analytical model).
  4. [§5.1 and §3.2] The central collective-speedup results come from ASTRA-sim 3.0, a simulator co-authored by this team (reference [98]), and the paper does not ship the simulator configuration, trace-generation tool, or validation data. Section 3.2's real-system profiling validates the existence of latency variation on current hardware, but it does not validate that the simulated 24x6 mesh with the chosen hop latencies reproduces those measurements. Without a calibration/validation experiment comparing simulated remote-store latency distributions against the measured MI300X values, the quantitative claims (1.8x for NAP+NAR, up to 28% TPOT improvement) cannot be independently checked. I would like the paper to include a validation section, or at minimum release the exact configuration and a comparison of simulated versus measured latency distributions.
minor comments (5)
  1. [§1, §4.1.2, §6.3] There are several typos: 'Specifcially' in the introduction, 'to to limit' in Section 4.1.2, and 'traffics' in Section 6.3.
  2. [§2.2, Figure 2] Figure 2 reports collective sizes derived from an in-house trace-generation tool, but the methodology is deferred to Section 5.2 and no validation is given for the trace shapes or collective sizes. A short validation sentence (e.g., comparison with known model configurations) would help the reader trust the prefill/decode CDFs.
  3. [§5.3, Table 2] The simulated system table lists '144 I/O ports (72 per die)' for a 4-compute-die socket, while Section 4.1.1 and Figure 5 describe a socket with four I/O ports. The figure is clearly illustrative, but the relationship between the illustrative four-port model and the simulated 144-port model should be stated explicitly to avoid confusion.
  4. [§6.2] The description of the UA latency histogram as 'the Gaussian distribution of individual path latencies in a mesh' is informal; the reported 920 ns average and 1,320 ns maximum are more informative. Reporting percentiles (e.g., p50/p99/max) directly in the text would make the comparison clearer.
  5. [§6.7] The phrase 'exposed communication ratio is approximately 4:1 for prefill and 5:1 for decode' is ambiguous: it is not clear whether this is compute-to-communication or communication-to-compute, and how the ratio is computed. Clarifying the definition would help the Amdahl's law argument.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the speedup claims are simulated outcomes from stated models, not predictions that reduce to their inputs by construction.

full rationale

The paper's central claims are empirical and simulation results conditional on stated parameters, not derivations that reduce to their inputs. Section 3.2 profiles real AMD Instinct GPUs (MI210, MI355X, MI300X) and measures latency variation, providing an independent anchor for the NUNA phenomenon. The 1.5x/1.8x collective speedups and 7%/28% TPOT gains are outputs of ASTRA-sim 3.0 simulations using explicit NoC mesh, hop latencies, I/O port counts, and scale-up switch parameters (Table 2, Section 5.3), rather than relabeled fits. NAR and NAP are evaluated against a round-robin/hash baseline; their benefit is computed by the simulator's cycle-level model, not assumed by construction. The use of ASTRA-sim 3.0 [98] is a self-citation, but it is used as a simulation tool with independently stated inputs, and the paper also presents real-system profiling; no load-bearing claim is justified solely by an unverified self-cited theorem. The paper's '1.8x/1.45x transfer slowdowns' in Section 1 are not arithmetically consistent with the 0.9 us NoC / 1 us scale-up numbers in Section 3.1.3, and the serial compute/communication model in Section 5.2 may inflate end-to-end gains, but these are correctness and robustness concerns, not circularity. No equation in the paper equates an output to an input by construction.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central results rest on a small number of assumed system parameters and modeling choices. The key free parameters are NAR granularity, resource budgets in Algorithm 1, and the unreported threshold for when to use NAR. The key axioms are that intra-socket NoC latency will be comparable to scale-up latency, that the representative 24x6 mesh plus Clos topology captures future systems, and that serial compute/communication modeling is a fair conservative assumption. No new physical entities are introduced.

free parameters (5)
  • NAR granularity (narmax) = 2-12 domains in sensitivity experiments; compiler-selected in practice
    Controls the latency vs bandwidth trade-off in NAR; no closed-form optimum provided; the paper says thresholds can be statically determined but does not give them.
  • CU outstanding request budget (Rmax CU) = 700 requests per CU
    Used in Algorithm 1 to limit co-located threadblocks; value assumed from hardware, not derived or validated.
  • HBM stack in-flight budget (bandwidth-delay product) = not specified numerically
    Rmax for memory in Algorithm 1; value is not reported, making reproduction difficult.
  • Number of threadblocks for collectives ('best-performing') = not reported
    Microbenchmarks in Section 5.1 use the best-performing number of threadblocks; the value is omitted, which could bias speedups.
  • NAR usage threshold (collective size) = unstated; 'statically determined' per architecture
    Footnote 1 says the optimal threshold can be statically determined, but the paper does not specify it.
assumptions (5)
  • domain assumption Intra-socket NoC transfer delay will be comparable to scale-up link latency (around 1 us) on next-generation GPUs, making intra-socket distance a first-order factor for inter-GPU communication.
    Invoked in Section 3.1.3 with values 0.9 us NoC vs 1 us scale-up; if scale-up latency dominates, NUNA speedups shrink.
  • domain assumption The simulated 24x6 mesh NoC with X-Y routing, 4 compute dies, and single-level Clos scale-up is representative of future multi-die GPU systems.
    Section 5.3 and Figure 8; no sensitivity analysis across alternative topologies or routing policies.
  • domain assumption Collectives up to roughly 100 MB are latency-bound rather than bandwidth-bound, so point-to-point latency reductions translate into collective speedups.
    Section 2.2 alpha-beta analysis; supports the focus on small collectives.
  • domain assumption Compute and communication do not overlap in inference execution, so reducing communication time directly reduces end-to-end time.
    Section 5.2 states serial modeling; conservative but may overstate communication exposure and hence NUNA benefit.
  • domain assumption The real-system latency variation in Figure 4 is caused by physical distance rather than by contention or hashing artifacts.
    The authors profile real GPUs and observe variation, but the attribution to spatial distance is an interpretation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems." pith.science (2026). https://pith.science/paper/XXVHHPR7

@misc{pith2026260800867,
  author       = {Pith},
  title        = {Pith review of: NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XXVHHPR7}},
  note         = {Machine review of arXiv:2608.00867}
}
read the original abstract

Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly. While previous research has optimized for compute and memory locality within a socket, the spatial impact on inter-GPU communication has not been well-studied. We introduce the term non-uniform network access (NUNA) to describe this emerging optimization dimension in multi-GPU systems. We specifically focus on latency-sensitive collective communication, common in machine learning inference. First, we highlight the need for NUNA-aware routing (NAR), choosing optimized, spatially-aware inter-GPU paths in large scale-up network topologies. Second, we introduce NUNA-aware placement (NAP), placing threadblocks and data near I/O to optimize the inter-GPU traffic. We demonstrate that the NAP optimizations alone offer up to 1.5x collective speedups over a locality-unaware baseline. Combining NAP with NAR yields up to 1.8x faster collectives over the locality-unaware baseline. This leads to 7% mean (28% max) time per output token speedup in machine learning inference.

Figures

Figures reproduced from arXiv: 2608.00867 by the authors.

Figure 1
Figure 1. Possible high and low latency inter-GPU commu [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Cumulative distribution of collective sizes (in out [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Representative scale-up system evaluated in this [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Real-system measured remote store latencies be [PITH_FULL_IMAGE:figures/full_fig_p003_4.png]
Figure 5
Figure 5. Figure 5: Static partitioning some of the physical address [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Example of (a) baseline and (b) granularity-2 NAR [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Visual example of logical allocation and physical placement steps for a two-GPU, two-channel All-Gather implemented [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Evaluated on-chip network. 24×6 mesh connects the CUs in compute dies (yellow). Mesh routers in vertical edge connect to 12 I/O ports (green). Routers in the horizontal edges connect to HBM stacks (purple) with full connectivity to all 16 memory channels within each st…
Figure 10
Figure 10. Figure 10: Latency distributions of remote stores in an eight [PITH_FULL_IMAGE:figures/full_fig_p009_10.png]
Figure 11
Figure 11. Figure 11: Isolation of techniques for small (100 kB) and [PITH_FULL_IMAGE:figures/full_fig_p009_11.png]
Figure 9
Figure 9. Figure 9: Execution speedups across pod and collective sizes [PITH_FULL_IMAGE:figures/full_fig_p009_9.png]
Figure 13
Figure 13. Figure 13: The speedup of NAR (UA+NAR) over the baseline [PITH_FULL_IMAGE:figures/full_fig_p010_13.png]
Figure 14
Figure 14. Figure 14: Impact of concurrent GEMM kernel on un￾aware baseline (UA) and NUNA-optimized configuration (NAP+NAR). Execution time relative to baseline (UA) of an eight-GPU, 1 MB All-Gather was measured for a sweep of matrix sizes. speed over a sweep of GEMM sizes for unaware (UA)…
Figure 15
Figure 15. Figure 15: End-to-end time (lower is better) of 64-GPU con [PITH_FULL_IMAGE:figures/full_fig_p011_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

115 extracted references · 52 canonical work pages

  1. [80]

    Sinclair

    Suchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena, and Matthew D. Sinclair. 2024. T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives. InProceedings of the 29th ACM International Confer- ence on Architectural Support for Programming Languages and Operating Systems, Volume 2(La Jolla, CA, USA)(ASPLOS ’24). Asso...

  2. [78]

    Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient generative llm inference using phase splitting. In2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 118–132

  3. [98]

    ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling

    William Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, and Bradford M. Beckmann. 2026. ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fi...

  4. [1]

    [n. d.]. GitHub - ROCm/rccl: ROCm Communication Collectives Library (RCCL) — github.com. https://github.com/ROCm/rccl. [Accessed 31-07-2025]

  5. [2]

    [n. d.]. NVLink & NVSwitch for Advanced Multi-GPU Communication — nvidia.com. https://www.nvidia.com/en-us/data-center/nvlink/. [Accessed 31-07-2025]

  6. [3]

    [n. d.]. SPCL - ATLAHS — spcl.inf.ethz.ch. https://spcl.inf.ethz.ch/Research/ Scalable_Networking/ATLAHS/. [Accessed 10-06-2026]

  7. [4]

    NVIDIA GB200 Interconnect Architecture Analysis: NVLink, InfiniBand, and Future Trends - NADDOD Blog — naddod.com

    2024. NVIDIA GB200 Interconnect Architecture Analysis: NVLink, InfiniBand, and Future Trends - NADDOD Blog — naddod.com. https://www.naddod.com/blog/nvidia-gb200-interconnect-architecture- analysis-nvlink-infiniband-and-future-trends. [Accessed 31-07-2025]

  8. [5]

    NVIDIA Collective Communications Library (NCCL)

    2025. NVIDIA Collective Communications Library (NCCL). https://developer. nvidia.com/nccl. [Accessed 31-07-2025]

Show all 115 references
  1. [6]

    Advanced Micro Devices, Inc. [n. d.].AMD Instinct TM MI250 microarchitecture — ROCm Documentation. Online documentation page

  2. [7]

    2021.Introducing AMD CDNA TM 2 Architecture

    Advanced Micro Devices, Inc. 2021.Introducing AMD CDNA TM 2 Architecture. White Paper. AMD. https://www.amd.com/content/dam/amd/en/documents/ instinct-business-docs/white-papers/amd-cdna2-white-paper.pdf Copyright notice shows 2021; document includes performance notes as of Ja...

  3. [8]

    2025.HIP documentation; HIP 6.4.43484 Documen- tation

    Advanced Micro Devices Inc. 2025.HIP documentation; HIP 6.4.43484 Documen- tation. [Accessed 01-08-2025]

  4. [9]

    Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In18th USENIX Symposium on Operating Systems Design and Implementat...

  5. [10]

    AMD. [n. d.]. AMD InstinctTM MI350X GPUs. https://www.amd.com/en/ products/accelerators/instinct/mi350/mi350x.html. [Accessed 31-07-2025]

  6. [11]

    AMD. 2025. AMD Advancing AI 2025. https://www.amd.com/en/corporate/ events/advancing-ai.html. [Accessed 31-07-2025]

  7. [12]

    Michael Andersch, Greg Palmer, Ronny Krashinsky, Nick Stam, Vishal Mehta, Gonzalo Brito, and Sridhar Ramaswamy. 2022. NVIDIA Hopper Architecture In-Depth. https://developer.nvidia.com/blog/nvidia-hopper-architecture-in- depth/. [Accessed 31-07-2025]

  8. [13]

    Akhil Arunkumar, Evgeny Bolotin, Benjamin Cho, Ugljesa Milic, Eiman Ebrahimi, Oreste Villa, Aamer Jaleel, Carole-Jean Wu, and David Nellans. 2017. MCM-GPU: Multi-chip-module GPUs for continued performance scalability. ACM SIGARCH Computer Architecture News45, 2 (2017), 320–332

  9. [14]

    Mohammad Banikazemi, Vijay Moorthy, and Dhabaleswar K Panda. 1998. Effi- cient collective communication on heterogeneous networks of workstations. InProceedings. 1998 International Conference on Parallel Processing (Cat. No. 98EX205). IEEE, 460–467

  10. [15]

    Jeff Barr. 2019. Amazon ec2 update–inf1 instances with AWS inferentia chips for high performance cost-effective inferencing.A WS News Blog(2019)

  11. [16]

    Bradford M Beckmann and David A Wood. 2004. Managing wire delay in large chip-multiprocessor caches. In37th International Symposium on Microarchitec- ture (MICRO-37’04). IEEE, 319–330

  12. [17]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901

  13. [18]

    Javier Cabezas, Lluís Vilanova, Isaac Gelado, Thomas B Jablin, Nacho Navarro, and Wen-mei W Hwu. 2015. Automatic parallelization of kernels in shared- memory multi-GPU nodes. InProceedings of the 29th ACM on International Conference on Supercomputing. 3–13

  14. [19]

    Li-Jhan Chen, Hsiang-Yun Cheng, Po-Han Wang, and Chia-Lin Yang. 2017. Improving GPGPU Performance via Cache Locality Aware Thread Block Sched- uling.IEEE Computer Architecture Letters16, 2 (2017), 127–131. https: //doi.org/10.1109/LCA.2017.2693371

  15. [20]

    Jaehong Cho, Hyunmin Choi, and Jongse Park. 2025. LLMServingSim2.0: A Unified Simulator for Heterogeneous Hardware and Serving Techniques in LLM Infrastructure .IEEE Computer Architecture Letters24, 02 (July 2025), 361–364. https://doi.org/10.1109/LCA.2025.3628325

  16. [21]

    Jack Choquette. 2022. Nvidia hopper gpu: Scaling performance. In2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1–46

  17. [22]

    Charles Clos. 1953. A study of non-blocking switching networks.Bell System Technical Journal32, 2 (1953), 406–424

  18. [23]

    Patrick H Coppock, Brian Zhang, Eliot H Solomon, Vasilis Kypriotis, Leon Yang, Bikash Sharma, Dan Schatzberg, Todd C Mowry, and Dimitrios Skarlatos

  19. [24]

    David Correa. 2025. Artificial Intelligence Chip Market Expected to Reach $460.9 Billion by 2034. https://www.einpresswire.com/article/866225614/artificial- intelligence-chip-market-expected-to-reach-460-9-billion-by-2034. [Accessed Nov. 11, 2025]

  20. [25]

    Meghan Cowan, Saeed Maleki, Madanlal Musuvathi, Olli Saarikivi, and Yifan Xiong. 2023. Mscclang: Microsoft collective communication language. InPro- ceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volu...

  21. [26]

    2004.Principles and practices of interconnection networks

    William James Dally and Brian Patrick Towles. 2004.Principles and practices of interconnection networks. Elsevier

  22. [27]

    Preyesh Dalmia, Rajesh Shashi Kumar, and Matthew D Sinclair. 2024. CPElide: Efficient multi-chiplet GPU implicit synchronization. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 700–717

  23. [28]

    2026.DeepSeek-V4: Towards highly efficient million-token context intelligence

    DeepSeek-AI. 2026.DeepSeek-V4: Towards highly efficient million-token context intelligence. Technical Report. DeepSeek AI. https://huggingface.co/deepseek- ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf

  24. [29]

    Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

    DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei L...

  25. [30]

    2018.Aluminum: An asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems

    Nikoli Dryden, Naoya Maruyama, Tim Moon, Tom Benson, Andy Yoo, Marc Snir, and Brian Van Essen. 2018.Aluminum: An asynchronous, GPU-aware communication library optimized for large-scale training of deep neural networks on HPC systems. Technical Report. Lawrence Livermore Nation...

  26. [31]

    Ege Erdil. 2025. Inference economics of language models.arXiv preprint arXiv:2506.04645(2025)

  27. [32]

    Amel Fatima, Yang Yang, Yifan Sun, Rachata Ausavarungnirun, and Adwait Jog. 2025. NetCrafter: Tailoring Network Traffic for Non-Uniform Bandwidth Multi-GPU Systems. InProceedings of the 52nd Annual International Symposium on Computer Architecture. 1064–1078

  28. [33]

    Adithya Gangidi, Rui Miao, Shengbao Zheng, Sai Jayesh Bondu, Guilherme Goes, Hany Morsy, Rohit Puri, Mohammad Riftadi, Ashmitha Jeevaraj Shetty, Jingyi Yang, et al. 2024. Rdma over ethernet for distributed training at meta scale. InProceedings of the ACM SIGCOMM 2024 Conference. 57–70

  29. [34]

    Krishnan Geeyarpuram. [n. d.]. Building Custom AI Infrastructure with NVLink Fusion. Presentation slides. https://hoti.org/assets/slides/2025_08_21_day2_ Invited_talk_NVIDIA_NVLink_Fusion.pdf

  30. [35]

    Raja Gond, Nipun Kwatra, and Ramachandran Ramjee. 2025. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference. arXiv preprint arXiv:2505.11329(2025)

  31. [36]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)

  32. [37]

    Fei Gui, Kaihui Gao, Li Chen, Dan Li, Vincent Liu, Ran Zhang, Hongbing Yang, and Dian Xiong. 2025. Accelerating design space exploration for LLM training systems with multi-experiment parallel simulation. InProceedings of the 22nd USENIX Symposium on Networked Systems Design a...

  33. [38]

    Nikos Hardavellas, Michael Ferdman, Babak Falsafi, and Anastasia Ailamaki

  34. [39]

    Ben Hawks, Gregor von Laszewski, Matthew D Sinclair, Marco Colombo, Shivaram Venkataraman, Rutwik Jain, Yiwei Jiang, Nhan Tran, and Geoffrey Fox. 2025. An MLCommons Scientific Benchmarks Ontology.arXiv preprint arXiv:2511.05614(2025)

  35. [40]

    Kevin Hsieh, Eiman Ebrahimi, Gwangsun Kim, Niladrish Chatterjee, Mike O’Connor, Nandita Vijaykumar, Onur Mutlu, and Stephen W Keckler. 2016. Transparent offloading and mapping (TOM) enabling programmer-transparent near-data processing in GPU systems.ACM SIGARCH Computer Archit...

  36. [41]

    Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, Joe Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong. 2023. Tutel: Adaptive mixture-of-experts at scale.arXiv:2206.0338(2023). https://doi.org/10.4855...

  37. [42]

    Ziyang Jia, Laxmi N Bhuyan, and Daniel Wong. 2024. Pccl: Energy-efficient llm training with power-aware collective communication. In2024 IEEE 42nd International Conference on Computer Design (ICCD). IEEE, 84–91

  38. [43]

    Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou 12 Hanna, Florian Bressand, et al . 2024. Mixtral of experts.arXiv preprint arXiv:2401.04088(2024)

  39. [44]

    Zhixian Jin, Christopher Rocca, Jiho Kim, Hans Kasan, Minsoo Rhu, Ali Bakhoda, Tor M Aamodt, and John Kim. 2024. Uncovering Real GPU NoC Characteristics: Implications on Interconnect Architecture. In2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 885–898

  40. [45]

    Myoungsoo Jung. 2025. Compute Can’t Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure.arXiv preprint arXiv:2507.07223(2025)

  41. [46]

    2025.Introducing UALink 200G 1.0 Specifica- tion

    Nathan Kalyanasundharam. 2025.Introducing UALink 200G 1.0 Specifica- tion. https://ualinkconsortium.org/wp-content/uploads/2025/04/UALink-1.0- White_Paper_FINAL.pdf

  42. [47]

    Mahmoud Khairy, Vadim Nikiforov, David Nellans, and Timothy G Rogers. 2020. Locality-centric data and threadblock management for massive GPUs. In2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1022–1036

  43. [48]

    Changkyu Kim, Doug Burger, and Stephen W Keckler. 2002. An adaptive, non- uniform cache structure for wire-delay dominated on-chip caches. InProceedings of the 10th international conference on Architectural support for programming languages and operating systems. 211–222

  44. [49]

    Hyojong Kim, Ramyad Hadidi, Lifeng Nai, Hyesoon Kim, Nuwan Jayasena, Yasuko Eckert, Onur Kayiran, and Gabriel H Loh. 2017. CODA: Enabling Co- location of Computation and Data for Near-Data Processing.arXiv preprint arXiv:1710.09517(2017)

  45. [50]

    Hyeonjin Kim and William J Song. 2023. LAS: Locality-aware scheduling for GEMM-accelerated convolutions in GPUs.IEEE Transactions on Parallel and Distributed Systems34, 5 (2023), 1479–1494

  46. [51]

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. InProceedings of the 29th Symposium on Operating Systems Principl...

  47. [52]

    Christoph Lameter. 2013. NUMA (Non-uniform memory access): An overview: NUMA becomes more common because memory controllers get close to execu- tion units on microprocessors.Queue11, 7 (2013), 40–51

  48. [53]

    Ang Li, Shuaiwen Leon Song, Jieyang Chen, Jiajia Li, Xu Liu, Nathan R Tallent, and Kevin J Barker. 2019. Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect.IEEE Transactions on Parallel and Distributed Systems31, 1 (2019), 94–110

  49. [54]

    Ang Li, Shuaiwen Leon Song, Weifeng Liu, Xu Liu, Akash Kumar, and Henk Corporaal. 2017. Locality-aware CTA clustering for modern GPUs.ACM SIGARCH Computer Architecture News45, 1 (2017), 297–311

  50. [55]

    Zonghang Li, Wenjiao Feng, Mohsen Guizani, and Hongfang Yu. 2024. TPI- LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices. arXiv:2410.00531 [cs.DC] https://arxiv.org/abs/2410.00531

  51. [56]

    Tim Luhnen, Tobias Marschner, and Sohan Lal. 2024. Benchmarking Thread Block Cluster. In2024 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1–7

  52. [57]

    Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, and Xiaowen Chu

  53. [58]

    Clemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl, and Volker Markl

  54. [59]

    Xiaoyu Ma and David Patterson. 2026. Challenges and Research Directions for Large Language Model Inference Hardware.arXiv preprint arXiv:2601.05047 (2026)

  55. [60]

    Brandi Martina and Liz Stine. [n. d.].AMD Unveils Strategy to Lead the $1 Trillion Compute Market and Accelerate Next Phase of Growth. Advanced Micro Devices, Inc. https://ir.amd.com/news-events/press-releases/detail/1266/amd- unveils-strategy-to-lead-the-1-trillion-compute-ma...

  56. [61]

    Philip K McKinley, Yih-jia Tsai, and David F Robinson. 1995. Collective com- munication in wormhole-routed massively parallel computers.Computer28, 12 (1995), 39–50

  57. [62]

    AI Meta. 2025. The llama 4 herd: The beginning of a new era of natively multi- modal ai innovation.https://ai. meta. com/blog/llama-4-multimodal-intelligence/, checked on4, 7 (2025), 2025

  58. [63]

    Ugljesa Milic, Oreste Villa, Evgeny Bolotin, Akhil Arunkumar, Eiman Ebrahimi, Aamer Jaleel, Alex Ramirez, and David Nellans. 2017. Beyond the socket: NUMA-aware GPUs. InProceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture(Cambridge, Massachusett...

  59. [64]

    Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tul- loch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, Liang Luo, Jie Amy Yang, Leon Gao, Dmytro Ivchenko, Aarti Basant, Yuxi Hu, Jiyan Yang, Ehsan K. Ardestani, Xiaodong Wang, Rakesh Kom...

  60. [65]

    NVIDIA. [n. d.].NVIDIA Fabric Manager NVIDIA Fabric Manager — docs.nvidia.com. [Accessed 31-07-2025]

  61. [66]

    2016.NVIDIA Tesla P100: The Most Advanced Datacen- ter Accelerator Ever Built — Featuring Pascal GP100, the World’s Fastest GPU

    NVIDIA Corporation. 2016.NVIDIA Tesla P100: The Most Advanced Datacen- ter Accelerator Ever Built — Featuring Pascal GP100, the World’s Fastest GPU. Whitepaper WP-08019-001 v01.1. NVIDIA. https://images.nvidia.com/content/ pdf/tesla/whitepaper/pascal-architecture-whitepaper.pdf

  62. [67]

    NVIDIA Corporation. 2022. The NVLink Network Switch. Presentation slides, Hot Chips 34 (HC34). https://hc34.hotchips.org/assets/program/conference/ day2/Network%20and%20Switches/NVSwitch%20HotChips%202022%20r5.pdf

  63. [68]

    NVIDIA Corporation

    NVIDIA Corporation 2023.NVIDIA NVLink SGXLS10 Switch Systems User Man- ual: Introduction. NVIDIA Corporation. https://docs.nvidia.com/networking/ display/sgxh100/introduction Last updated Dec. 13, 2023

  64. [69]

    2026.NVIDIA CUDA C++ Programming Guide

    NVIDIA Corporation. 2026.NVIDIA CUDA C++ Programming Guide. NVIDIA Corporation. Version 13.3

  65. [70]

    NVIDIA Corporation

    NVIDIA Corporation 2026.NVIDIA Multi-Instance GPU User Guide. NVIDIA Corporation. https://docs.nvidia.com/datacenter/tesla/mig-user-guide/latestn

  66. [71]

    OpenAI, :, Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Altman, Andy Applebaum, Edwin Arbus, Rahul K. Arora, Yu Bai, Bowen Baker, Haiming Bao, Boaz Barak, Ally Bennett, Tyler Bertao, Nivedita Brett, Eugene Brevdo, Greg Brockman, Sebastien Bubeck, Che Chang, Kai Chen, Mark Chen,...

  67. [72]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  68. [73]

    Muhammad Osama, Ryan Swann, Karthik Sangaiah, Sonali Singh, and Ganesh Dasika. [n. d.]. Deep dive into the MI300 compute and memory partition modes; ROCm Blogs — rocm.blogs.amd.com. https://rocm.blogs.amd.com/software- tools-optimization/compute-memory-modes/README.html. [Acce...

  69. [74]

    Myle Ott, Sam Shleifer, Min Xu, Priya Goyal, Quentin Duval, and Vittorio Caggiano. [n. d.].Fully Sharded Data Parallel: faster AI training with fewer GPUs. https://engineering.fb.com/2021/07/15/open-source/fsdp/

  70. [75]

    Saptadeep Pal, Daniel Petrisko, Matthew Tomei, Puneet Gupta, Subramanian S Iyer, and Rakesh Kumar. 2019. Architecting waferscale processors-a gpu case study. In2019 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 250–263

  71. [76]

    Dylan Patel, Xie Myron, Daniel Nishball, Ivan Chiam, Patrick Zhou, Doug OLaughlin, and Wega Chu. [n. d.]. NVIDIA GTC 2025 – Built For Reason- ing, Vera Rubin, Kyber, CPO, Dynamo Inference, Jensen Math, Feynman. https://semianalysis.com/2025/03/19/nvidia-gtc-2025-built-for-reas...

  72. [77]

    Dylan Patel, Myron Xie, and Gerald Wong. 2023. AI Capacity Con- straints—CoWoS and HBM Supply Chain.SemiAnalysis.[Online]. A vailable: https://www. semianalysis. com/p/ai-capacity-constraints-cowos-and(2023)

  73. [79]

    2024.Cross-Stack Optimizations for Sequence-Based Models on GPUs

    Suchita Pati. 2024.Cross-Stack Optimizations for Sequence-Based Models on GPUs. The University of Wisconsin-Madison

  74. [81]

    David Patterson, Joseph Gonzalez, Urs Hölzle, Quoc Le, Chen Liang, Lluis- Miquel Munguia, Daniel Rothchild, David R So, Maud Texier, and Jeff Dean

  75. [82]

    Le Qin, Junwei Cui, Weilin Cai, Meng Niu, Yan Yang, and Jiayi Huang. 2025. Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks. InProceedings of the 58th IEEE/ACM International Symposium on Microarchitecture. 659–674

  76. [83]

    Gabin Schieffer, Ruimin Shi, Stefano Markidis, Andreas Herten, Jennifer Faj, and Ivy Peng. 2024. Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric. InSC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage a...

  77. [84]

    David Schor. 2021. 5th Gen CoWoS-S Extends 3 Reticle Size. https://fuse. wikichip.org/news/6031/5th-gen-cowos-s-extends-3-reticle-size/. [Accessed 31-07-2025]

  78. [85]

    Aashaka Shah, Abhinav Jangda, Binyang Li, Caio Rocha, Changho Hwang, Jithin Jose, Madan Musuvathi, Olli Saarikivi, Peng Cheng, Qinghua Zhou, et al. 2025. MSCCL++: Rethinking GPU Communication Abstractions for Cutting-edge AI Applications.arXiv preprint arXiv:2504.09014(2025)

  79. [86]

    2025.Optimizing for Low-Latency Communication in Inference Workloads with JAX and XLA

    Jaya Shankar, TJ Xu, and Tejash Shah. 2025.Optimizing for Low-Latency Communication in Inference Workloads with JAX and XLA. NVIDIA Tech- nical Blog. https://developer.nvidia.com/blog/optimizing-for-low-latency- communication-in-inference-workloads-with-jax-and-xla/ Blog post

  80. [87]

    Siyuan Shen, Tommaso Bonato, Zhiyi Hu, Pasquale Jordan, Tiancheng Chen, and Torsten Hoefler. 2025. ATLAHS: An Application-centric Network Sim- ulator Toolchain for AI, HPC, and Distributed Storage. InProceedings of the International Conference for High Performance Computing, N...

  81. [88]

    Siddharth Singh, Mahua Singh, and Abhinav Bhatele. 2025. The Big Send-off: High Performance Collectives on GPU-based Supercomputers.arXiv preprint arXiv:2504.18658(2025)

  82. [89]

    Alan Smith, Eric Chapman, Chintan Patel, Raja Swaminathan, John Wuu, Tyrone Huang, Wonjun Jung, Alexander Kaganov, Hugh McIntyre, and Ramon Man- gaser. 2024. 11.1 AMD InstinctTM MI300 series modular chiplet package–HPC and AI accelerator for exa-class systems. In2024 IEEE Inte...

  83. [90]

    Alan Smith, Gabriel H Loh, John Wuu, Samuel Naffziger, Tyrone Huang, Hugh McIntyre, Ramon Mangaser, Wonjun Jung, and Raja Swaminathan. 2024. AMD Instinct™MI300X Accelerator: Packaging and Architecture Co-Optimization. In2024 IEEE Symposium on VLSI Technology and Circuits (VLSI...

  84. [91]

    Srinivas Sridharan, Taekyung Heo, Louis Feng, Zhaodong Wang, Matt Bergeron, Wenyin Fu, Shengbao Zheng, Brian Coutinho, Saeed Rashidi, Changhai Man, et al. 2023. Chakra: Advancing performance benchmarking and co-design using standardized execution traces.arXiv preprint arXiv:23...

  85. [92]

    Andrew Tee, Nicholas Curtis, Noah Wolfe, and Daniel Wong. 2025. The MALL is Open: Exploring Shared Caches and Latency in AMD CDNA™3 GPUs. In Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. 1110–1116

  86. [93]

    Ajay Tirumala and Raymond Wong. 2024. Nvidia blackwell platform: Advancing generative ai and accelerated computing. In2024 IEEE Hot Chips 36 Symposium (HCS). IEEE Computer Society, 1–33

  87. [94]

    Walker J Turner, John W Poulton, John M Wilson, Xi Chen, Stephen G Tell, Matthew Fojtik, Thomas H Greer, Brian Zimmer, Sanquan Song, Nikola Nedovic, et al. 2018. Ground-referenced signaling for intra-chip and short-reach chip-to- chip interconnects. In2018 IEEE Custom Integrat...

  88. [95]

    Ben Verghese, Scott Devine, Anoop Gupta, and Mendel Rosenblum. 1996. Oper- ating system support for improving data locality on CC-NUMA compute servers. SIGPLAN Not.31, 9 (Sept. 1996), 279–289. https://doi.org/10.1145/248209.237205

  89. [96]

    Nandita Vijaykumar, Eiman Ebrahimi, Kevin Hsieh, Phillip B Gibbons, and Onur Mutlu. 2018. The locality descriptor: A holistic cross-layer abstraction to express data locality in GPUs. In2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). IEEE, 829–842

  90. [97]

    Xizheng Wang, Qingxu Li, Yichi Xu, Gang Lu, Dan Li, Li Chen, Heyang Zhou, Linkang Zheng, Sen Zhang, Yikai Zhu, Yang Liu, Pengcheng Zhang, Kun Qian, Kunling He, Jiaqi Gao, Ennan Zhai, Dennis Cai, and Binzhang Fu. 2025. SimAI: unifying architecture design and performance tuning ...

  91. [99]

    Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al. 2022. Sustainable ai: Environmental implications, challenges and opportunities.Pro- ceedings of machine learning and systems4 (20...

  92. [100]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025). 14

  93. [101]

    Jie Amy Yang, Jongsoo Park, Srinivas Sridharan, and Ping Tak Peter Tang

  94. [102]

    Jialiang Zhang, Michael Swift, and Jing Li. 2022. Software-defined address mapping: a case on 3d memory. InProceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. 70–83

  95. [103]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068 (2022)

  96. [104]

    Xia Zhao, Magnus Jahre, Yuhua Tang, Guangda Zhang, and Lieven Eeckhout

  97. [105]

    Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. 2024. HeteGen: Efficient Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices. InProceedings of Ma- chine Learning and Systems, P. Gibbons, G. Pekhimenko, and C...

  98. [106]

    Gonza- lez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonza- lez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. InAdvances in Neural Infor...

  99. [107]

    InConference on Knowledge Discovery and Data Mining (KDD), Vol

    Training deep learning recommendation model with quantized collective communications. InConference on Knowledge Discovery and Data Mining (KDD), Vol. 95

  100. [108]

    Shizhuo Zhu, Illia Shkirko, Jacob Levinson, Zhengrong Wang, and Tony Nowatzki. 2024. SPGPU: Spatially Programmed GPU.IEEE Computer Ar- chitecture Letters(2024). 15

  101. [114]

    Tianhao Zheng, David Nellans, Arslan Zulfiqar, Mark Stephenson, and Stephen W Keckler. 2016. Towards high performance paged memory for GPUs. In2016 IEEE International Symposium on High Performance Computer Architec- ture (HPCA). IEEE, 345–357

  102. [2009]

    InProceedings of the 36th annual international symposium on Computer architecture

    Reactive NUCA: near-optimal block placement and replication in dis- tributed caches. InProceedings of the 36th annual international symposium on Computer architecture. 184–195

  103. [2020]

    InProceedings of the 2020 ACM SIGMOD International Con- ference on Management of Data(Portland, OR, USA)(SIGMOD ’20)

    Pump Up the Volume: Processing Large Data on GPUs with Fast Interconnects. InProceedings of the 2020 ACM SIGMOD International Con- ference on Management of Data(Portland, OR, USA)(SIGMOD ’20). Associ- ation for Computing Machinery, New York, NY, USA, 1633–1649. https: //doi.or...

  104. [2022]

    The carbon footprint of machine learning training will plateau, then shrink.Computer55, 7 (2022), 18–28

  105. [2023]

    InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2

    NUBA: Non-uniform bandwidth GPUs. InProceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2. 544–559

  106. [2024]

    In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS)

    Benchmarking and dissecting the nvidia hopper gpu architecture. In2024 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 656–667

  107. [2025]

    In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles

    LithOS: An operating system for efficient machine learning on GPUs. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles. 1–17

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.