REVIEW 4 major objections 5 minor 86 references
Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper claims that inter-core connected NPUs can be fully virtualized into arbitrary virtual topologies with near-zero overhead.
desk verdict A genuinely new virtualization design for inter-core connected NPUs, with a solid prototype and sane overhead numbers; the headline MIG comparison is real but mostly reflects partition granularity, not superior translation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the virtual topology, backed by two hardware tables. The routing table lives at the NPU controller and in each core's NoC engine: it maps every virtual core ID to a physical core ID and, for irregular shapes, stores a direction for each relay hop so network packets stay inside the virtual boundary. The Range Translation Table (RTT) maps virtual address ranges to physical ranges and, instead of a costly search, advances a current-entry pointer RTT_CUR within an iteration and stores a last_v hint that jumps back to the right entry at the start of the next iteration, matching the paper's three observed memory access patterns. Around these tables, a minimum topology-edit-distance search (with connectedness and deduplication pruning) turns a requested virtual shape into a best-effort physical allocation.
What would settle it
Instrument the vNPU prototype or simulator with a graph neural network workload whose memory access is random at small granularity, and count range-TLB misses per DMA transfer. If the average translation overhead approaches the 9.2% the paper reports for a 32-entry page TLB rather than the under-4.3% claimed for vChunk, then the central performance claim holds only for the three access patterns, not for NPU workloads in general.
Extended reading notes
Core claim
The paper's central claim is that an inter-core connected NPU can be fully virtualized by virtualizing three things at once: the instruction stream, the network-on-chip, and the special SRAM-centric memory. vNPU adds a vRouter that uses a routing table to rewrite virtual core IDs to physical core IDs, both for instruction dispatch and for NoC packets, with direction fields so packets never cross into another virtual NPU's territory. It adds vChunk, which replaces fixed-size page translation with range-based translation entries and uses per-core pointers to exploit the monotonic, repeated, tensor-granular address streams typical of weight loading, cutting translation overhead to under 4.3 percent with only four hardware entries. It also adds a best-effort topology allocation algorithm based on minimum topology-edit distance, which chooses a connected physical core set that resembles the requested shape when an exact match is impossible. The evaluation claims up to 1.92x and 1.28x speedups over MIG-based partitions for Transformer and ResNet, and under 1% overhead relative to bare metal.
Load-bearing premise
vChunk's speedup relies on the assumption that NPU global-memory traffic comes in large chunk-sized transfers whose addresses increase monotonically within an iteration and repeat across iterations; a workload like a graph neural network that accesses memory in small random pieces would miss the range cache and lose most of the benefit, a case the paper itself points to when recommending page-level translation.
Editorial extensions
If this is right
- A cloud provider can run many small models concurrently on one large inter-core NPU, allocating exactly the core count each model needs instead of fixed slices.
- Dataflow-style models such as Transformers gain most from virtual topology routing, up to 1.92x over MIG, because intermediate results stay on the NoC instead of round-tripping through global memory.
- The hardware cost of adding virtualization, about 2% of resources and under 1% end-to-end slowdown, is small enough that the feature could plausibly ship in production NPUs.
- Because page-level translation remains available for irregular workloads, a hybrid design could switch between range-based and page-based translation depending on the access pattern.
Reading between the lines
- The range-translation mechanism is not tied to inter-core NPUs; any SRAM-centric accelerator that loads data by bulk DMA could adopt vChunk-style translation to avoid TLB pressure.
- The same topology-edit-distance idea could be applied dynamically: a hypervisor could reshape a virtual NPU's topology between inference phases as traffic patterns change, which the paper leaves unexplored.
- If the three memory patterns hold for streaming LLM workloads, vChunk could also simplify dynamic KV-cache management and offloading, a future-work item the paper acknowledges.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents vNPU, a proposed virtualization architecture for inter-core connected neural processing units, comprising three components: vRouter for virtualizing instruction dispatch and NoC data flows, vChunk for range-based memory virtualization that leverages observed NPU memory access patterns, and a best-effort topology mapping strategy based on topology edit distance. The authors report implementations on an FPGA prototype (Chipyard/FireSim) and a cycle-exact simulator (DCRA), with microbenchmarks and end-to-end ML workload evaluations. Their headline claims are performance improvements of up to 1.92x and 1.28x over a MIG-based virtual NPU for Transformer and ResNet workloads, respectively, a 2.29x improvement over a UVM-based virtual NPU for Transformer, and hardware overhead of about 2% in LUTs/FFs with less than 1% end-to-end performance degradation.
Significance. If the results hold, vNPU would be a useful step toward practical multi-tenancy on inter-core connected NPUs by enabling flexible virtual topologies, a capability absent from prior NPU virtualization work such as Aurora and V10. The paper's strengths include an actual RTL-level prototype, cycle-exact simulation, explicit recognition of workload-dependent limitations in Section 7, and microbenchmarks that attempt to isolate the overhead of routing and memory translation. However, the significance is tempered by the fact that the strongest headline numbers depend on a MIG configuration that cannot allocate the requested core count, and on workload access-pattern assumptions that the authors themselves concede do not cover graph workloads.
major comments (4)
- [§6.3.2 and Abstract] The headline 1.92x improvement over MIG is not a direct measure of vNPU's virtualization efficiency. As the text explains, GPT-large requires 36 cores while the MIG configuration can provision at most 24, so MIG must time-multiplex a single physical core among multiple virtual cores. The reported speedup therefore largely reflects vNPU's ability to allocate exactly 36 cores rather than lower translation overhead or better virtual topology routing. To support the central claim, the paper needs a sensitivity analysis varying both the workload core-count requirement and the set of MIG partition sizes, including a configuration in which MIG can also allocate the exact requested core count; without this, the abstract's 'up to 1.92x' claim is overstated.
- [§4.2 and §7] The vChunk performance benefits are load-bearing on the three memory access patterns stated in Section 4.2: tensor-granular transfers, monotonic address increase within an iteration, and repeated address sets across iterations. The paper itself states in Section 7 that for graph workloads such as GNNs, which involve random information retrieval, page-level translation is recommended. This means the demonstrated memory-virtualization improvement is conditional on a workload class; the paper should either quantify, using its own DCRA traces, the fraction of representative workloads that conform to Patterns 1-3, or explicitly scope the performance claims to such workloads. Currently, the central claim 'vChunk maximizes memory bandwidth' is not established for workloads outside the three patterns.
- [§6.3.1] The 2.29x improvement over the UVM-based virtual NPU is acknowledged as 'somewhat unfair' because UVM lacks inter-core connection support altogether. While the authors state their intent to highlight the importance of virtual topology, the evaluation does not isolate the contribution of vRouter from the inherent architectural benefit of inter-core communication versus memory synchronization. A cleaner comparison would hold the physical NPU architecture fixed and toggle only the virtualization features (vRouter/vChunk) on and off, as is done for the bare-metal comparison in §6.3.3; reporting such an ablation would make the central claim more convincing.
- [§4.3 and §6.3.5] The topology mapping algorithm computes minimum topology edit distance over candidate subgraphs, and Section 4.3 states the underlying problem is NP-hard. Algorithm 1 prunes candidates with connectivity and deduplication checks, but the paper provides no empirical or analytical evaluation of the algorithm's runtime for larger chips (e.g., IPU-scale with thousands of cores). Since topology mapping is invoked during virtual NPU creation, the lack of scalability data leaves open whether the proposed best-effort mapping can be applied at realistic NPU sizes; the current evaluation at 9-28 cores is insufficient to establish this.
minor comments (5)
- [Abstract vs. §6.3.2] The abstract states improvements 'of up to 1.92x and 1.28x for the Transformer and ResNet models', but the abstract also later says 'up to a 2x performance improvement across various ML models'; using 1.92x and 1.28x consistently would avoid confusion.
- [Figure 16] The upper half of Figure 16 uses small diagrams to illustrate TDM and exact allocation, but the correspondence between the diagrams and the bar/dot data in the lower half is not explicitly labeled; a legend or sub-caption clarifying which bars correspond to which MIG partition configuration would improve readability.
- [Figure 13] The x-axis labels such as 'comp 1-1' and 'comp 1-4' are cryptic; spelling out 'computation time' and 'sender-to-receiver ratio' as axis annotations or a caption would make the figure self-contained.
- [Table 3] The microbenchmark in Table 3 reports a 1%-2% overhead for vRouter, but no error bars or run-to-run variance are provided; given the small absolute overheads (342 vs. 309 clocks for Send with 2 packets), confirming statistical stability with multiple runs would strengthen the claim.
- [§6.4] The hardware cost analysis compares vNPU against Kim's solution, but the paper does not break out the cost of vChunk versus vRouter separately on the NPU core; reporting these components individually would help readers understand where the 2% LUT/FF overhead comes from.
Circularity Check
No significant circularity: vNPU's results are empirical measurements; the vChunk access-pattern assumption is honestly scoped, and the only same-author citation is not load-bearing.
full rationale
vNPU does not derive quantitative outcomes from fitted parameters. The headline claim of up to 1.92x/1.28x improvement over MIG is a measured comparison on FireSim/DCRA, and the paper explicitly attributes the 1.92x figure to MIG's 24-core provisioning ceiling versus GPT-large's 36-core demand (Section 6.3.2), which is a baseline-granularity artifact, not a circular reduction. The vChunk mechanism is motivated by observed NPU memory access patterns (Section 4.2), and Section 7 candidly restricts its scope by recommending page-level translation for random-access graph workloads such as GNNs; this is an honest scoping of a workload-driven mechanism rather than a self-fulfilling prediction. The only same-author citation, [20] (sNPU), is used together with the external NeuMMU [36] citation to motivate coarse-grain DMA versus page-level translation, and the paper independently confirms the translation overhead in its own Figure 14 measurements, so the self-citation is not load-bearing. No uniqueness theorem or ansatz is imported from prior author work; the topology mapping relies on standard graph edit distance with explicitly stated pruning and cost heuristics. Overall, the paper is self-consistent empirical systems work with no significant circularity.
Assumptions & free parameters
free parameters (1)
- Topology edit distance penalty weights (NodeMatch, EdgeMatch costs) =
User-specified; no default or sensitivity analysis
assumptions (4)
- domain assumption The physical NPU uses a regular 2D mesh NoC with dimension-order routing, and any connected subgraph can be made a virtual topology with direction hints in the routing table.
- domain assumption NPU global-memory access is characterized by Patterns 1-3: tensor-granular weight loads, monotonic address increase within an iteration, and repeated addresses across iterations.
- domain assumption A hyper-mode NPU controller can be trusted to enforce isolation by restricting writes to routing tables and range translation tables.
- domain assumption Memory bandwidth can be throttled per virtual NPU via access counters in each core.
invented entities (3)
-
vRouter module (NPU controller instruction router plus NoC engine rewrite)
-
vChunk range translation table (RTT) with last_v/RTT_CUR indexing
-
Routing table (RT) with standard and 2D-mesh compressed encodings
Cite this review
Pith. "Pith review of Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units." pith.science (2026). https://pith.science/paper/KBDNG5F2
@misc{pith2026250611446,
author = {Pith},
title = {Pith review of: Topology-Aware Virtualization over Inter-Core Connected Neural Processing Units},
year = {2026},
howpublished = {\url{https://pith.science/paper/KBDNG5F2}},
note = {Machine review of arXiv:2506.11446}
}
read the original abstract
With the rapid development of artificial intelligence (AI) applications, an emerging class of AI accelerators, termed Inter-core Connected Neural Processing Units (NPU), has been adopted in both cloud and edge computing environments, like Graphcore IPU, Tenstorrent, etc. Despite their innovative design, these NPUs often demand substantial hardware resources, leading to suboptimal resource utilization due to the imbalance of hardware requirements across various tasks. To address this issue, prior research has explored virtualization techniques for monolithic NPUs, but has neglected inter-core connected NPUs with the hardware topology. This paper introduces vNPU, the first comprehensive virtualization design for inter-core connected NPUs, integrating three novel techniques: (1) NPU route virtualization, which redirects instruction and data flow from virtual NPU cores to physical ones, creating a virtual topology; (2) NPU memory virtualization, designed to minimize translation stalls for SRAM-centric and NoC-equipped NPU cores, thereby maximizing the memory bandwidth; and (3) Best-effort topology mapping, which determines the optimal mapping from all candidate virtual topologies, balancing resource utilization with end-to-end performance. We have developed a prototype of vNPU on both an FPGA platform (Chipyard+FireSim) and a simulator (DCRA). Evaluation results indicate that, compared to other virtualization approaches such as unified virtual memory and MIG, vNPU achieves up to a 2x performance improvement across various ML models, with only 2% hardware cost.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Dennis Abts, Garrin Kimmell, Andrew Ling, John Kim, Matt Boyd, Andrew Bitar, Sahil Parmar, Ibrahim Ahmed, Roberto DiCecco, David Han, John Thompson, Michael Bye, Jennifer Hwang, Jeremy Fowers, Peter Lillian, Ashwin Murthy, Elyas Mehtabuddin, Chetan Tekur, Thomas Sohmers, Kris Kang, Stephen Maresh, and Jonathan Ross. 2022. A software-defined tensor streami...
arXiv 2022
-
[2]
Alibaba. 2024. Qwen/Qwen2-0.5B. https://huggingface.co/Qwen/Qwen2-0.5B. Referenced January 2024
2024
-
[3]
Alon Amid, David Biancolin, Abraham Gonzalez, Daniel Grubb, Sagar Karandikar, Harrison Liew, Albert Magyar, Howard Mao, Albert Ou, Nathan Pemberton, Paul Rigge, Colin Schmidt, John Wright, Jerry Zhao, Yakun Sophia Shao, Krste Asanović, and Borivoje Nikolić. 2020. Chipyard: Integrated Design, Simulation, and Implementation Framework for Custom SoCs. IEEE M...
arXiv 2020
-
[4]
Ardalan Amiri Sani, Kevin Boos, Shaopu Qin, and Lin Zhong. 2014. I/O paravir- tualization at the device file boundary. ACM SIGARCH Computer Architecture News 42, 1 (2014), 319–332
2014
-
[5]
Ardalan Amiri Sani, Kevin Boos, Min Hong Yun, and Lin Zhong. 2014. Rio: a system solution for sharing i/o between mobile systems. In Proceedings of the 12th annual international conference on Mobile systems, applications, and services . 259–272
2014
-
[6]
AWS. 2024. GNeuronCore-v2 Architecture. https://awsdocs-neuron.readthedocs- hosted.com/en/latest/general/arch/neuron-hardware/neuron-core-v2.html. Ref- erenced January 2024
2024
-
[7]
AWS. 2024. Use Amazon SageMaker Built-in Algorithms or Pre-trained Mod- els. https://docs.aws.amazon.com/sagemaker/latest/dg/algos.html. Referenced January 2024
2024
-
[8]
Amnon Barak, Tal Ben-Nun, Ely Levy, and Amnon Shiloh. 2010. A package for OpenCL based heterogeneous computing on clusters with many GPU devices. In 2010 IEEE international conference on cluster computing workshops and posters (CLUSTER WORKSHOPS). IEEE, 1–7
2010
Show all 86 references
-
[9]
Paul Barham, Boris Dragovic, Keir Fraser, Steven Hand, Tim Harris, Alex Ho, Rolf Neugebauer, Ian Pratt, and Andrew Warfield. 2003. Xen and the art of virtualization. ACM SIGOPS operating systems review 37, 5 (2003), 164–177
2003
-
[10]
Dongwei Chen, Dong Tong, Chun Yang, Jiangfang Yi, and Xu Cheng. 2023. FlexPointer: Fast Address Translation Based on Range TLB and Tagged Pointers. ACM Trans. Archit. Code Optim. 20, 2, Article 30 (March 2023), 24 pages. https: //doi.org/10.1145/3579854
2023 doi
-
[11]
Quan Chen, Hailong Yang, Minyi Guo, Ram Srivatsa Kannan, Jason Mars, and Lingjia Tang. 2017. Prophet: Precise qos prediction on non-preemptive accelera- tors to improve utilization in warehouse-scale computers. In Proceedings of the Twenty-Second International Conference on Ar...
2017
-
[12]
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv:1512.01274 [cs.DC]
2015 arXiv
-
[13]
Alibaba Cloud. 2024. High GPU Utilization with cGPU. https://www.alibabacloud. com/en/solutions/cgpu?_p_lc=1. Referenced January 2024
2024
-
[14]
CNCF. 2024. HAMi: Heterogeneous AI Computing Virtualization Middleware. https://github.com/Project-HAMi/HAMi. Referenced November 2024
2024
-
[15]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[16]
Micah Dowty and Jeremy Sugerman. 2008. GPU virtualization on VMware’s hosted I/O architecture. ACM SIGOPS Oper. Syst. Rev. 43 (2008), 73–82. https: //api.semanticscholar.org/CorpusID:228328
2008
-
[17]
José Duato, Francisco D Igual, Rafael Mayo, Antonio J Pena, Enrique S Quintana- Ortí, and Federico Silla. 2010. An efficient implementation of GPU virtualization in high performance clusters. In Euro-Par 2009–Parallel Processing Workshops: HPPC, HeteroPar, PROPER, ROIA, UNICOR...
2010
-
[18]
Jose Duato, Antonio J Pena, Federico Silla, Juan C Fernandez, Rafael Mayo, and Enrique S Quintana-Orti. 2011. Enabling CUDA acceleration within virtual machines using rCUDA. In2011 18th International Conference on High Performance Computing. IEEE, 1–10
2011
-
[19]
Peña, Federico Silla, Rafael Mayo, and Enrique S
José Duato, Antonio J. Peña, Federico Silla, Rafael Mayo, and Enrique S. Quintana- Ortí. 2010. rCUDA: Reducing the number of GPU-based accelerators in high performance clusters. In 2010 International Conference on High Performance Com- puting and Simulation. 224–231. https://d...
2010
-
[20]
Erhu Feng, Dahu Feng, Dong Du, Yubin Xia, and Haibo Chen. 2024. sNPU: Trusted Execution Environments on Integrated NPUs. (2024)
2024
-
[21]
Amin Firoozshahian, Joel Coburn, Roman Levenstein, Rakesh Nattoji, Ashwin Kamath, Olivia Wu, Gurdeepak Grewal, Harish Aepala, Bhasker Jakka, and Bob Dreyer. 2023. Mtia: First generation silicon targeting meta’s recommendation systems. In Proceedings of the 50th Annual Internat...
2023
-
[22]
Hill, Kathryn S
Jayneel Gandhi, Vasileios Karakostas, Furkan Ayar, Adrián Cristal, Mark D. Hill, Kathryn S. McKinley, Mario Nemirovsky, Michael M. Swift, and Osman S. Ünsal
-
[23]
Hasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali, Vighnesh Iyer, Pranav Prakash, Jerry Zhao, Daniel Grubb, Harrison Liew, Howard Mao, Albert Ou, Colin Schmidt, Samuel Steffl, John Wright, Ion Stoica, Jonathan Ragan-Kelley, Krste Asanovic, Borivoje Nikolic, and Yakun Sophia Shao....
2021
-
[24]
Giulio Giunta, Raffaele Montella, Giuseppe Agrillo, and Giuseppe Coviello. 2010. A GPGPU transparent virtualization component for high performance computing clouds. In Euro-Par 2010-Parallel Processing: 16th International Euro-Par Conference, Ischia, Italy, August 31-September...
2010
-
[25]
Google. 2024. Extract insights from images, documents, and videos. https: //cloud.google.com/vision. Referenced January 2024
2024
-
[26]
Google. 2024. TPU/TPU v6e. https://cloud.google.com/tpu/docs/v6e. Referenced February 2025
2024
-
[27]
Graphcore. 2024. Intelligence processing unit. https://www.graphcore.ai/ products/ipu. Referenced January 2024
2024
-
[28]
Graphcore. 2024. IPU/ipu-programmers-guide. https://docs.graphcore.ai/ projects/ipu-programmers-guide/en/latest/programming_tools.html. Referenced February 2025
2024
-
[29]
Vishakha Gupta, Ada Gavrilovska, Karsten Schwan, Harshvardhan Kharche, Niraj Tolia, Vanish Talwar, and Parthasarathy Ranganathan. 2009. GViM: GPU- accelerated virtual machines. In Proceedings of the 3rd ACM Workshop on System- level virtualization for High Performance Computin...
2009
-
[30]
Vishakha Gupta, Karsten Schwan, Niraj Tolia, Vanish Talwar, and Parthasarathy Ranganathan. 2011. Pegasus: Coordinated scheduling for virtualized accelerator- based systems. In 2011 USENIX Annual Technical Conference (USENIX ATC 11)
2011
-
[31]
Peter E Hart, Nils J Nilsson, and Bertram Raphael. 1968. A formal basis for the heuristic determination of minimum cost paths. IEEE transactions on Systems Science and Cybernetics 4, 2 (1968), 100–107
1968
-
[32]
Zhang, Shaoqing Ren, and Jian Sun
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015), 770–778. https://api.semanticscholar.org/CorpusID: 206594692
2015
-
[33]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[34]
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)
2017 arXiv
-
[35]
Rachel Huang, Jonathan Pedoeem, and Cuixian Chen. 2018. YOLO-LITE: a real- time object detection algorithm optimized for non-GPU computers. In 2018 IEEE international conference on big data (big data) . IEEE, 2503–2510
2018
-
[36]
Bongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim, and Minsoo Rhu
-
[37]
Intel. 2023. Intel Virtualization Technology for Directed I/O Architecture Specifi- cation. https://cdrdv2-public.intel.com/671081/vt-directed-io-spec.pdf. Refer- enced April 2023
2023
-
[38]
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazi...
2017
-
[39]
Sagar Karandikar, Howard Mao, Donggyu Kim, David Biancolin, Alon Amid, Dayeol Lee, Nathan Pemberton, Emmanuel Amaro, Colin Schmidt, Aditya Chopra, Qijing Huang, Kyle Kovacs, Borivoje Nikolic, Randy Katz, Jonathan Bachrach, and Krste Asanović. 2018. FireSim: FPGA-accelerated Cy...
2018
-
[40]
Jungwon Kim, Sangmin Seo, Jun Lee, Jeongho Nah, Gangwon Jo, and Jaejin Lee
-
[41]
Seah Kim, Jerry Zhao, Krste Asanović, Borivoje Nikolić, and Yakun Sophia Shao
-
[42]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classifi- cation with deep convolutional neural networks. Advances in neural information processing systems 25 (2012)
2012
-
[43]
Yossi Kuperman, Eyal Moscovici, Joel Nider, Razya Ladelsky, Abel Gordon, and Dan Tsafrir. 2016. Paravirtual remote i/o. ACM SIGARCH Computer Architecture News 44, 2 (2016), 49–65
2016
-
[44]
K. Li, H. Chen, J. Sun, and L. Shi. 2012. vCUDA: GPU-Accelerated High- Performance Computing in Virtual Machines. IEEE Trans. Comput. 61, 06 (jun 2012), 804–816. https://doi.org/10.1109/TC.2011.112
2012 doi
-
[45]
Teng Li, Vikram K Narayana, Esam El-Araby, and Tarek El-Ghazawi. 2011. GPU resource sharing and virtualization on high performance computing systems. In 2011 International Conference on Parallel Processing . IEEE, 733–742
2011
-
[46]
Tyng-Yeu Liang and Yu-Wei Chang. 2011. GridCuda: a grid-enabled CUDA programming toolkit. In 2011 IEEE Workshops of International Conference on Advanced Information Networking and Applications . IEEE, 141–146
2011
-
[47]
Sean Lie. 2023. Cerebras architecture deep dive: First look inside the hardware/- software co-design for deep learning. IEEE Micro 43, 3 (2023), 18–30
2023
-
[48]
Zhen Lin, Lars Nyland, and Huiyang Zhou. 2016. Enabling efficient preemption for SIMT architectures with lightweight context switching. In SC’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 898–908
2016
-
[49]
Yiqi Liu, Yuqi Xue, Yu Cheng, Lingxiao Ma, Ziming Miao, Jilong Xue, and Jian Huang. 2024. Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor. arXiv preprint arXiv:2408.04808 (2024)
2024 arXiv
-
[50]
morenes. 2024. DCRA/DCRA ae. https://github.com/morenes/dcra. Referenced February 2025
2024
-
[51]
Michel Neuhaus, Kaspar Riesen, and Horst Bunke. 2006. Fast suboptimal algo- rithms for the computation of graph edit distance. In Structural, Syntactic, and Statistical Pattern Recognition: Joint IAPR International Workshops, SSPR 2006 and SPR 2006, Hong Kong, China, August 17...
2006
-
[52]
NVIDIA. 2023. NVIDIA Multi-Instance GPU. https://www.nvidia.com/en-us/ technologies/multi-instance-gpu/. Referenced April 2023
2023
-
[53]
NVIDIA. 2024. MULTI-PROCESS SERVICE. https://docs.nvidia.com/deploy/pdf/ CUDA_Multi_Process_Service_Overview.pdf. Referenced January 2024
2024
-
[54]
NVIDIA. 2024. Unlock Next Level Performance with Virtual GPUs. https://www. nvidia.com/en-us/data-center/virtual-solutions/. Referenced January 2024
2024
-
[55]
OpenAI. 2024. Introducing ChatGPT. https://openai.com/index/chatgpt/. Refer- enced January 2024
2024
-
[56]
Marcelo Orenes-Vera, Esin Tureci, Margaret Martonosi, and David Wentzlaff. 2024. DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications. arXiv:2311.15443 [cs.AR] https://arxiv.org/abs/2311.15443
2024 arXiv
-
[57]
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024. Splitwise: Efficient generative LLM inference using phase splitting. arXiv:2311.18677 [cs.AR]
2024 arXiv
-
[58]
Raghu Prabhakar, Sumti Jairath, and Jinuk Luke Shin. 2022. Sambanova sn10 rdu: A 7nm dataflow architecture to accelerate software 2.0. In2022 IEEE International Solid-State Circuits Conference (ISSCC) , Vol. 65. IEEE, 350–352
2022
-
[59]
Carlos Reaño, Antonio J Peña, Federico Silla, José Duato, Rafael Mayo, and Enrique S Quintana-Ortí. 2012. CU2rCU: Towards the complete rCUDA remote GPU virtualization and sharing solution. In 2012 19th International Conference on High Performance Computing. IEEE, 1–10
2012
-
[60]
Kaspar Riesen and Horst Bunke. 2009. Approximate graph edit distance compu- tation by means of bipartite graph matching. Image and Vision computing 27, 7 (2009), 950–959
2009
-
[61]
Kaspar Riesen, Stefan Fankhauser, and Horst Bunke. 2007. Speeding up graph edit distance computation with a bipartite heuristic.. In MLG. Citeseer, 21–24
2007
-
[62]
Tell, Yanqing Zhang, William J
Yakun Sophia Shao, Jason Clemons, Rangharajan Venkatesan, Brian Zimmer, Matthew Fojtik, Nan Jiang, Ben Keller, Alicia Klinefelter, Nathaniel Pinckney, Priyanka Raina, Stephen G. Tell, Yanqing Zhang, William J. Dally, Joel Emer, C. Thomas Gray, Brucek Khailany, and Stephen W. K...
2019
-
[63]
Yusuke Suzuki, Shinpei Kato, Hiroshi Yamada, and Kenji Kono. 2016. GPUvm: GPU Virtualization at the Hypervisor. IEEE Trans. Comput. 65 (2016), 2752–2766. https://api.semanticscholar.org/CorpusID:6941728
2016
-
[64]
Michael M Swift, Brian N Bershad, and Henry M Levy. 2003. Improving the reliability of commodity operating systems. In Proceedings of the nineteenth ACM symposium on Operating systems principles . 207–222
2003
-
[65]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2014. Going Deeper with Convolutions. arXiv:1409.4842 [cs.CV]
2014 arXiv
-
[66]
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015. Going deeper with convolutions. InProceedings of the IEEE conference on computer vision and pattern recognition . 1–9
2015
-
[67]
Emil Talpes, Debjit Das Sarma, Ganesh Venkataramanan, Peter Bannon, Bill McGee, Benjamin Floering, Ankit Jalote, Christopher Hsiong, Sahil Arora, Atchyuth Gorti, and Gagandeep S. Sachdev. 2020. Compute Solution for Tesla’s Full Self-Driving Computer. IEEE Micro 40, 2 (2020), 2...
2020
-
[68]
Emil Talpes, Douglas Williams, and Debjit Das Sarma. 2022. Dojo: The microar- chitecture of tesla’s exa-scale computer. In 2022 IEEE Hot Chips 34 Symposium (HCS). IEEE Computer Society, 1–28
2022
-
[69]
tenstorrent. 2024. Tenstorrent - Scalable and Efficient Hardware for Deep Learn- ing. https://tenstorrent.com/. Referenced January 2024
2024
-
[70]
Cowperthwaite
Kun Tian, Yaozu Dong, and David J. Cowperthwaite. 2014. A Full GPU Virtu- alization Solution with Mediated Pass-Through. In USENIX Annual Technical Conference. https://api.semanticscholar.org/CorpusID:12658735
2014
-
[71]
Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M
Hugo Touvron, Louis Martin, Kevin R. Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Daniel M. Bikel, Lukas Blecher, Cristian Cantón Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jer...
2023 arXiv
-
[72]
Dimitrios Vasilas, Stefanos Gerangelos, and Nectarios Koziris. 2016. VGVM: Efficient GPU capabilities in virtual machines. In 2016 International Conference on High Performance Computing & Simulation (HPCS) . IEEE, 637–644
2016
-
[73]
Gibbons, and Onur Mutlu
Nandita Vijaykumar, Kevin Hsieh, Gennady Pekhimenko, Samira Manabi Khan, Ashish Shrestha, Saugata Ghose, Adwait Jog, Phillip B. Gibbons, and Onur Mutlu
-
[74]
Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration. arXiv:2406.01014 [cs.CL] https://arxiv.org/abs/2406.01014
2024 arXiv
-
[75]
2017.{vCorfu}: A{Cloud-Scale} Object Store on a Shared Log
Michael Wei, Amy Tai, Christopher J Rossbach, Ittai Abraham, Maithem Munshed, Medhavi Dhawan, Jim Stabile, Udi Wieder, Scott Fritchie, and Steven Swanson. 2017.{vCorfu}: A{Cloud-Scale} Object Store on a Shared Log. In 14th USENIX Symposium on Networked Systems Design and Imple...
2017
-
[76]
Shucai Xiao, Pavan Balaji, James Dinan, Qian Zhu, Rajeev Thakur, Susan Coghlan, Heshan Lin, Gaojin Wen, Jue Hong, and Wu-chun Feng. 2012. Transparent accelerator migration in a virtualized GPU environment. In 2012 12th IEEE/ACM International Symposium on Cluster, Cloud and Gri...
2012
-
[77]
Yuqi Xue, Yiqi Liu, and Jian Huang. 2023. System Virtualization for Neural Processing Units. In Proceedings of the 19th Workshop on Hot Topics in Oper- ating Systems (Providence, RI, USA) (HOTOS ’23). Association for Computing Machinery, New York, NY, USA, 80–86. https://doi.o...
2023
-
[78]
2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) (2016), 1–14
Zorua: A holistic approach to resource virtualization in GPUs. 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) (2016), 1–14. https://api.semanticscholar.org/CorpusID:2310493
2016
-
[79]
Tsung Tai Yeh, Amit Sabne, Putt Sakdhnagool, Rudolf Eigenmann, and Timothy G. Rogers. 2017. Pagoda: Fine-Grained GPU Resource Virtualization for Narrow ISCA ’25, June 21–25, 2025, Tokyo, Japan Feng et al. Tasks. Proceedings of the 22nd ACM SIGPLAN Symposium on Principles and P...
2017
-
[80]
Hangchen Yu, Arthur Michener Peters, Amogh Akshintala, and Christopher J Rossbach. 2020. AvA: Accelerated virtualization of accelerators. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Program- ming Languages and Operating Systems . 807–825
2020
-
[81]
Kai Zhang, Bingsheng He, Jiayu Hu, Ze ke Wang, Bei Hua, Jiayi Meng, and Lishan Yang. 2018. G-NET: Effective GPU Sharing in NFV Systems. In Symposium on Networked Systems Design and Implementation . https://api.semanticscholar.org/ CorpusID:4493567
2018
-
[83]
Yuqi Xue, Yiqi Liu, Lifeng Nai, and Jian Huang. 2023. V10: Hardware-Assisted NPU Multi-tenancy for Improved Resource Utilization and Fairness. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA ’23). Association for Comp...
2023
-
[2012]
In Proceedings of the 26th ACM international conference on Supercomputing
SnuCL: an OpenCL framework for heterogeneous CPU/GPU clusters. In Proceedings of the 26th ACM international conference on Supercomputing. 341–352
-
[2016]
IEEE Micro 36, 3 (2016), 118–126
Range Translations for Fast Virtual Memory. IEEE Micro 36, 3 (2016), 118–126. https://doi.org/10.1109/MM.2016.10
2016 doi
-
[2019]
Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems(2019)
NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing Units. Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems(2019). https://api.semanticscholar.org/CorpusID:208139570
2019
-
[2023]
2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) (2023)
AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant Work- loads. 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) (2023)
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.