REVIEW 5 major objections 7 minor 4 cited by
RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing
T0 review · 5 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read RRTO makes transparent offloading fast by recording a model's fixed operator sequence and replaying it on the GPU server, cutting per-inference RPCs from 5,895 to 11.
desk verdict RRTO's record/replay idea is a genuine step forward for transparent offloading, and the measurements largely back it up; the unstated scope on overlapping inference is the main caveat. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Operator Sequence Search algorithm, which reconstructs the exact per-inference operator sequence from raw interception logs. It relies on three observations: the same operator sequence repeats across inferences; each inference is bracketed by a single host-to-device memory copy at the start and a single device-to-host copy at the end; and every operator's inputs must satisfy data-dependency constraints against the raw input, prior operator outputs, or model parameters. To keep the search tractable on logs with tens of thousands of entries, the algorithm first runs FastCheck on compact category tags to prune candidates by repeated occurrence, then runs FullCheck on the survivors to realign start/end markers, verify data dependencies, and confirm exact record-level repetition across the whole log.
What would settle it
Run RRTO under an execution engine that uses pre-pinned asynchronous DMA, persistent CUDA graphs, stream pooling, or two concurrently interleaved inference streams; if Operator Sequence Search then fails to find a valid repeated sequence and the system falls back to per-operator RPCs, or if replay produces wrong outputs, the central speedup claim is falsified for that setting.
Extended reading notes
Core claim
The paper's central discovery is that the per-operator RPC cost that dominates transparent offloading is not inherent to transparency but an artifact of reacting to operator calls as they arrive. For models whose operator sequence is fixed, the sequence can be discovered once and then replayed: the offloading client intercepts CUDA calls, records them during an initial recording phase, searches the accumulated log for the repeating inference pattern, and thereafter returns cached status results to the application while the server executes the whole recorded sequence in one shot. The paper argues this is the first demonstration that transparent offloading can match non-transparent offloading in both latency and energy, while requiring zero source-code modification.
Load-bearing premise
Every inference must begin with exactly one host-to-device copy and end with exactly one device-to-host copy, with no overlapping inferences, so that the recorded log contains clean, repeated brackets around the operator sequence.
Editorial extensions
If this is right
- Transparent offloading becomes a default candidate for static-activation models on resource-constrained mobiles: the same code that runs locally runs remotely without edits, at performance comparable to hand-modified offloading.
- Because the search identifies the operator sequence and its data dependencies, it exposes model architecture information that layer-partitioning and operator-level scheduling techniques can consume from inside a transparent system.
- The sharp drop in remote calls raises GPU-server utilization (from about 1.1% to 27.5% in the reported experiment), which implies shared edge servers can serve more concurrent transparent clients at lower per-inference energy.
- When a model's sequence changes, RRTO falls back to the standard transparent path and re-searches, so the downside of unsupported dynamic models is bounded by the old slow path.
Reading between the lines
- A natural extension the paper leaves implicit is applying the same record/replay idea to any fixed-sequence GPU workload, such as robotics control loops or signal-processing pipelines, not just ML inference.
- The benefit is a function of round-trip time: on low-latency data-center links the per-operator RPC penalty shrinks, so RRTO's advantage over the transparent baseline should be largest exactly in the high-RTT wireless MEC settings the paper targets.
- One testable consequence: if a model's first inference is unusual (e.g., lazy initialization or just-in-time compilation), the timing of the recording phase determines how quickly the system converges; a stress test varying the number of recording inferences could quantify initialization-noise robustness.
- The single-bracket assumption could be probed by measuring whether batched or multi-stream inference engines break sequence detection in practice, which would motivate a future variant that handles interleaved host-to-device and device-to-host markers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes RRTO, a transparent offloading system for mobile edge computing that eliminates per-operator RPC communication by recording CUDA runtime API calls during an initial phase, identifying the repeated operator sequence of a static-activation model via a new Operator Sequence Search algorithm (FastCheck/FullCheck), and then replaying that sequence on the GPU server for subsequent inferences. The system is implemented on top of Cricket and evaluated on a Jetson Xavier NX robot for KAPAO and several Torchvision models, reporting large reductions in inference latency and energy compared with Cricket and performance comparable to native non-transparent offloading (NNTO), all without source-code modification.
Significance. If the central claim holds, RRTO is a meaningful advance: it shows that transparent offloading can approach the performance of non-transparent offloading for static-activation models, removing a long-standing trade-off between compatibility and performance in MEC inference. The paper's direct measurements, the RPC micro-analysis (5895 RPCs per inference for Cricket versus 11 for RRTO), and the released code are concrete strengths that support reproducibility. The scope is honestly limited to static-activation models, with a fallback for dynamic ones. However, the evaluation lacks statistical rigor in a highly variable wireless environment, the headline abstract number is not supported by the reported results, and the sequence-extraction algorithm relies on a non-overlap assumption that is not stated as a limitation or tested against common concurrent/overlapped execution patterns. These issues are addressable but currently leave the central parity claim less firmly established than the text suggests.
major comments (5)
- [Abstract and Sec. V-A, Fig. 10] The abstract claims reductions of 'up to 98%' in both per-inference latency and energy, but the evaluation section reports only up to 95% latency reduction and up to 94% energy reduction for KAPAO, and no model in the text is shown to reach 98%. Please either correct the abstract to match the reported measurements or point to the specific experiment supporting the 98% figure.
- [Sec. III-B2, Observation 2 and Alg. 1, Lines 2-3, 7-8] The Operator Sequence Search assumes that each inference is cleanly bracketed by a single cudaMemcpyHtoD at the start and a single cudaMemcpyDtoH at the end, with no overlapping inferences. Many real inference engines use double buffering, multi-stream pipelines, or CUDA graphs, which can interleave the HtoD of inference n+1 with the DtoH of inference n, making the extracted 'sequence' a cross-inference composite or cyclic shift that fails FullCheck's data-dependency test. The paper does not state that such concurrent or overlapped execution is out of scope, and the abstract's general transparency claim implies broad applicability. Please either add a robustness experiment with overlapped/double-buffered inference or explicitly restrict the scope and temper the claims.
- [Alg. 1 and Sec. V] The algorithm's correctness depends on two undisclosed inputs: R (minimum number of repeats required by FastCheck/FullCheck) and N (the number of recorded inferences used as input to Operator Sequence Search). These are never given concrete values, and there is no sensitivity analysis showing how the results change with R and N. Without this information, a reader cannot judge whether the reported success is robust or a fortunate choice of parameters, nor can the experiments be reproduced exactly.
- [Sec. V, Fig. 3 and Fig. 10] The evaluation reports average latency and energy in a wireless environment that the paper itself shows to be highly variable (Fig. 3 has near-zero bandwidth drops outdoors), yet no error bars, confidence intervals, or number of repeated trials are reported. The central parity claim is empirical, so statistical support is load-bearing. Please report per-condition distributions or confidence intervals over multiple runs.
- [Alg. 3, Lines 19-20 and Alg. 4] During replay, the client returns the recorded 'ret' value for all functions other than HtoD and DtoH, and the server ignores the return values of intermediate operators. If a server-side CUDA kernel fails asynchronously during replay, the error is not propagated to the client until the final DtoH, so a transient GPU fault can silently produce wrong inference outputs while the application sees cudaSuccess for every kernel launch. The paper should either implement and evaluate error propagation in replay or state this as a correctness limitation.
minor comments (7)
- [Sec. II-B] There is a typo: 'sreduce' should be 'reduce' in the sentence discussing TorchScript.
- [Sec. IV-A] 'CUDA-acclerated' should be 'CUDA-accelerated'.
- [Sec. V-C] The model name 'RetainNet' appears to be a typo for 'RetinaNet'.
- [Sec. II-A] The phrase 'compared than' is ungrammatical and should be 'compared to' or 'than'.
- [Sec. V, Fig. 10 and Fig. 12] The figure captions should state whether the plotted points are single measurements or means over multiple trials, and the number of trials should be given.
- [References, [27]] The companion paper 'Intra-dp' is listed but not discussed in the body; the relationship between RRTO and that work should be clarified to avoid novelty ambiguity.
- [Sec. I and Sec. III-B1] The phrase 'the first high-performance transparent offloading system' is a strong claim; suggesting 'a' or adding a comparison to prior high-performance transparent systems would be more defensible, especially given the implicit scope restriction to static-activation models.
Circularity Check
No significant circularity: RRTO's headline results come from direct measurements against external baselines, and its core sequence-search algorithm is an engineering procedure rather than a fitted predictor.
full rationale
RRTO's central claims are supported by direct experimental measurements on a physical robot against external baselines (Cricket and NNTO), not by fitting a model to data and then predicting the same data. The Operator Sequence Search algorithm (Alg. 1-2) takes raw CUDA API logs and outputs a repeated operator sequence; its correctness is an implementation property, and its output is validated by end-to-end inference results. No equation in the paper reduces to a fitted quantity or to a self-citation. The only self-citation of note is reference [27], the authors' companion paper 'Intra-DP', which appears in the reference list but is not load-bearing for any argument in this paper. The known limitations, such as reliance on Observation 2's clean HtoD/DtoH bracketing and non-overlapping inferences (Sec. III-B2), are assumptions about execution environments, not circular derivations. Therefore the circularity score is low.
Assumptions & free parameters
free parameters (2)
- R (minimum number of repeats required by FastCheck/FullCheck) =
not disclosed
- N (number of recorded inferences used in Operator Sequence Search) =
not disclosed
assumptions (4)
- domain assumption Static-activation models (SAMs) dominate mobile workloads, making them the primary target of RRTO.
- domain assumption A static model repeats the exact same operator sequence on every inference (Observation 1).
- domain assumption Each inference is bracketed by a cudaMemcpyHtoD at the start and cudaMemcpyDtoH at the end, with no overlapping inferences (Observation 2).
- domain assumption Data dependencies can be tracked by CUDA memory address equality: inputs come from raw input, a prior operator's output at the same address, or model parameters (Observation 3).
Cite this review
Pith. "Pith review of RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing." pith.science (2026). https://pith.science/paper/74ZZTH3Q
@misc{pith2026250721739,
author = {Pith},
title = {Pith review of: RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing},
year = {2026},
howpublished = {\url{https://pith.science/paper/74ZZTH3Q}},
note = {Machine review of arXiv:2507.21739}
}
read the original abstract
Deploying Machine Learning (ML) applications on resource-constrained mobile devices remains challenging due to limited computational resources and poor platform compatibility. While Mobile Edge Computing (MEC) offers offloading-based inference paradigm using GPU servers, existing approaches are divided into non-transparent and transparent methods, with the latter necessitating modifications to the source code. Non-transparent offloading achieves high performance but requires intrusive code modification, limiting compatibility with diverse applications. Transparent offloading, in contrast, offers wide compatibility but introduces significant transmission delays due to per-operator remote procedure calls (RPCs). To overcome this limitation, we propose RRTO, the first high-performance transparent offloading system tailored for MEC inference. RRTO introduces a record/replay mechanism that leverages the static operator sequence in ML models to eliminate repetitive RPCs. To reliably identify this sequence, RRTO integrates a novel Operator Sequence Search algorithm that detects repeated patterns, filters initialization noise, and accelerates matching via a two-level strategy. Evaluation demonstrates that RRTO achieves substantial reductions of up to 98% in both per-inference latency and energy consumption compared to state-of-the-art transparent methods and yields results comparable to non-transparent approaches, all without necessitating any source code modification.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 4 Pith papers
-
Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification
FogFool creates fog-based adversarial perturbations using Perlin noise optimization to achieve high black-box transferability (83.74% TASR) and robustness to defenses in remote sensing classification.
-
Dual-Envelope Constrained Nonlinear MPC for Distributed Drive Electric Vehicles Drifting Under Bounded Steering and Direct Yaw-Moment Control
The extended dual-envelope NMPC enables smoother drifting convergence and cuts steady-state tracking errors in speed, sideslip angle, and yaw rate by 33%, 71%, and 31% respectively in hardware tests.
-
SL-FAC: A Communication-Efficient Split Learning Framework with Frequency-Aware Compression
SL-FAC reduces communication in split learning via frequency-aware compression of activations and gradients while aiming to preserve training-critical information.
-
SwarmSense-DNN: A Trustworthy and Decentralized Neural Framework for Proactive Anomaly Defense in Consumer IoT
SwarmSense-DNN is a proposed decentralized neural framework that integrates swarm intelligence with hierarchical federated learning and graph neural networks to achieve 95.44% anomaly detection accuracy and 67% reduce...
Reference graph
Works this paper leans on
-
[27]
Intra-dp: A high performance collabora- tive inference system for mobile edge computing,
Z. Sun, X. Guan, Z. Lin, Z. Fang, X. Cai, Z. Chen, F. Liu, H. Cui, J. Xiong, W. Ni et al. , “Intra-dp: A high performance collabora- tive inference system for mobile edge computing,” arXiv preprint arXiv:2507.05829, 2025
-
[1]
Binarized Neural Network for Edge Intelligence of Sensor-Based Human Activity Recognition,
F. Luo, S. Khan, Y . Huang, and K. Wu, “Binarized Neural Network for Edge Intelligence of Sensor-Based Human Activity Recognition,” IEEE Trans. Mobile Comput. , vol. 22, no. 3, pp. 1356–1368, Mar. 2023
2023
-
[2]
MERIT: Multimodal Wearable Vital Sign Waveform Monitoring,
Y . Tang, Z. Chen, A. Li, T. Zheng, Z. Lin, J. Xu, P. Lv, Z. Sun, and Y . Gao, “MERIT: Multimodal Wearable Vital Sign Waveform Monitoring,” arXiv preprint arXiv:2410.00392 , 2024
arXiv 2024
-
[3]
DNN Patching: Progressive Fixing and Augmenting the Functionalities of DNNs for Autonomous Vehicles,
A. N. Saridena and A. Choromanska, “DNN Patching: Progressive Fixing and Augmenting the Functionalities of DNNs for Autonomous Vehicles,” IEEE Robot. Autom. Lett. , vol. 7, no. 2, pp. 3257–3264, Apr. 2022
2022
-
[4]
Channel Power Gain Estimation for Terahertz Vehicle-to-Infrastructure Networks,
Z. Lin, L. Wang, J. Ding, B. Tan, and S. Jin, “Channel Power Gain Estimation for Terahertz Vehicle-to-Infrastructure Networks,”IEEE Commun. Lett., vol. 27, no. 1, pp. 155–159, 2022
2022
-
[5]
IC3M: In-Car Multimodal Multi-Object Monitoring for Abnormal Status of Both Driver and Passengers,
Z. Fang, Z. Lin, S. Hu, H. Cao, Y . Deng, X. Chen, and Y . Fang, “IC3M: In-Car Multimodal Multi-Object Monitoring for Abnormal Status of Both Driver and Passengers,” arXiv preprint arXiv:2410.02592 , 2024
arXiv 2024
-
[6]
Fedsn: A federated learning framework over heterogeneous leo satellite networks,
Z. Lin, Z. Chen, Z. Fang, X. Chen, X. Wang, and Y . Gao, “Fedsn: A federated learning framework over heterogeneous leo satellite networks,” IEEE Transactions on Mobile Computing , 2024
2024
-
[7]
A Novel Framework of Three-Hierarchical Offloading Optimization for MEC in Industrial IoT Networks,
Z. Zhao, R. Zhao, J. Xia, X. Lei, D. Li, C. Yuen, and L. Fan, “A Novel Framework of Three-Hierarchical Offloading Optimization for MEC in Industrial IoT Networks,” IEEE Trans. Ind. Informat., vol. 16, no. 8, pp. 5424–5434, 2019
2019
Show all 92 references
-
[8]
SigChord: Sniffing Wide Non-Sparse Multiband Signals for Terrestrial and Non- Terrestrial Wireless Networks,
J. Peng, J. Duan, Z. Lin, H. Yuan, Y . Gao, and Z. Chen, “SigChord: Sniffing Wide Non-Sparse Multiband Signals for Terrestrial and Non- Terrestrial Wireless Networks,” arXiv preprint arXiv:2504.06587, 2025
2025 arXiv
-
[9]
Constructing 4D Radio Map in LEO Satellite Networks with Limited Samples,
H. Yuan, Z. Chen, Z. Lin, J. Peng, Y . Zhong, X. Hu, S. Xue, W. Li, and Y . Gao, “Constructing 4D Radio Map in LEO Satellite Networks with Limited Samples,” IEEE INFOCOM, 2025
2025
-
[10]
RF- Based Human Activity Recognition Using Signal Adapted Convolutional Neural Network,
Z. Chen, C. Cai, T. Zheng, J. Luo, J. Xiong, and X. Wang, “RF- Based Human Activity Recognition Using Signal Adapted Convolutional Neural Network,” IEEE Trans. Mobile Comput., vol. 22, no. 1, pp. 487– 499, 2021
2021
-
[11]
Spatial-spectral terahertz networks,
Z. Lin, L. Wang, B. Tan, and X. Li, “Spatial-spectral terahertz networks,” IEEE Transactions on Wireless Communications , vol. 21, no. 6, pp. 3881–3892, 2021
2021
-
[12]
Fedac: An adaptive clustered federated learning framework for heterogeneous data,
Y . Zhang, H. Chen, Z. Lin, Z. Chen, and J. Zhao, “Fedac: An adaptive clustered federated learning framework for heterogeneous data,” arXiv preprint arXiv:2403.16460, 2024
2024 arXiv
-
[13]
SatSense: Multi-Satellite Collaborative Framework for Spectrum Sensing,
H. Yuan, Z. Chen, Z. Lin, J. Peng, Z. Fang, Y . Zhong, Z. Song, and Y . Gao, “SatSense: Multi-Satellite Collaborative Framework for Spectrum Sensing,” IEEE Trans. Cogn. Commun. Netw. , 2025
2025
-
[14]
SUMS: Sniffing Unknown Multiband Signals under Low Sampling Rates,
J. Peng, Z. Chen, Z. Lin, H. Yuan, Z. Fang, L. Bao, Z. Song, Y . Li, J. Ren, and Y . Gao, “SUMS: Sniffing Unknown Multiband Signals under Low Sampling Rates,” IEEE Trans. Mobile Comput. , 2024
2024
-
[15]
LEO Satellite Networks Assisted Geo-Distributed Data Processing,
Z. Zhao, Z. Chen, Z. Lin, W. Zhu, K. Qiu, C. You, and Y . Gao, “LEO Satellite Networks Assisted Geo-Distributed Data Processing,” IEEE Wireless Commun. Lett., 2024
2024
-
[16]
Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,
Z. Lin, W. Wei, Z. Chen, C.-T. Lam, X. Chen, Y . Gao, and J. Luo, “Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,” IEEE Trans. Mobile Comput. , 2025
2025
-
[17]
Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi- Person Human Pose Estimation,
W. McNally, K. Vats, A. Wong, and J. McPhee, “Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi- Person Human Pose Estimation,” in Proc. ECCV, Oct. 2022, pp. 37–54
2022
-
[18]
Efficient Parallel Split Learning over Resource-Constrained Wireless Edge Networks,
Z. Lin, G. Zhu, Y . Deng, X. Chen, Y . Gao, K. Huang, and Y . Fang, “Efficient Parallel Split Learning over Resource-Constrained Wireless Edge Networks,” IEEE Trans. Mobile Comput. , vol. 23, no. 10, pp. 9224–9239, 2024
2024
-
[19]
AGRNav: Efficient and Energy-Saving Au- tonomous Navigation for Air-Ground Robots in Occlusion-Prone En- vironments,
J. Wang, Z. Sun, X. Guan, T. Shen, Z. Zhang, T. Duan, D. Huang, S. Zhao, and H. Cui, “AGRNav: Efficient and Energy-Saving Au- tonomous Navigation for Air-Ground Robots in Occlusion-Prone En- vironments,” in Proc. ICRA, May 2024, pp. 4494–4501
2024
-
[20]
Rethinking adversarial attacks in reinforcement learning from policy distribution perspective,
T. Duan, Z. Zhang, Z. Lin, Y . Gao, L. Xiong, Y . Cui, H. Liang, X. Chen, H. Cui, and D. Huang, “Rethinking adversarial attacks in reinforcement learning from policy distribution perspective,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Pr...
2025
-
[21]
Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,
Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, and K. Huang, “Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,” arXiv preprint arXiv:2309.16739 , 2023
2023 arXiv
-
[22]
State-aware perturbation optimization for robust deep reinforcement learning,
Z. Zhang, T. Duan, Z. Lin, D. Huang, Z. Fang, Z. Sun, L. Xiong, H. Liang, H. Cui, and Y . Cui, “State-aware perturbation optimization for robust deep reinforcement learning,” arXiv preprint arXiv:2503.20613 , 2025
2025 arXiv
-
[23]
HASFL: Heterogeneity- Aware Split Federated Learning over Edge Computing Systems,
Z. Lin, Z. Chen, X. Chen, W. Ni, and Y . Gao, “HASFL: Heterogeneity- Aware Split Federated Learning over Edge Computing Systems,” arXiv preprint arXiv:2506.08426, 2025
2025 arXiv
-
[24]
MonoScene: Monocular 3D Semantic Scene Completion,
A.-Q. Cao and R. de Charette, “MonoScene: Monocular 3D Semantic Scene Completion,” in Proc. CVPR, Jun. 2022, pp. 3991–4001
2022
-
[25]
Optimal resource allocation for u-shaped parallel split learning,
S. Lyu, Z. Lin, G. Qu, X. Chen, X. Huang, and P. Li, “Optimal resource allocation for u-shaped parallel split learning,” in 2023 IEEE Globecom Workshops (GC Wkshps), 2023, pp. 197–202
2023
-
[26]
SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models,
Z. Lin, X. Hu, Y . Zhang, Z. Chen, Z. Fang, X. Chen, A. Li, P. Vepakomma, and Y . Gao, “SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models,” arXiv preprint arXiv:2407.00952, 2024
2024 arXiv
-
[28]
Adaptsfl: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,
Z. Lin, G. Qu, W. Wei, X. Chen, and K. K. Leung, “Adaptsfl: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,” IEEE Trans. Netw., 2024
2024
-
[29]
Mobile Edge Computing: A Survey,
N. Abbas, Y . Zhang, A. Taherkordi, and T. Skeie, “Mobile Edge Computing: A Survey,” IEEE Internet Things J. , vol. 5, no. 1, pp. 450– 465, Feb. 2018
2018
-
[30]
The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSL,
B. Qiao, M. A. ¨Ozkan, J. Teich, and F. Hannig, “The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSL,” in Proc. DAC, 2020, pp. 1–6
2020
-
[31]
Cricket: A Virtual- ization Layer for Distributed Execution of CUDA Applications with Checkpoint/Restart Support,
N. Eiling, J. Baude, S. Lankes, and A. Monti, “Cricket: A Virtual- ization Layer for Distributed Execution of CUDA Applications with Checkpoint/Restart Support,” Concurrency Comput. Pract. Exp., vol. 34, no. 14, p. e6474, May 2022
2022
-
[32]
TensorRT Inference with TensorFlow,
P. Davoodi, C. Gwon, G. Lai, and T. Morris, “TensorRT Inference with TensorFlow,” in Proc. GPU Technol. Conf. (GTC) , Mar. 2019
2019
-
[33]
A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture,
X. Yi, “A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture,” arXiv preprint arXiv:2409.10661 , 2024
2024 arXiv
-
[34]
Infer-EDGE: Dynamic DNN Inference Optimization in ‘Just-in-Time’ Edge-AI Implementations,
M. Mounesan, X. Zhang, and S. Debroy, “Infer-EDGE: Dynamic DNN Inference Optimization in ‘Just-in-Time’ Edge-AI Implementations,” arXiv preprint arXiv:2501.18842 , 2025
2025 arXiv
-
[35]
TorchScript: Optimized Execution of PyTorch Programs,
Z. DeVito, “TorchScript: Optimized Execution of PyTorch Programs,” Retrieved January, 2022
2022
-
[36]
PyTorch: An Imperative Style, High- Performance Deep Learning Library,
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An Imperative Style, High- P...
1912 arXiv
-
[37]
Implementing Open-Source CUDA Runtime,
S. Kato, “Implementing Open-Source CUDA Runtime,” in Proc. of the 54the Programming Symposium, Hakone, Japan, Jan. 2013, pp. 111–118
2013
-
[38]
Evaluation of Communication Technologies for Distributed Industrial Control Sys- tems: Concept and Evaluation of 5G and WiFi 6,
M. Schneider, F. Haag, A. K. Khalil, and D. A. Breunig, “Evaluation of Communication Technologies for Distributed Industrial Control Sys- tems: Concept and Evaluation of 5G and WiFi 6,” Procedia CIRP, vol. 107, pp. 588–593, 2022
2022
-
[39]
TensorFlow: A System for Large-Scale Machine Learning,
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “TensorFlow: A System for Large-Scale Machine Learning,” in Proc. OSDI , Nov. 2016, pp. 265– 283
2016
-
[40]
The OpenCL Specification,
A. Munshi, “The OpenCL Specification,” in Proc. IEEE Hot Chips 21 Symp. (HCS), Aug. 2009
2009
-
[41]
CoDL: Efficient CPU-GPU Co-Execution for Deep Learning Inference on Mobile Devices,
F. Jia, D. Zhang, T. Cao, S. Jiang, Y . Liu, J. Ren, and Y . Zhang, “CoDL: Efficient CPU-GPU Co-Execution for Deep Learning Inference on Mobile Devices,” in Proc. MobiSys, Jun. 2022, pp. 209–221
2022
-
[42]
Accelerating Robot Dynamics Gradients on a CPU, GPU, and FPGA,
B. Plancher, S. M. Neuman, T. Bourgeat, S. Kuindersma, S. Devadas, and V . J. Reddi, “Accelerating Robot Dynamics Gradients on a CPU, GPU, and FPGA,” IEEE Robot. Autom. Lett. , vol. 6, no. 2, pp. 2335– 2342, Apr. 2021. 15
2021
-
[43]
DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices,
N. D. Lane, S. Bhattacharya, P. Georgiev, C. Forlivesi, L. Jiao, L. Qen- dro, and F. Kawsar, “DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices,” in Proc. IPSN, Apr. 2016, pp. 1–12
2016
-
[44]
REACT: Streaming Video Analytics on the Edge with Asynchronous Cloud Support,
A. Ghosh, S. Iyengar, S. Lee, A. Rathore, and V . N. Padmanabhan, “REACT: Streaming Video Analytics on the Edge with Asynchronous Cloud Support,” in Proc. IoTDI, May 2023, pp. 222–235
2023
-
[45]
MLPerf Mobile Benchmarks,
“MLPerf Mobile Benchmarks,” https://mlcommons.org/working-groups/ benchmarks/mobile/, 2023
2023
-
[46]
Raspberry Pi,
RaspberryPi, “Raspberry Pi,” https://www.raspberrypi.com/ documentation/computers/configuration.html\#power-consumption, 2022
2022
-
[47]
Jetson Xavier NX Series: The World’s Smallest AI Supercomputer,
NVIDIA, “Jetson Xavier NX Series: The World’s Smallest AI Supercomputer,” https://www.nvidia.com/en-us/autonomous-machines/ embedded-systems/jetson-xavier-nx-developer-kit/, 2024
2024
-
[48]
Power Consumption Benchmark for Embedded AI Inference,
Z. Ning, M. Vandersteegen, K. Van Beeck, T. Goedem ´e, and P. Vande- walle, “Power Consumption Benchmark for Embedded AI Inference,” in Proc. Int. Conf. Appl. Comput. WWW/Internet (AC) , Jan. 2024, pp. 3–10
2024
-
[49]
Local Feature Selection for Data Classification,
N. Armanfard, J. P. Reilly, and M. Komeili, “Local Feature Selection for Data Classification,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 38, no. 6, pp. 1217–1227, 2015
2015
-
[50]
Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,
Y . Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,” in Proc. ASPLOS, Apr. 2017, pp. 615–629
2017
-
[51]
Computation Offloading Toward Edge Computing,
L. Lin, X. Liao, H. Jin, and P. Li, “Computation Offloading Toward Edge Computing,” Proc. IEEE , vol. 107, no. 8, pp. 1584–1607, Aug. 2019
2019
-
[52]
QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural Networks,
Z. Fang, T. Yu, O. J. Mengshoel, and R. K. Gupta, “QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural Networks,” in Proc. CIKM, Nov. 2017, pp. 2067–2070
2017
-
[53]
InfiniBand Networking Solutions,
NVIDIA, “InfiniBand Networking Solutions,” https://www.nvidia.com/ en-us/networking/products/infiniband/, 2024
2024
-
[54]
A First Look at Wi-Fi 6 in Action: Throughput, Latency, Energy Efficiency, and Security,
R. Liu and N. Choi, “A First Look at Wi-Fi 6 in Action: Throughput, Latency, Energy Efficiency, and Security,” in Proc. ACM Meas. Anal. Comput. Syst., vol. 7, no. 1, Mar. 2023, pp. 1–25
2023
-
[55]
Mobile Access Bandwidth in Practice: Measurement, Analysis, and Implications,
X. Yang, H. Lin, Z. Li, F. Qian, X. Li, Z. He, X. Wu, X. Wang, Y . Liu, Z. Liao et al. , “Mobile Access Bandwidth in Practice: Measurement, Analysis, and Implications,” in Proc. SIGCOMM, Aug. 2022, pp. 114– 128
2022
-
[56]
Throughput Comparison between The New HEW 802.11 ax Standard and 802.11 n/ac Standards in Selected Distance Windows,
A. Masiukiewicz, “Throughput Comparison between The New HEW 802.11 ax Standard and 802.11 n/ac Standards in Selected Distance Windows,” Int. J. Electron. Telecommun. , vol. 65, no. 1, pp. 79–84, 2019
2019
-
[57]
Performance Impact of LoS and NLoS Transmissions in Dense Cellular Networks,
M. Ding, P. Wang, D. L ´opez-P´erez, G. Mao, and Z. Lin, “Performance Impact of LoS and NLoS Transmissions in Dense Cellular Networks,” IEEE Trans. Wireless Commun. , vol. 15, no. 3, pp. 2365–2380, Mar. 2016
2016
-
[58]
Proportional and Preemption-Enabled Traffic Offloading for IP Flow Mobility: Algo- rithms and Performance Evaluation,
Y . Ren, C.-W. Tung, J.-C. Chen, and F. Y . Li, “Proportional and Preemption-Enabled Traffic Offloading for IP Flow Mobility: Algo- rithms and Performance Evaluation,” IEEE Trans. Veh. Technol., vol. 67, no. 12, pp. 12 095–12 108, Dec. 2018
2018
-
[59]
iPerf - Download iPerf3 and Original iPerf Pre-Compiled Binaries,
“iPerf - Download iPerf3 and Original iPerf Pre-Compiled Binaries,” https://iperf.fr/iperf-download.php, 2024
2024
-
[60]
TCP in Wireless Environments: Problems and Solutions,
Y . Tian, K. Xu, and N. Ansari, “TCP in Wireless Environments: Problems and Solutions,” IEEE Commun. Mag., vol. 43, no. 3, pp. S27– S32, Mar. 2005
2005
-
[61]
NVIDIA Nsight Compute Documentation,
NVIDIA, “NVIDIA Nsight Compute Documentation,” https://docs. nvidia.com/nsight-compute/index.html, 2024
2024
-
[62]
libtirpc: Transport Independent RPC library,
S. Dickson, “libtirpc: Transport Independent RPC library,” https://git. linux-nfs.org/?p=steved/libtirpc.git, 2024
2024
-
[63]
Cliquemap: Productionizing an RMA-Based Distributed Caching Sys- tem,
A. Singhvi, A. Akella, M. Anderson, R. Cauble, H. Deshmukh, D. Gib- son, M. M. Martin, A. Strominger, T. F. Wenisch, and A. Vahdat, “Cliquemap: Productionizing an RMA-Based Distributed Caching Sys- tem,” in Proc. SIGCOMM. ACM, 2021, pp. 93–105
2021
-
[64]
Dagger: Efficient and Fast RPCs in Cloud Microservices With Near-Memory Reconfigurable NICs,
N. Lazarev, S. Xiang, N. Adit, Z. Zhang, and C. Delimitrou, “Dagger: Efficient and Fast RPCs in Cloud Microservices With Near-Memory Reconfigurable NICs,” in Proc. ASPLOS. ACM, 2021, pp. 36–51
2021
-
[65]
Efficient Asynchronous RPC Calls for Mi- croservices: DeathStarBench Study,
S. Eyerman and I. Hur, “Efficient Asynchronous RPC Calls for Mi- croservices: DeathStarBench Study,” arXiv preprint arXiv:2209.13265 , 2022
2022 arXiv
-
[66]
Interpretable and Accurate Fine-Grained Recog- nition via Region Grouping,
Z. Huang and Y . Li, “Interpretable and Accurate Fine-Grained Recog- nition via Region Grouping,” in Proc. CVPR, 2020, pp. 8662–8672
2020
-
[67]
GMC: Graph-Based Multi-View Clustering,
H. Wang, Y . Yang, and B. Liu, “GMC: Graph-Based Multi-View Clustering,” IEEE Trans. Knowledge Data Eng., vol. 32, no. 6, pp. 1116– 1129, 2019
2019
-
[68]
Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschel et al. , “Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks,” in Proc. NeurIPS, vol. 33, 2020, pp. 9459–9474
2020
-
[69]
BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint arXiv:1810.04805, Oct. 2018
2018 arXiv
-
[70]
An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling,
S. Bai, J. Z. Kolter, and V . Koltun, “An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling,” arXiv preprint arXiv:1803.01271, 2018
2018 arXiv
-
[71]
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in Proc. ICLR, 2020
2020
-
[72]
Language Models Are Few-Shot Learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language Models Are Few-Shot Learners,” in Proc. NeurIPS , vol. 33, 2020, pp. 1877– 1901
2020
-
[73]
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” in Proc. ICLR, 2017
2017
-
[74]
Sequence to Sequence Learning with Neural Networks,
I. Sutskever, O. Vinyals, and Q. V . Le, “Sequence to Sequence Learning with Neural Networks,” in Proc. NeurIPS, vol. 27, 2014
2014
-
[75]
XGBoost: A Scalable Tree Boosting System,
T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proc. KDD, 2016, pp. 785–794
2016
-
[76]
Semi-Supervised Learning with Graph Learning-Convolutional Networks,
B. Jiang, Z. Zhang, D. Lin, J. Tang, and B. Luo, “Semi-Supervised Learning with Graph Learning-Convolutional Networks,” in Proc. CVPR, 2019, pp. 11 313–11 320
2019
-
[77]
Efficient Algorithms for Device Placement of DNN Graph Operators,
J. M. Tarnawski, A. Phanishayee, N. Devanur, D. Mahajan, and F. N. Paravecino, “Efficient Algorithms for Device Placement of DNN Graph Operators,” in Proc. NeurIPS, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33, Dec. 2020, pp. 15 451–15 463
2020
-
[78]
Faster Algorithms for Longest Common Substring,
P. Charalampopoulos, T. Kociumaka, S. P. Pissis, and J. Radoszewski, “Faster Algorithms for Longest Common Substring,” arXiv preprint arXiv:2105.03106, 2021
2021
-
[79]
Torchvision the Machine-Vision Package of Torch,
S. Marcel and Y . Rodriguez, “Torchvision the Machine-Vision Package of Torch,” in Proc. ACM Multimedia, Oct. 2010, pp. 1485–1488
2010
-
[80]
Deep Residual Learning for Image Recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” arXiv preprint arXiv:1512.03385 , Dec. 2015
2015 arXiv
-
[81]
A ConvNet for the 2020s,
Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” arXiv preprint arXiv:2201.03545 , Jan. 2022
2022 arXiv
-
[82]
Fully Convolutional Networks for Semantic Segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” arXiv preprint arXiv:1411.4038 , Nov. 2015
2015 arXiv
-
[83]
Rethinking Atrous Convolution for Semantic Image Segmentation,
L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking Atrous Convolution for Semantic Image Segmentation,” arXiv preprint arXiv:1706.05587, Jun. 2017
2017 arXiv
-
[84]
Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,” arXiv preprint arXiv:1506.01497, Jun. 2015
2015 arXiv
-
[85]
Focal Loss for Dense Object Detection,
T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal Loss for Dense Object Detection,” arXiv preprint arXiv:1708.02002, Aug. 2017
2017 arXiv
-
[86]
Datacenter Traffic Control: Understanding Techniques and Tradeoffs,
M. Noormohammadpour and C. S. Raghavendra, “Datacenter Traffic Control: Understanding Techniques and Tradeoffs,” IEEE Commun. Surv. Tutor., vol. 20, no. 2, pp. 1492–1525, Jun. 2018
2018
-
[87]
ultralytics/yolov5: v7.0 - YOLOv5 SOTA Realtime Instance Segmentation,
G. Jocher, A. Chaurasia, A. Stoken, J. Borovec, NanoCode012, Y . Kwon, K. Michael, T. Xie, J. Fang, imyhxy, Lorna, Z. Yifu, C. Wong, A. V , D. Montes, Z. Wang, C. Fati, J. Nadar, Laughing, UnglvKitDe, V . Sonck, tkianai, yxNONG, P. Skalski, A. Hogan, D. Nair, M. Strobel, and M...
2022
-
[88]
pynvml: Python Bindings for the NVIDIA Management Library,
NVIDIA, “pynvml: Python Bindings for the NVIDIA Management Library,” https://developer.nvidia.com/management-library-nvml/, 2024
2024
-
[89]
Optimal Path Planning of Au- tonomous Marine Vehicles in Stochastic Dynamic Ocean Flows Using a GPU-Accelerated Algorithm,
R. Chowdhury and D. Subramani, “Optimal Path Planning of Au- tonomous Marine Vehicles in Stochastic Dynamic Ocean Flows Using a GPU-Accelerated Algorithm,” IEEE J. Ocean. Eng. , vol. 47, no. 4, pp. 864–879, Oct. 2022
2022
-
[90]
Query Processing on Hetero- geneous CPU/GPU Systems,
V . Rosenfeld, S. Breß, and V . Markl, “Query Processing on Hetero- geneous CPU/GPU Systems,” ACM Comput. Surv. , vol. 55, no. 1, pp. 1–38, Jan. 2023
2023
-
[91]
VecQ: Minimal Loss DNN Model Compression with Vectorized Weight Quantization,
C. Gong, Y . Chen, Y . Lu, T. Li, C. Hao, and D. Chen, “VecQ: Minimal Loss DNN Model Compression with Vectorized Weight Quantization,” IEEE Trans. Comput. , vol. 70, no. 5, pp. 696–710, May 2021
2021
-
[92]
Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks,
L. Wang and K.-J. Yoon, “Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 6, pp. 3048–3068, Jun. 2022
2022
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.