Pith. sign in

REVIEW 5 major objections 7 minor 4 cited by

RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing

T0 review · 5 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read RRTO makes transparent offloading fast by recording a model's fixed operator sequence and replaying it on the GPU server, cutting per-inference RPCs from 5,895 to 11.

desk verdict RRTO's record/replay idea is a genuine step forward for transparent offloading, and the measurements largely back it up; the unstated scope on overlapping inference is the main caveat. read the letter →

arxiv 2507.21739 v1 pith:74ZZTH3Q submitted 2025-07-29 cs.NI

classification cs.NI
keywords transparentoffloadingmobileedgecomputingrecord/replayoperatorsequencesearchRPCeliminationCUDAAPIinterceptionstaticactivationmodelsmodelinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

RRTO claims that transparent offloading—running a model's CUDA calls on a remote GPU without touching application source code—no longer has to be slow. Because static-activation models execute the same operator sequence on every inference, RRTO records that sequence during the first few inferences, identifies it from raw call logs, and then replays it directly on the GPU server so that only the input and output cross the network. The paper reports that this cuts per-inference remote calls from 5,895 to 11 for the main test application, and reduces latency by roughly 94–95% and energy by 93–94% relative to the state-of-the-art transparent baseline. The claim matters because it would remove the compatibility-versus-performance trade-off that currently forces mobile developers to choose between source-code modifications and slow inference.

What carries the argument

The load-bearing mechanism is the Operator Sequence Search algorithm, which reconstructs the exact per-inference operator sequence from raw interception logs. It relies on three observations: the same operator sequence repeats across inferences; each inference is bracketed by a single host-to-device memory copy at the start and a single device-to-host copy at the end; and every operator's inputs must satisfy data-dependency constraints against the raw input, prior operator outputs, or model parameters. To keep the search tractable on logs with tens of thousands of entries, the algorithm first runs FastCheck on compact category tags to prune candidates by repeated occurrence, then runs FullCheck on the survivors to realign start/end markers, verify data dependencies, and confirm exact record-level repetition across the whole log.

What would settle it

Run RRTO under an execution engine that uses pre-pinned asynchronous DMA, persistent CUDA graphs, stream pooling, or two concurrently interleaved inference streams; if Operator Sequence Search then fails to find a valid repeated sequence and the system falls back to per-operator RPCs, or if replay produces wrong outputs, the central speedup claim is falsified for that setting.

Watch

Extended reading notes

Core claim

The paper's central discovery is that the per-operator RPC cost that dominates transparent offloading is not inherent to transparency but an artifact of reacting to operator calls as they arrive. For models whose operator sequence is fixed, the sequence can be discovered once and then replayed: the offloading client intercepts CUDA calls, records them during an initial recording phase, searches the accumulated log for the repeating inference pattern, and thereafter returns cached status results to the application while the server executes the whole recorded sequence in one shot. The paper argues this is the first demonstration that transparent offloading can match non-transparent offloading in both latency and energy, while requiring zero source-code modification.

Load-bearing premise

Every inference must begin with exactly one host-to-device copy and end with exactly one device-to-host copy, with no overlapping inferences, so that the recorded log contains clean, repeated brackets around the operator sequence.

Editorial extensions

If this is right

  • Transparent offloading becomes a default candidate for static-activation models on resource-constrained mobiles: the same code that runs locally runs remotely without edits, at performance comparable to hand-modified offloading.
  • Because the search identifies the operator sequence and its data dependencies, it exposes model architecture information that layer-partitioning and operator-level scheduling techniques can consume from inside a transparent system.
  • The sharp drop in remote calls raises GPU-server utilization (from about 1.1% to 27.5% in the reported experiment), which implies shared edge servers can serve more concurrent transparent clients at lower per-inference energy.
  • When a model's sequence changes, RRTO falls back to the standard transparent path and re-searches, so the downside of unsupported dynamic models is bounded by the old slow path.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is applying the same record/replay idea to any fixed-sequence GPU workload, such as robotics control loops or signal-processing pipelines, not just ML inference.
  • The benefit is a function of round-trip time: on low-latency data-center links the per-operator RPC penalty shrinks, so RRTO's advantage over the transparent baseline should be largest exactly in the high-RTT wireless MEC settings the paper targets.
  • One testable consequence: if a model's first inference is unusual (e.g., lazy initialization or just-in-time compilation), the timing of the recording phase determines how quickly the system converges; a stress test varying the number of recording inferences could quantify initialization-noise robustness.
  • The single-bracket assumption could be probed by measuring whether batched or multi-stream inference engines break sequence detection in practice, which would motivate a future variant that handles interleaved host-to-device and device-to-host markers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. The paper proposes RRTO, a transparent offloading system for mobile edge computing that eliminates per-operator RPC communication by recording CUDA runtime API calls during an initial phase, identifying the repeated operator sequence of a static-activation model via a new Operator Sequence Search algorithm (FastCheck/FullCheck), and then replaying that sequence on the GPU server for subsequent inferences. The system is implemented on top of Cricket and evaluated on a Jetson Xavier NX robot for KAPAO and several Torchvision models, reporting large reductions in inference latency and energy compared with Cricket and performance comparable to native non-transparent offloading (NNTO), all without source-code modification.

Significance. If the central claim holds, RRTO is a meaningful advance: it shows that transparent offloading can approach the performance of non-transparent offloading for static-activation models, removing a long-standing trade-off between compatibility and performance in MEC inference. The paper's direct measurements, the RPC micro-analysis (5895 RPCs per inference for Cricket versus 11 for RRTO), and the released code are concrete strengths that support reproducibility. The scope is honestly limited to static-activation models, with a fallback for dynamic ones. However, the evaluation lacks statistical rigor in a highly variable wireless environment, the headline abstract number is not supported by the reported results, and the sequence-extraction algorithm relies on a non-overlap assumption that is not stated as a limitation or tested against common concurrent/overlapped execution patterns. These issues are addressable but currently leave the central parity claim less firmly established than the text suggests.

major comments (5)
  1. [Abstract and Sec. V-A, Fig. 10] The abstract claims reductions of 'up to 98%' in both per-inference latency and energy, but the evaluation section reports only up to 95% latency reduction and up to 94% energy reduction for KAPAO, and no model in the text is shown to reach 98%. Please either correct the abstract to match the reported measurements or point to the specific experiment supporting the 98% figure.
  2. [Sec. III-B2, Observation 2 and Alg. 1, Lines 2-3, 7-8] The Operator Sequence Search assumes that each inference is cleanly bracketed by a single cudaMemcpyHtoD at the start and a single cudaMemcpyDtoH at the end, with no overlapping inferences. Many real inference engines use double buffering, multi-stream pipelines, or CUDA graphs, which can interleave the HtoD of inference n+1 with the DtoH of inference n, making the extracted 'sequence' a cross-inference composite or cyclic shift that fails FullCheck's data-dependency test. The paper does not state that such concurrent or overlapped execution is out of scope, and the abstract's general transparency claim implies broad applicability. Please either add a robustness experiment with overlapped/double-buffered inference or explicitly restrict the scope and temper the claims.
  3. [Alg. 1 and Sec. V] The algorithm's correctness depends on two undisclosed inputs: R (minimum number of repeats required by FastCheck/FullCheck) and N (the number of recorded inferences used as input to Operator Sequence Search). These are never given concrete values, and there is no sensitivity analysis showing how the results change with R and N. Without this information, a reader cannot judge whether the reported success is robust or a fortunate choice of parameters, nor can the experiments be reproduced exactly.
  4. [Sec. V, Fig. 3 and Fig. 10] The evaluation reports average latency and energy in a wireless environment that the paper itself shows to be highly variable (Fig. 3 has near-zero bandwidth drops outdoors), yet no error bars, confidence intervals, or number of repeated trials are reported. The central parity claim is empirical, so statistical support is load-bearing. Please report per-condition distributions or confidence intervals over multiple runs.
  5. [Alg. 3, Lines 19-20 and Alg. 4] During replay, the client returns the recorded 'ret' value for all functions other than HtoD and DtoH, and the server ignores the return values of intermediate operators. If a server-side CUDA kernel fails asynchronously during replay, the error is not propagated to the client until the final DtoH, so a transient GPU fault can silently produce wrong inference outputs while the application sees cudaSuccess for every kernel launch. The paper should either implement and evaluate error propagation in replay or state this as a correctness limitation.
minor comments (7)
  1. [Sec. II-B] There is a typo: 'sreduce' should be 'reduce' in the sentence discussing TorchScript.
  2. [Sec. IV-A] 'CUDA-acclerated' should be 'CUDA-accelerated'.
  3. [Sec. V-C] The model name 'RetainNet' appears to be a typo for 'RetinaNet'.
  4. [Sec. II-A] The phrase 'compared than' is ungrammatical and should be 'compared to' or 'than'.
  5. [Sec. V, Fig. 10 and Fig. 12] The figure captions should state whether the plotted points are single measurements or means over multiple trials, and the number of trials should be given.
  6. [References, [27]] The companion paper 'Intra-dp' is listed but not discussed in the body; the relationship between RRTO and that work should be clarified to avoid novelty ambiguity.
  7. [Sec. I and Sec. III-B1] The phrase 'the first high-performance transparent offloading system' is a strong claim; suggesting 'a' or adding a comparison to prior high-performance transparent systems would be more defensible, especially given the implicit scope restriction to static-activation models.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: RRTO's headline results come from direct measurements against external baselines, and its core sequence-search algorithm is an engineering procedure rather than a fitted predictor.

full rationale

RRTO's central claims are supported by direct experimental measurements on a physical robot against external baselines (Cricket and NNTO), not by fitting a model to data and then predicting the same data. The Operator Sequence Search algorithm (Alg. 1-2) takes raw CUDA API logs and outputs a repeated operator sequence; its correctness is an implementation property, and its output is validated by end-to-end inference results. No equation in the paper reduces to a fitted quantity or to a self-citation. The only self-citation of note is reference [27], the authors' companion paper 'Intra-DP', which appears in the reference list but is not load-bearing for any argument in this paper. The known limitations, such as reliance on Observation 2's clean HtoD/DtoH bracketing and non-overlapping inferences (Sec. III-B2), are assumptions about execution environments, not circular derivations. Therefore the circularity score is low.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

No new physical entities, mediators, or fitted constants are introduced. The system rests on domain assumptions about model staticity, CUDA copy boundaries, and address-traceable data dependencies; the hidden algorithm parameters R and N are free parameters that are not disclosed.

free parameters (2)
  • R (minimum number of repeats required by FastCheck/FullCheck) = not disclosed
    Input to Alg. 1; controls the minimum number of repetitions before a candidate sequence is accepted. The paper does not report this threshold or how it was chosen; it directly affects search runtime and the risk of false positives.
  • N (number of recorded inferences used in Operator Sequence Search) = not disclosed
    Alg. 1 says 'Logs: list of OperatorInfo entries from the first N inferences'; N is never specified, nor is it stated whether the search is run once or incrementally.
assumptions (4)
  • domain assumption Static-activation models (SAMs) dominate mobile workloads, making them the primary target of RRTO.
    Sec. III-B1 states 'Given the predominance of SAMs in mobile applications, these fallback events are expected to be infrequent'. If dynamic-activation models become the common case, RRTO's benefits shrink to the fallback path.
  • domain assumption A static model repeats the exact same operator sequence on every inference (Observation 1).
    This is the basis of repeated-pattern detection in Alg. 1 and Alg. 2. It is an assumption about model execution, not a derived result; JIT optimizations or runtime kernel selection can break it.
  • domain assumption Each inference is bracketed by a cudaMemcpyHtoD at the start and cudaMemcpyDtoH at the end, with no overlapping inferences (Observation 2).
    Alg. 1 Lines 2-3, 7-8 and Alg. 2 Line 11 rely on these boundaries. Persistent CUDA graphs, asynchronous DMA engines, or pipelined concurrent inferences would violate this assumption.
  • domain assumption Data dependencies can be tracked by CUDA memory address equality: inputs come from raw input, a prior operator's output at the same address, or model parameters (Observation 3).
    Alg. 2 Line 11 uses this in the data-dependency check. Memory pooling or address reuse could create false alias matches and cause an incorrect sequence to be accepted.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing." pith.science (2026). https://pith.science/paper/74ZZTH3Q

@misc{pith2026250721739,
  author       = {Pith},
  title        = {Pith review of: RRTO: A High-Performance Transparent Offloading System for Model Inference in Mobile Edge Computing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/74ZZTH3Q}},
  note         = {Machine review of arXiv:2507.21739}
}
read the original abstract

Deploying Machine Learning (ML) applications on resource-constrained mobile devices remains challenging due to limited computational resources and poor platform compatibility. While Mobile Edge Computing (MEC) offers offloading-based inference paradigm using GPU servers, existing approaches are divided into non-transparent and transparent methods, with the latter necessitating modifications to the source code. Non-transparent offloading achieves high performance but requires intrusive code modification, limiting compatibility with diverse applications. Transparent offloading, in contrast, offers wide compatibility but introduces significant transmission delays due to per-operator remote procedure calls (RPCs). To overcome this limitation, we propose RRTO, the first high-performance transparent offloading system tailored for MEC inference. RRTO introduces a record/replay mechanism that leverages the static operator sequence in ML models to eliminate repetitive RPCs. To reliably identify this sequence, RRTO integrates a novel Operator Sequence Search algorithm that detects repeated patterns, filters initialization noise, and accelerates matching via a two-level strategy. Evaluation demonstrates that RRTO achieves substantial reductions of up to 98% in both per-inference latency and energy consumption compared to state-of-the-art transparent methods and yields results comparable to non-transparent approaches, all without necessitating any source code modification.

Figures

Figures reproduced from arXiv: 2507.21739 by the authors.

Figure 1
Figure 1. The performance of VGG-16 under device-only inference across different mobile devices [45]–[48]. B. Non-Transparent Offloading Non-transparent offloading strategies relocate model com￾putations to GPU servers by mandating application source code modifications, a requirement that inherently increases engineering complexity and restricts application diversity. This non-transparency severely limits their use, particula… view at source ↗
Figure 2
Figure 2. Workflow of Transparent Offloading System for Model Inference in MEC. detailed steps of the transparent offloading process (depicted in the right part in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The wireless transmission instability of TCP between our robot and the base station in MEC networks. measurements exhibiting higher fluctuations and occasional near-zero drops due to obstacles and reduced signal reflections. This indicates that transparent offloading is hard to sustain in real-world wireless networks, as the RTT for operator-level RPCs in traditional transparent offloading systems is usually higher … view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Architecture of RRTO, with key components highlighted in red boxes. When RRTO intercepts CUDA kernel function calls originat￾ing from upper-layer ML applications, its recorder component logs these invoked functions, including their parameters and return values. Concurr…
Figure 5
Figure 5. Figure 5: Illustration of Operator Sequence Search. single missing or extra operator disrupts the end-to-end data flow and produces incorrect results. To achieve this under realistic conditions, Operator Sequence Search must meet three fundamental challenges: i) Transparency: RR…
Figure 6
Figure 6. Figure 6: Workflow of RRTO during the replaying phase. To address the communication costs in the transparent offloading systems, RRTO introduces an automatic recording and replay mechanism. Given that ML models in mobile appli￾cations often exhibit static and predictable sequenc…
Figure 7
Figure 7. Figure 7: The detailed composition of the robot platform. inference communication standby Energy (Watt) 13.35 4.25 4.04 TABLE II: Power draw (Watt) of our robot in different states [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: Kapao [17], a real-time people-tracking application on our four-wheeled robot with a CNN-based human keypoint detection model. functions in Cricket, allowing for seamless integration and efficient operation of the record/replay mechanism. Hardware. The evaluation was c…
Figure 10
Figure 10. Figure 10: Performance of Kapao in different environments with various systems. In terms of inference time, RRTO reduced inference time by an average of 72% compared to local computation and 95% compared to Cricket in the indoors scenario; the reductions in the outdoors scenario…
Figure 11
Figure 11. Figure 11: Semi-RRTO: only applying Caching [63] specifically to the RPCs of “cudaGetDevice” and “cudaGetLastError” in RRTO, effectively eliminating their transmission requirements. speeds observed with NNTO. This is evident from the fact that “cudaLaunchKernel” still represents…
Figure 12
Figure 12. Figure 12: Performance of Torchvision models in different environments with various systems. entail complex logic and branching are generally more suited for CPU rather than GPU execution [90], and optimizing inference for DAMs with changing operator sequences remains a pervasiv…

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Physically-Induced Atmospheric Adversarial Perturbations: Enhancing Transferability and Robustness in Remote Sensing Image Classification

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    FogFool creates fog-based adversarial perturbations using Perlin noise optimization to achieve high black-box transferability (83.74% TASR) and robustness to defenses in remote sensing classification.

  2. Dual-Envelope Constrained Nonlinear MPC for Distributed Drive Electric Vehicles Drifting Under Bounded Steering and Direct Yaw-Moment Control

    eess.SY 2026-04 unverdicted novelty 6.0 of 10

    The extended dual-envelope NMPC enables smoother drifting convergence and cuts steady-state tracking errors in speed, sideslip angle, and yaw rate by 33%, 71%, and 31% respectively in hardware tests.

  3. SL-FAC: A Communication-Efficient Split Learning Framework with Frequency-Aware Compression

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    SL-FAC reduces communication in split learning via frequency-aware compression of activations and gradients while aiming to preserve training-critical information.

  4. SwarmSense-DNN: A Trustworthy and Decentralized Neural Framework for Proactive Anomaly Defense in Consumer IoT

    cs.CR 2026-06 unverdicted novelty 3.0 of 10

    SwarmSense-DNN is a proposed decentralized neural framework that integrates swarm intelligence with hierarchical federated learning and graph neural networks to achieve 95.44% anomaly detection accuracy and 67% reduce...

Reference graph

Works this paper leans on

92 extracted references · 63 canonical work pages · cited by 4 Pith papers

  1. [27]

    Intra-dp: A high performance collabora- tive inference system for mobile edge computing,

    Z. Sun, X. Guan, Z. Lin, Z. Fang, X. Cai, Z. Chen, F. Liu, H. Cui, J. Xiong, W. Ni et al. , “Intra-dp: A high performance collabora- tive inference system for mobile edge computing,” arXiv preprint arXiv:2507.05829, 2025

  2. [1]

    Binarized Neural Network for Edge Intelligence of Sensor-Based Human Activity Recognition,

    F. Luo, S. Khan, Y . Huang, and K. Wu, “Binarized Neural Network for Edge Intelligence of Sensor-Based Human Activity Recognition,” IEEE Trans. Mobile Comput. , vol. 22, no. 3, pp. 1356–1368, Mar. 2023

  3. [2]

    MERIT: Multimodal Wearable Vital Sign Waveform Monitoring,

    Y . Tang, Z. Chen, A. Li, T. Zheng, Z. Lin, J. Xu, P. Lv, Z. Sun, and Y . Gao, “MERIT: Multimodal Wearable Vital Sign Waveform Monitoring,” arXiv preprint arXiv:2410.00392 , 2024

  4. [3]

    DNN Patching: Progressive Fixing and Augmenting the Functionalities of DNNs for Autonomous Vehicles,

    A. N. Saridena and A. Choromanska, “DNN Patching: Progressive Fixing and Augmenting the Functionalities of DNNs for Autonomous Vehicles,” IEEE Robot. Autom. Lett. , vol. 7, no. 2, pp. 3257–3264, Apr. 2022

  5. [4]

    Channel Power Gain Estimation for Terahertz Vehicle-to-Infrastructure Networks,

    Z. Lin, L. Wang, J. Ding, B. Tan, and S. Jin, “Channel Power Gain Estimation for Terahertz Vehicle-to-Infrastructure Networks,”IEEE Commun. Lett., vol. 27, no. 1, pp. 155–159, 2022

  6. [5]

    IC3M: In-Car Multimodal Multi-Object Monitoring for Abnormal Status of Both Driver and Passengers,

    Z. Fang, Z. Lin, S. Hu, H. Cao, Y . Deng, X. Chen, and Y . Fang, “IC3M: In-Car Multimodal Multi-Object Monitoring for Abnormal Status of Both Driver and Passengers,” arXiv preprint arXiv:2410.02592 , 2024

  7. [6]

    Fedsn: A federated learning framework over heterogeneous leo satellite networks,

    Z. Lin, Z. Chen, Z. Fang, X. Chen, X. Wang, and Y . Gao, “Fedsn: A federated learning framework over heterogeneous leo satellite networks,” IEEE Transactions on Mobile Computing , 2024

  8. [7]

    A Novel Framework of Three-Hierarchical Offloading Optimization for MEC in Industrial IoT Networks,

    Z. Zhao, R. Zhao, J. Xia, X. Lei, D. Li, C. Yuen, and L. Fan, “A Novel Framework of Three-Hierarchical Offloading Optimization for MEC in Industrial IoT Networks,” IEEE Trans. Ind. Informat., vol. 16, no. 8, pp. 5424–5434, 2019

Show all 92 references
  1. [8]

    SigChord: Sniffing Wide Non-Sparse Multiband Signals for Terrestrial and Non- Terrestrial Wireless Networks,

    J. Peng, J. Duan, Z. Lin, H. Yuan, Y . Gao, and Z. Chen, “SigChord: Sniffing Wide Non-Sparse Multiband Signals for Terrestrial and Non- Terrestrial Wireless Networks,” arXiv preprint arXiv:2504.06587, 2025

  2. [9]

    Constructing 4D Radio Map in LEO Satellite Networks with Limited Samples,

    H. Yuan, Z. Chen, Z. Lin, J. Peng, Y . Zhong, X. Hu, S. Xue, W. Li, and Y . Gao, “Constructing 4D Radio Map in LEO Satellite Networks with Limited Samples,” IEEE INFOCOM, 2025

  3. [10]

    RF- Based Human Activity Recognition Using Signal Adapted Convolutional Neural Network,

    Z. Chen, C. Cai, T. Zheng, J. Luo, J. Xiong, and X. Wang, “RF- Based Human Activity Recognition Using Signal Adapted Convolutional Neural Network,” IEEE Trans. Mobile Comput., vol. 22, no. 1, pp. 487– 499, 2021

  4. [11]

    Spatial-spectral terahertz networks,

    Z. Lin, L. Wang, B. Tan, and X. Li, “Spatial-spectral terahertz networks,” IEEE Transactions on Wireless Communications , vol. 21, no. 6, pp. 3881–3892, 2021

  5. [12]

    Fedac: An adaptive clustered federated learning framework for heterogeneous data,

    Y . Zhang, H. Chen, Z. Lin, Z. Chen, and J. Zhao, “Fedac: An adaptive clustered federated learning framework for heterogeneous data,” arXiv preprint arXiv:2403.16460, 2024

  6. [13]

    SatSense: Multi-Satellite Collaborative Framework for Spectrum Sensing,

    H. Yuan, Z. Chen, Z. Lin, J. Peng, Z. Fang, Y . Zhong, Z. Song, and Y . Gao, “SatSense: Multi-Satellite Collaborative Framework for Spectrum Sensing,” IEEE Trans. Cogn. Commun. Netw. , 2025

  7. [14]

    SUMS: Sniffing Unknown Multiband Signals under Low Sampling Rates,

    J. Peng, Z. Chen, Z. Lin, H. Yuan, Z. Fang, L. Bao, Z. Song, Y . Li, J. Ren, and Y . Gao, “SUMS: Sniffing Unknown Multiband Signals under Low Sampling Rates,” IEEE Trans. Mobile Comput. , 2024

  8. [15]

    LEO Satellite Networks Assisted Geo-Distributed Data Processing,

    Z. Zhao, Z. Chen, Z. Lin, W. Zhu, K. Qiu, C. You, and Y . Gao, “LEO Satellite Networks Assisted Geo-Distributed Data Processing,” IEEE Wireless Commun. Lett., 2024

  9. [16]

    Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,

    Z. Lin, W. Wei, Z. Chen, C.-T. Lam, X. Chen, Y . Gao, and J. Luo, “Hi- erarchical Split Federated Learning: Convergence Analysis and System Optimization,” IEEE Trans. Mobile Comput. , 2025

  10. [17]

    Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi- Person Human Pose Estimation,

    W. McNally, K. Vats, A. Wong, and J. McPhee, “Rethinking Keypoint Representations: Modeling Keypoints and Poses as Objects for Multi- Person Human Pose Estimation,” in Proc. ECCV, Oct. 2022, pp. 37–54

  11. [18]

    Efficient Parallel Split Learning over Resource-Constrained Wireless Edge Networks,

    Z. Lin, G. Zhu, Y . Deng, X. Chen, Y . Gao, K. Huang, and Y . Fang, “Efficient Parallel Split Learning over Resource-Constrained Wireless Edge Networks,” IEEE Trans. Mobile Comput. , vol. 23, no. 10, pp. 9224–9239, 2024

  12. [19]

    AGRNav: Efficient and Energy-Saving Au- tonomous Navigation for Air-Ground Robots in Occlusion-Prone En- vironments,

    J. Wang, Z. Sun, X. Guan, T. Shen, Z. Zhang, T. Duan, D. Huang, S. Zhao, and H. Cui, “AGRNav: Efficient and Energy-Saving Au- tonomous Navigation for Air-Ground Robots in Occlusion-Prone En- vironments,” in Proc. ICRA, May 2024, pp. 4494–4501

  13. [20]

    Rethinking adversarial attacks in reinforcement learning from policy distribution perspective,

    T. Duan, Z. Zhang, Z. Lin, Y . Gao, L. Xiong, Y . Cui, H. Liang, X. Chen, H. Cui, and D. Huang, “Rethinking adversarial attacks in reinforcement learning from policy distribution perspective,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Pr...

  14. [21]

    Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,

    Z. Lin, G. Qu, Q. Chen, X. Chen, Z. Chen, and K. Huang, “Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities,” arXiv preprint arXiv:2309.16739 , 2023

  15. [22]

    State-aware perturbation optimization for robust deep reinforcement learning,

    Z. Zhang, T. Duan, Z. Lin, D. Huang, Z. Fang, Z. Sun, L. Xiong, H. Liang, H. Cui, and Y . Cui, “State-aware perturbation optimization for robust deep reinforcement learning,” arXiv preprint arXiv:2503.20613 , 2025

  16. [23]

    HASFL: Heterogeneity- Aware Split Federated Learning over Edge Computing Systems,

    Z. Lin, Z. Chen, X. Chen, W. Ni, and Y . Gao, “HASFL: Heterogeneity- Aware Split Federated Learning over Edge Computing Systems,” arXiv preprint arXiv:2506.08426, 2025

  17. [24]

    MonoScene: Monocular 3D Semantic Scene Completion,

    A.-Q. Cao and R. de Charette, “MonoScene: Monocular 3D Semantic Scene Completion,” in Proc. CVPR, Jun. 2022, pp. 3991–4001

  18. [25]

    Optimal resource allocation for u-shaped parallel split learning,

    S. Lyu, Z. Lin, G. Qu, X. Chen, X. Huang, and P. Li, “Optimal resource allocation for u-shaped parallel split learning,” in 2023 IEEE Globecom Workshops (GC Wkshps), 2023, pp. 197–202

  19. [26]

    SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models,

    Z. Lin, X. Hu, Y . Zhang, Z. Chen, Z. Fang, X. Chen, A. Li, P. Vepakomma, and Y . Gao, “SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models,” arXiv preprint arXiv:2407.00952, 2024

  20. [28]

    Adaptsfl: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,

    Z. Lin, G. Qu, W. Wei, X. Chen, and K. K. Leung, “Adaptsfl: Adaptive Split Federated Learning in Resource-Constrained Edge Networks,” IEEE Trans. Netw., 2024

  21. [29]

    Mobile Edge Computing: A Survey,

    N. Abbas, Y . Zhang, A. Taherkordi, and T. Skeie, “Mobile Edge Computing: A Survey,” IEEE Internet Things J. , vol. 5, no. 1, pp. 450– 465, Feb. 2018

  22. [30]

    The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSL,

    B. Qiao, M. A. ¨Ozkan, J. Teich, and F. Hannig, “The Best of Both Worlds: Combining CUDA Graph with an Image Processing DSL,” in Proc. DAC, 2020, pp. 1–6

  23. [31]

    Cricket: A Virtual- ization Layer for Distributed Execution of CUDA Applications with Checkpoint/Restart Support,

    N. Eiling, J. Baude, S. Lankes, and A. Monti, “Cricket: A Virtual- ization Layer for Distributed Execution of CUDA Applications with Checkpoint/Restart Support,” Concurrency Comput. Pract. Exp., vol. 34, no. 14, p. e6474, May 2022

  24. [32]

    TensorRT Inference with TensorFlow,

    P. Davoodi, C. Gwon, G. Lai, and T. Morris, “TensorRT Inference with TensorFlow,” in Proc. GPU Technol. Conf. (GTC) , Mar. 2019

  25. [33]

    A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture,

    X. Yi, “A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture,” arXiv preprint arXiv:2409.10661 , 2024

  26. [34]

    Infer-EDGE: Dynamic DNN Inference Optimization in ‘Just-in-Time’ Edge-AI Implementations,

    M. Mounesan, X. Zhang, and S. Debroy, “Infer-EDGE: Dynamic DNN Inference Optimization in ‘Just-in-Time’ Edge-AI Implementations,” arXiv preprint arXiv:2501.18842 , 2025

  27. [35]

    TorchScript: Optimized Execution of PyTorch Programs,

    Z. DeVito, “TorchScript: Optimized Execution of PyTorch Programs,” Retrieved January, 2022

  28. [36]

    PyTorch: An Imperative Style, High- Performance Deep Learning Library,

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “PyTorch: An Imperative Style, High- P...

  29. [37]

    Implementing Open-Source CUDA Runtime,

    S. Kato, “Implementing Open-Source CUDA Runtime,” in Proc. of the 54the Programming Symposium, Hakone, Japan, Jan. 2013, pp. 111–118

  30. [38]

    Evaluation of Communication Technologies for Distributed Industrial Control Sys- tems: Concept and Evaluation of 5G and WiFi 6,

    M. Schneider, F. Haag, A. K. Khalil, and D. A. Breunig, “Evaluation of Communication Technologies for Distributed Industrial Control Sys- tems: Concept and Evaluation of 5G and WiFi 6,” Procedia CIRP, vol. 107, pp. 588–593, 2022

  31. [39]

    TensorFlow: A System for Large-Scale Machine Learning,

    M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard et al. , “TensorFlow: A System for Large-Scale Machine Learning,” in Proc. OSDI , Nov. 2016, pp. 265– 283

  32. [40]

    The OpenCL Specification,

    A. Munshi, “The OpenCL Specification,” in Proc. IEEE Hot Chips 21 Symp. (HCS), Aug. 2009

  33. [41]

    CoDL: Efficient CPU-GPU Co-Execution for Deep Learning Inference on Mobile Devices,

    F. Jia, D. Zhang, T. Cao, S. Jiang, Y . Liu, J. Ren, and Y . Zhang, “CoDL: Efficient CPU-GPU Co-Execution for Deep Learning Inference on Mobile Devices,” in Proc. MobiSys, Jun. 2022, pp. 209–221

  34. [42]

    Accelerating Robot Dynamics Gradients on a CPU, GPU, and FPGA,

    B. Plancher, S. M. Neuman, T. Bourgeat, S. Kuindersma, S. Devadas, and V . J. Reddi, “Accelerating Robot Dynamics Gradients on a CPU, GPU, and FPGA,” IEEE Robot. Autom. Lett. , vol. 6, no. 2, pp. 2335– 2342, Apr. 2021. 15

  35. [43]

    DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices,

    N. D. Lane, S. Bhattacharya, P. Georgiev, C. Forlivesi, L. Jiao, L. Qen- dro, and F. Kawsar, “DeepX: A Software Accelerator for Low-Power Deep Learning Inference on Mobile Devices,” in Proc. IPSN, Apr. 2016, pp. 1–12

  36. [44]

    REACT: Streaming Video Analytics on the Edge with Asynchronous Cloud Support,

    A. Ghosh, S. Iyengar, S. Lee, A. Rathore, and V . N. Padmanabhan, “REACT: Streaming Video Analytics on the Edge with Asynchronous Cloud Support,” in Proc. IoTDI, May 2023, pp. 222–235

  37. [45]

    MLPerf Mobile Benchmarks,

    “MLPerf Mobile Benchmarks,” https://mlcommons.org/working-groups/ benchmarks/mobile/, 2023

  38. [46]

    Raspberry Pi,

    RaspberryPi, “Raspberry Pi,” https://www.raspberrypi.com/ documentation/computers/configuration.html\#power-consumption, 2022

  39. [47]

    Jetson Xavier NX Series: The World’s Smallest AI Supercomputer,

    NVIDIA, “Jetson Xavier NX Series: The World’s Smallest AI Supercomputer,” https://www.nvidia.com/en-us/autonomous-machines/ embedded-systems/jetson-xavier-nx-developer-kit/, 2024

  40. [48]

    Power Consumption Benchmark for Embedded AI Inference,

    Z. Ning, M. Vandersteegen, K. Van Beeck, T. Goedem ´e, and P. Vande- walle, “Power Consumption Benchmark for Embedded AI Inference,” in Proc. Int. Conf. Appl. Comput. WWW/Internet (AC) , Jan. 2024, pp. 3–10

  41. [49]

    Local Feature Selection for Data Classification,

    N. Armanfard, J. P. Reilly, and M. Komeili, “Local Feature Selection for Data Classification,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 38, no. 6, pp. 1217–1227, 2015

  42. [50]

    Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,

    Y . Kang, J. Hauswald, C. Gao, A. Rovinski, T. Mudge, J. Mars, and L. Tang, “Neurosurgeon: Collaborative Intelligence Between the Cloud and Mobile Edge,” in Proc. ASPLOS, Apr. 2017, pp. 615–629

  43. [51]

    Computation Offloading Toward Edge Computing,

    L. Lin, X. Liao, H. Jin, and P. Li, “Computation Offloading Toward Edge Computing,” Proc. IEEE , vol. 107, no. 8, pp. 1584–1607, Aug. 2019

  44. [52]

    QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural Networks,

    Z. Fang, T. Yu, O. J. Mengshoel, and R. K. Gupta, “QoS-Aware Scheduling of Heterogeneous Servers for Inference in Deep Neural Networks,” in Proc. CIKM, Nov. 2017, pp. 2067–2070

  45. [53]

    InfiniBand Networking Solutions,

    NVIDIA, “InfiniBand Networking Solutions,” https://www.nvidia.com/ en-us/networking/products/infiniband/, 2024

  46. [54]

    A First Look at Wi-Fi 6 in Action: Throughput, Latency, Energy Efficiency, and Security,

    R. Liu and N. Choi, “A First Look at Wi-Fi 6 in Action: Throughput, Latency, Energy Efficiency, and Security,” in Proc. ACM Meas. Anal. Comput. Syst., vol. 7, no. 1, Mar. 2023, pp. 1–25

  47. [55]

    Mobile Access Bandwidth in Practice: Measurement, Analysis, and Implications,

    X. Yang, H. Lin, Z. Li, F. Qian, X. Li, Z. He, X. Wu, X. Wang, Y . Liu, Z. Liao et al. , “Mobile Access Bandwidth in Practice: Measurement, Analysis, and Implications,” in Proc. SIGCOMM, Aug. 2022, pp. 114– 128

  48. [56]

    Throughput Comparison between The New HEW 802.11 ax Standard and 802.11 n/ac Standards in Selected Distance Windows,

    A. Masiukiewicz, “Throughput Comparison between The New HEW 802.11 ax Standard and 802.11 n/ac Standards in Selected Distance Windows,” Int. J. Electron. Telecommun. , vol. 65, no. 1, pp. 79–84, 2019

  49. [57]

    Performance Impact of LoS and NLoS Transmissions in Dense Cellular Networks,

    M. Ding, P. Wang, D. L ´opez-P´erez, G. Mao, and Z. Lin, “Performance Impact of LoS and NLoS Transmissions in Dense Cellular Networks,” IEEE Trans. Wireless Commun. , vol. 15, no. 3, pp. 2365–2380, Mar. 2016

  50. [58]

    Proportional and Preemption-Enabled Traffic Offloading for IP Flow Mobility: Algo- rithms and Performance Evaluation,

    Y . Ren, C.-W. Tung, J.-C. Chen, and F. Y . Li, “Proportional and Preemption-Enabled Traffic Offloading for IP Flow Mobility: Algo- rithms and Performance Evaluation,” IEEE Trans. Veh. Technol., vol. 67, no. 12, pp. 12 095–12 108, Dec. 2018

  51. [59]

    iPerf - Download iPerf3 and Original iPerf Pre-Compiled Binaries,

    “iPerf - Download iPerf3 and Original iPerf Pre-Compiled Binaries,” https://iperf.fr/iperf-download.php, 2024

  52. [60]

    TCP in Wireless Environments: Problems and Solutions,

    Y . Tian, K. Xu, and N. Ansari, “TCP in Wireless Environments: Problems and Solutions,” IEEE Commun. Mag., vol. 43, no. 3, pp. S27– S32, Mar. 2005

  53. [61]

    NVIDIA Nsight Compute Documentation,

    NVIDIA, “NVIDIA Nsight Compute Documentation,” https://docs. nvidia.com/nsight-compute/index.html, 2024

  54. [62]

    libtirpc: Transport Independent RPC library,

    S. Dickson, “libtirpc: Transport Independent RPC library,” https://git. linux-nfs.org/?p=steved/libtirpc.git, 2024

  55. [63]

    Cliquemap: Productionizing an RMA-Based Distributed Caching Sys- tem,

    A. Singhvi, A. Akella, M. Anderson, R. Cauble, H. Deshmukh, D. Gib- son, M. M. Martin, A. Strominger, T. F. Wenisch, and A. Vahdat, “Cliquemap: Productionizing an RMA-Based Distributed Caching Sys- tem,” in Proc. SIGCOMM. ACM, 2021, pp. 93–105

  56. [64]

    Dagger: Efficient and Fast RPCs in Cloud Microservices With Near-Memory Reconfigurable NICs,

    N. Lazarev, S. Xiang, N. Adit, Z. Zhang, and C. Delimitrou, “Dagger: Efficient and Fast RPCs in Cloud Microservices With Near-Memory Reconfigurable NICs,” in Proc. ASPLOS. ACM, 2021, pp. 36–51

  57. [65]

    Efficient Asynchronous RPC Calls for Mi- croservices: DeathStarBench Study,

    S. Eyerman and I. Hur, “Efficient Asynchronous RPC Calls for Mi- croservices: DeathStarBench Study,” arXiv preprint arXiv:2209.13265 , 2022

  58. [66]

    Interpretable and Accurate Fine-Grained Recog- nition via Region Grouping,

    Z. Huang and Y . Li, “Interpretable and Accurate Fine-Grained Recog- nition via Region Grouping,” in Proc. CVPR, 2020, pp. 8662–8672

  59. [67]

    GMC: Graph-Based Multi-View Clustering,

    H. Wang, Y . Yang, and B. Liu, “GMC: Graph-Based Multi-View Clustering,” IEEE Trans. Knowledge Data Eng., vol. 32, no. 6, pp. 1116– 1129, 2019

  60. [68]

    Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. K ¨uttler, M. Lewis, W.-t. Yih, T. Rockt ¨aschel et al. , “Retrieval- Augmented Generation for Knowledge-Intensive NLP Tasks,” in Proc. NeurIPS, vol. 33, 2020, pp. 9459–9474

  61. [69]

    BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint arXiv:1810.04805, Oct. 2018

  62. [70]

    An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling,

    S. Bai, J. Z. Kolter, and V . Koltun, “An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling,” arXiv preprint arXiv:1803.01271, 2018

  63. [71]

    An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in Proc. ICLR, 2020

  64. [72]

    Language Models Are Few-Shot Learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language Models Are Few-Shot Learners,” in Proc. NeurIPS , vol. 33, 2020, pp. 1877– 1901

  65. [73]

    Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,

    N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean, “Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer,” in Proc. ICLR, 2017

  66. [74]

    Sequence to Sequence Learning with Neural Networks,

    I. Sutskever, O. Vinyals, and Q. V . Le, “Sequence to Sequence Learning with Neural Networks,” in Proc. NeurIPS, vol. 27, 2014

  67. [75]

    XGBoost: A Scalable Tree Boosting System,

    T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proc. KDD, 2016, pp. 785–794

  68. [76]

    Semi-Supervised Learning with Graph Learning-Convolutional Networks,

    B. Jiang, Z. Zhang, D. Lin, J. Tang, and B. Luo, “Semi-Supervised Learning with Graph Learning-Convolutional Networks,” in Proc. CVPR, 2019, pp. 11 313–11 320

  69. [77]

    Efficient Algorithms for Device Placement of DNN Graph Operators,

    J. M. Tarnawski, A. Phanishayee, N. Devanur, D. Mahajan, and F. N. Paravecino, “Efficient Algorithms for Device Placement of DNN Graph Operators,” in Proc. NeurIPS, H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33, Dec. 2020, pp. 15 451–15 463

  70. [78]

    Faster Algorithms for Longest Common Substring,

    P. Charalampopoulos, T. Kociumaka, S. P. Pissis, and J. Radoszewski, “Faster Algorithms for Longest Common Substring,” arXiv preprint arXiv:2105.03106, 2021

  71. [79]

    Torchvision the Machine-Vision Package of Torch,

    S. Marcel and Y . Rodriguez, “Torchvision the Machine-Vision Package of Torch,” in Proc. ACM Multimedia, Oct. 2010, pp. 1485–1488

  72. [80]

    Deep Residual Learning for Image Recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” arXiv preprint arXiv:1512.03385 , Dec. 2015

  73. [81]

    A ConvNet for the 2020s,

    Z. Liu, H. Mao, C.-Y . Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” arXiv preprint arXiv:2201.03545 , Jan. 2022

  74. [82]

    Fully Convolutional Networks for Semantic Segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” arXiv preprint arXiv:1411.4038 , Nov. 2015

  75. [83]

    Rethinking Atrous Convolution for Semantic Image Segmentation,

    L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking Atrous Convolution for Semantic Image Segmentation,” arXiv preprint arXiv:1706.05587, Jun. 2017

  76. [84]

    Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards Real- Time Object Detection with Region Proposal Networks,” arXiv preprint arXiv:1506.01497, Jun. 2015

  77. [85]

    Focal Loss for Dense Object Detection,

    T.-Y . Lin, P. Goyal, R. Girshick, K. He, and P. Doll ´ar, “Focal Loss for Dense Object Detection,” arXiv preprint arXiv:1708.02002, Aug. 2017

  78. [86]

    Datacenter Traffic Control: Understanding Techniques and Tradeoffs,

    M. Noormohammadpour and C. S. Raghavendra, “Datacenter Traffic Control: Understanding Techniques and Tradeoffs,” IEEE Commun. Surv. Tutor., vol. 20, no. 2, pp. 1492–1525, Jun. 2018

  79. [87]

    ultralytics/yolov5: v7.0 - YOLOv5 SOTA Realtime Instance Segmentation,

    G. Jocher, A. Chaurasia, A. Stoken, J. Borovec, NanoCode012, Y . Kwon, K. Michael, T. Xie, J. Fang, imyhxy, Lorna, Z. Yifu, C. Wong, A. V , D. Montes, Z. Wang, C. Fati, J. Nadar, Laughing, UnglvKitDe, V . Sonck, tkianai, yxNONG, P. Skalski, A. Hogan, D. Nair, M. Strobel, and M...

  80. [88]

    pynvml: Python Bindings for the NVIDIA Management Library,

    NVIDIA, “pynvml: Python Bindings for the NVIDIA Management Library,” https://developer.nvidia.com/management-library-nvml/, 2024

  81. [89]

    Optimal Path Planning of Au- tonomous Marine Vehicles in Stochastic Dynamic Ocean Flows Using a GPU-Accelerated Algorithm,

    R. Chowdhury and D. Subramani, “Optimal Path Planning of Au- tonomous Marine Vehicles in Stochastic Dynamic Ocean Flows Using a GPU-Accelerated Algorithm,” IEEE J. Ocean. Eng. , vol. 47, no. 4, pp. 864–879, Oct. 2022

  82. [90]

    Query Processing on Hetero- geneous CPU/GPU Systems,

    V . Rosenfeld, S. Breß, and V . Markl, “Query Processing on Hetero- geneous CPU/GPU Systems,” ACM Comput. Surv. , vol. 55, no. 1, pp. 1–38, Jan. 2023

  83. [91]

    VecQ: Minimal Loss DNN Model Compression with Vectorized Weight Quantization,

    C. Gong, Y . Chen, Y . Lu, T. Li, C. Hao, and D. Chen, “VecQ: Minimal Loss DNN Model Compression with Vectorized Weight Quantization,” IEEE Trans. Comput. , vol. 70, no. 5, pp. 696–710, May 2021

  84. [92]

    Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks,

    L. Wang and K.-J. Yoon, “Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 44, no. 6, pp. 3048–3068, Jun. 2022

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.