REVIEW 4 major objections 4 minor 47 references
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper claims that a single SYCL codebase, migrated from a CUDA Smith-Waterman alignment suite, matches CUDA performance on NVIDIA GPUs, reaches similar architectural efficiency on AMD and Intel GPUs in most cases, and runs across CPUs…
desk verdict A broad, useful SYCL portability dataset undercut by an unvalidated peak model and an overclaiming abstract, but the functional portability and CUDA-parity findings are solid. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the theoretical peak performance model of Eqs. 4–6, which computes each device's peak GCUPS as (Clock Rate × Throughput × Lanes) / 12, where 12 is the assumed instruction count for one Smith-Waterman cell update and Throughput is a per-architecture instruction issue factor (3 for Maxwell/Pascal, 2 for RDNA2 and GCN, 8 for Intel Xe, 1 for CPUs). Dividing the measured GCUPS by this peak gives the architectural efficiency, and averaging those efficiencies over a platform set, using the reformulated performance portability metric, yields the portability values that carry the comparison. The model is what converts raw runtimes into the normalized numbers that support every conclusion about SYCL versus CUDA and about cross-vendor parity.
What would settle it
Profile the SW# SYCL kernel on one tested GPU (for example, the RTX 3090) and count the actual instructions retired per cell update with the vendor's profiler. If the count is not 12, recompute that device's theoretical peak GCUPS with the corrected count; a corrected peak that falls below the reported achieved GCUPS (288.6 for protein search) would expose an inconsistency in the model. Alternatively, run a microbenchmark of pure add/subtract/max throughput on the same device to verify the chosen throughput factor.
Extended reading notes
Core claim
The paper's finding is that SYCL's performance portability is a measured property: for protein database search the same SYCL code achieves 42.2% average architectural efficiency across six NVIDIA GPUs (versus 42.0% for CUDA), 48.9% across four Intel GPUs, comparable rates on AMD GPUs, and 40.3% across nine CPUs from both vendors. For pairwise DNA alignment, the NVIDIA values are 53.7% for SYCL and 51.4% for CUDA. In hybrid CPU-GPU configurations, functional portability is complete but performance portability falls to 19.3%, and the paper identifies the cause: the SW# scheduler distributes query sequences without considering each device's speed, leaving the faster device idle. The authors also show that a scalar kernel used for long sequences defeats compiler vectorization, explaining why CPU architectural efficiency for pairwise alignment drops to 5–17% while the vectorizable protein kernel reaches 21–51%. Their conclusion treats these shortfalls as fixable workload and code issues, not fundamental SYCL limits.
Load-bearing premise
The results stand or fall on the assumed theoretical peaks: if the fixed 12-instruction count per cell update or the per-vendor throughput factors are wrong, every architectural efficiency and portability percentage changes, which could overturn the conclusion that SYCL is equally efficient across vendors.
Editorial extensions
If this is right
- A single SYCL source can replace separate CUDA, HIP, and Level Zero code paths for Smith-Waterman workloads, with no more than a small portability loss on NVIDIA hardware.
- The protein benchmark's CPU portability of 40.3% shows that SYCL can target desktop, server, and mobile CPUs from Intel and AMD without per-vendor rewrites, provided the kernel vectorizes.
- Device-aware scheduling would immediately improve multi-GPU and CPU-GPU results, since the paper traces those losses to workload distribution rather than SYCL.
- Programmers targeting both CPUs and GPUs must keep kernels vectorizable; the pairwise long-sequence kernel's scalar code is the concrete cause of the 11.5% CPU portability figure.
- The paper's own planned optimizations, such as instruction reordering and lower-precision integers, would change the cell-update instruction count, so future kernel versions will need new peak estimates.
Reading between the lines
- If the assumed 12-instruction count is off, the absolute portability values shift, but the ranking between CUDA and SYCL on NVIDIA hardware is robust because both are measured with the same runtime and the same model.
- The gap between Intel's Gen iGPUs (up to 75.7% efficiency) and Xe GPUs (23.4%) suggests driver maturity rather than SYCL overhead limits Intel's showing, so future driver releases could materially raise Intel's portability numbers.
- The same measurement recipe could be applied to other dynamic-programming kernels with the same dependency pattern, such as edit distance or alignment with different gap penalties, providing a cheap test of whether SYCL's portability generalizes beyond SW#.
- The paper itself cautions that the CPU comparison rests on a small sample (one AMD CPU), so the Intel-versus-AMD portability difference should not be over-read.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates SYCL performance portability for a Smith-Waterman sequence alignment code (SW#) across a wide range of hardware: six NVIDIA GPUs, two AMD GPUs, four Intel GPUs, nine CPUs, and several multi-GPU and CPU-GPU hybrid configurations. The study compares a CUDA implementation against a pure SYCL port on NVIDIA hardware, and reports architectural efficiency and performance portability metrics (Pennycook/Marowka) for the other platforms. The main claims are that CUDA and SYCL perform comparably on NVIDIA GPUs, that SYCL achieves similar architectural efficiency on AMD and Intel GPUs in most cases, that SYCL runs effectively on CPUs including Intel's hybrid architectures, and that CPU-GPU hybrid configurations are functionally portable but suffer from workload-distribution limitations.
Significance. The manuscript assembles an unusually broad experimental matrix and provides a valuable head-to-head CUDA/SYCL comparison on six NVIDIA GPUs with paired runs; the functional portability evidence across many devices and vendors is credible and the code is made publicly available. However, the quantitative portability conclusions, especially the cross-vendor architectural efficiencies and the CPU results, rest on a theoretical peak model (Eqs. 4-6) whose per-device throughput values and 12-instruction cell update are asserted rather than validated. Until that denominator is validated or the claims are reframed in model-independent terms (e.g., raw GCUPS), the strong cross-vendor efficiency statements are conditional. If the model issue is resolved, the paper would be a useful reference for SYCL adoption in heterogeneous bioinformatics workloads.
major comments (4)
- [Section 3.2, Eqs. (4)-(6), Tables 1-2] The theoretical peak model is the basis for every architectural efficiency and portability value in Tables 4-13, but the per-vendor instruction throughput values are assigned without validation. In particular, Table 2 sets CPU instruction throughput to 1 for every core, yet modern Intel and AMD cores can issue two or more vector integer add/sub/max instructions per cycle. If the true throughput is higher, the CPU architectural efficiencies in Tables 9-11 are inflated and the claim of 'remarkable versatility and effectiveness across CPUs' is not quantitatively supported. Please validate the throughput values with microbenchmarks or authoritative documentation, or restrict the portability conclusions to raw GCUPS and the NVIDIA CUDA/SYCL comparison.
- [Section 3.2.1, Algorithm 1, Eq. (6)] The model assumes exactly 12 instructions per cell update for every architecture and both kernels, but the generated instruction stream depends on vectorization, predication, and compiler scheduling, and memory, address, and loop overhead is ignored. Eq. (6) is therefore an algorithmic abstraction rather than a real device peak, and efficiency values are not directly comparable across vendors. The paper should either demonstrate that the 12-instruction count is representative on each device (e.g., by inspecting SASS/assembly or running microbenchmarks) or explicitly discuss this as a limitation preventing cross-vendor efficiency comparisons.
- [Section 4.4.1, Table 9 and Figure 2] The conclusion that SYCL shows 'remarkable versatility and effectiveness across CPUs' is based on architectural efficiencies computed from the unvalidated CPU peak. With a realistic throughput greater than 1 instruction/cycle, the efficiencies would be substantially lower, and the comparison of 40.3% CPU portability with 42.2% NVIDIA GPU portability would not be supported by the presented numbers. This is a load-bearing issue for the CPU contribution, not a minor presentation concern.
- [Section 4.5, Tables 12-13] The hybrid CPU-GPU efficiency and portability values are equally model-dependent. For example, the 6% efficiency for the Xeon E5-1620 v3 + RX 6700 XT combination in Table 12 and the 6% portability in Table 13 are driven by the same unvalidated denominator. The discussion attributing these results mainly to workload distribution would be more convincing if the model were independently validated or if the analysis were repeated with raw GCUPS.
minor comments (4)
- [Section 4.1] Each test was run 20 times and averaged, but no variance, standard deviation, or min/max is reported; since the CUDA/SYCL differences in Table 4 are often around 1-2%, error bars or statistical tests would materially strengthen the parity claim.
- [Table 1] The #Lanes row appears to contain only five values for the 13 GPUs, and the 'Inst. throu.' row is misaligned; please check the table formatting for readability and correctness.
- [Table 11] The Xeon E5-1620 v3 row lists '0.9 3.9 9.7%' with a peak of 0.9, but Table 2 gives a peak of 9.3; the columns appear to be swapped or mistyped, and the efficiency value is inconsistent with Table 9.
- [General] Reference [19] is spelled 'Penycook' in the text; the author's name is 'Pennycook', and the reference formatting for [26] and [27] is inconsistent with the rest of the bibliography.
Circularity Check
No significant circularity: measured CUDA/SYCL comparison and the spec-based normalization are independent evidence.
full rationale
The paper's derivation chain does not reduce any claimed result to its own inputs. The theoretical peak in Eqs. 4-6 is assembled from published hardware specifications (clock rate, core count, SIMD lanes, and per-vendor instruction throughput) together with a stated 12-instruction cell update from Algorithm 1; no parameter is fitted to the measured GCUPS values. Achieved GCUPS are timed measurements, architectural efficiency is their ratio to the modeled peak, and portability is the Pennycook/Marowka average of those efficiencies. Thus every reported efficiency and portability number is either a measurement or a definitional normalization of a measurement, not a prediction forced by a fitted parameter. The direct CUDA/SYCL performance parity on NVIDIA GPUs is independent measured evidence and would survive even if the peak model were revised. The hand-assigned throughput values (3 for Maxwell/Pascal, 2 for RDNA2, 8 for Xe, 1 for CPUs) and the uniform 12-instruction count are modeling assumptions whose correctness affects the quantitative portability conclusions, but this is a validity/sensitivity concern, not circularity. The self-citations to the authors' prior works [5]-[7] supply the SYCL implementation and earlier context, while the current benchmark data, platform tables, and timing methodology are presented independently in this paper. No equation is defined in terms of the quantity it is used to establish, and no fitted quantity is later renamed as a prediction.
Assumptions & free parameters
free parameters (5)
- Instruction count per cell update (12 instructions) =
12
- NVIDIA instruction throughput =
2 or 3 depending on compute capability
- AMD RDNA2 instruction throughput =
2
- CPU SIMD instruction throughput =
1 per core
- Intel GPU instruction throughput =
8 (Xe) or derived from EU count
assumptions (3)
- domain assumption The theoretical peak capability of a device is well approximated by Capability = Clock Rate x Throughput x Lanes (Eq. 4).
- ad hoc to paper The 12-instruction cell update (Algorithm 1) is representative across all tested architectures and both scalar and vectorized kernels.
- standard math The performance portability metric of Pennycook et al. and Marowka is an appropriate summary statistic for cross-platform performance.
Cite this review
Pith. "Pith review of Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment." pith.science (2026). https://pith.science/paper/OK5ZOJOF
@misc{pith2026241208308,
author = {Pith},
title = {Pith review of: Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment},
year = {2026},
howpublished = {\url{https://pith.science/paper/OK5ZOJOF}},
note = {Machine review of arXiv:2412.08308}
}
read the original abstract
The high-performance computing (HPC) landscape is undergoing rapid transformation, with an increasing emphasis on energy-efficient and heterogeneous computing environments. This comprehensive study extends our previous research on SYCL's performance portability by evaluating its effectiveness across a broader spectrum of computing architectures, including CPUs, GPUs, and hybrid CPU-GPU configurations from NVIDIA, Intel, and AMD. Our analysis covers single-GPU, multi-GPU, single-CPU, and CPU-GPU hybrid setups, using two common, bioinformatic applications as a case study. The results demonstrate SYCL's versatility across different architectures, maintaining comparable performance to CUDA on NVIDIA GPUs while achieving similar architectural efficiency rates on AMD and Intel GPUs in the majority of cases tested. SYCL also demonstrated remarkable versatility and effectiveness across CPUs from various manufacturers, including the latest hybrid architectures from Intel. Although SYCL showed excellent functional portability in hybrid CPU-GPU configurations, performance varied significantly based on specific hardware combinations. Some performance limitations were identified in multi-GPU and CPU-GPU configurations, primarily attributed to workload distribution strategies rather than SYCL-specific constraints. These findings position SYCL as a promising unified programming model for heterogeneous computing environments, particularly for bioinformatic applications.
Figures
Reference graph
Works this paper leans on
- [1]
-
[2]
Shilov, Discrete GPU sales increase as Intel’s share drops to 0, Tom’s Hardware (9 2024)
A. Shilov, Discrete GPU sales increase as Intel’s share drops to 0, Tom’s Hardware (9 2024). URL https://www.tomshardware.com/pc-components/gpus/d iscrete-gpu-sales-increase-as-intels-share-drops-to-0
work page 2024
-
[3]
Norem, Intel has reportedly lost all its discrete GPU market share, Ex- tremeTech (9 2024)
J. Norem, Intel has reportedly lost all its discrete GPU market share, Ex- tremeTech (9 2024). URL https://www.extremetech.com/gaming/intel-has-repor tedly-lost-all-its-discrete-gpu-market-share
work page 2024
-
[4]
URL {https://registry.khronos.org/SYCL/specs/sycl-202 0/html/sycl-2020.html}
Khronos SYCL working group, Sycl 2020 specification (2023). URL {https://registry.khronos.org/SYCL/specs/sycl-202 0/html/sycl-2020.html}
work page 2023
-
[5]
M. Costanzo, E. Rucci, C. Garc ´ıa-S´anchez, M. Naiouf, M. Prieto-Mat´ıas, Migrating cuda to oneapi: A smith-waterman case study, in: I. Rojas, O. Valenzuela, F. Rojas, L. J. Herrera, F. Ortu ˜no (Eds.), Bioinformatics and Biomedical Engineering, Springer International Publishing, Cham, 2022, pp. 103–116. doi:10.1007/978-3-031-07802-6 9
-
[6]
M. Costanzo, E. Rucci, C. Garcia-Sanchez, M. Naiouf, M. Prieto-Matias, Comparing performance and portability between cuda and sycl for protein database search on nvidia, amd, and intel gpus, in: 2023 IEEE 35th In- ternational Symposium on Computer Architecture and High Performance Computing (SBAC-PAD), IEEE Computer Society, Los Alamitos, CA, USA, 2023, p...
-
[7]
M. Costanzo, E. Rucci, C. G. S ´anchez, M. R. Naiouf, M. Prieto, Assess- ing opportunities of sycl for biological sequence alignment on gpu-based systems, J. Supercomput. 80 (2022) 12599–12622. doi:10.1007 /s11227- 024-05907-2
work page 2022
-
[8]
H. Lan, W. Liu, Y . Liu, B. Schmidt, SWhybrid: A hybrid-parallel frame- work for large-scale protein sequence database search, in: Parallel and Distributed Processing Symposium (IPDPS), 2017 IEEE International, IEEE, 2017, pp. 42–51. doi:10.1109/IPDPS.2017.42. 12
Show all 47 references
-
[9]
D. B. Kirk, W.-m. W. Hwu, Programming Massively Parallel Processors: A Hands-on Approach, 1st Edition, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2010. doi:10.1016/C2015-0-02431-5
2010 doi
-
[10]
J. E. Stone, D. Gohara, G. Shi, Opencl: A parallel programming standard for heterogeneous computing systems, Computing in Science and Engg. 12 (3) (2010) 66–73. doi:10.1109/MCSE.2010.69
2010 doi
-
[11]
OpenMP ARB, The OpenMP Specification, https: //www.openmp.org/
-
[12]
OpenACC organization, The OpenACC Specification, https://www.openacc.org/
-
[13]
Farber, Parallel Programming with OpenACC, 1st Edition, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2016
R. Farber, Parallel Programming with OpenACC, 1st Edition, Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 2016
2016
-
[15]
E. F. D. O. Sandes, A. Boukerche, A. C. M. A. D. Melo, Parallel optimal pairwise biological sequence comparison: Algorithms, plat- forms, and classification, ACM Comput. Surv. 48 (4) (Mar. 2016). doi:10.1145/2893488. URL https://doi.org/10.1145/2893488
2016 doi
-
[16]
T. F. Smith, M. S. Waterman, Identification of common molecular subsequences, Journal of Molecular Biology 147 (1) (1981) 195–197. doi:10.1016/0022-2836(81)90087-5
1981 doi
-
[17]
Gotoh, An improved algorithm for matching biological sequences, in: Journal of Molecular Biology, V ol
O. Gotoh, An improved algorithm for matching biological sequences, in: Journal of Molecular Biology, V ol. 162, 1981, pp. 705–708. doi:10.1016/0022-2836(82)90398-9
1981 doi
-
[18]
Rognes, Faster Smith-Waterman database searches with inter- sequence SIMD parallelization, BMC Bioinformatics 12:221 (2011)
T. Rognes, Faster Smith-Waterman database searches with inter- sequence SIMD parallelization, BMC Bioinformatics 12:221 (2011). doi:10.1186/1471-2105-12-221
2011 doi
-
[19]
Pennycook, J
S. Pennycook, J. Sewall, V . Lee, Implications of a metric for performance portability, Future Generation Computer Systems 92 (2019) 947–958. doi:10.1016/j.future.2017.08.007
2019 doi
-
[20]
Marowka, Reformulation of the performance portability met- ric, Software: Practice and Experience 52 (1) (2022) 154–171
A. Marowka, Reformulation of the performance portability met- ric, Software: Practice and Experience 52 (1) (2022) 154–171. doi:10.1002/spe.3002
2022 doi
-
[21]
Korpar, M
M. Korpar, M. Sikic, SW# - GPU-enabled exact alignments on genome scale., Bioinformatics 29 (19) (2013) 2494–2495. doi:10.1093/bioinformatics/btt410
2013 doi
-
[22]
Korpar, M
M. Korpar, M. Sosic, D. Blazeka, M. Sikic, SWdb: GPU-Accelerated Exact Sequence Similarity Database Search, PLOS ONE 10 (12) (2016) 1–11. doi:10.1371/journal.pone.0145857
2016 doi
-
[23]
Rucci, C
E. Rucci, C. Garc ´ıa, G. Botella, A. De Giusti, M. Naiouf, M. Prieto- Mat´ıas, State-of-the-Art in Smith-Waterman Protein Database Search on HPC Platforms, Springer International Publishing, Cham, 2016, pp. 197–
2016
-
[24]
MJ Rutter, Intel’s Variable Clock Speeds and Benchmarking, https://www.mjr19.org.uk/ IT/clocks.html (2023)
2023
-
[25]
URL {https://www.intel.la/content/www/xl/es/products/ \\platforms/details/alder-lake-s.html}
Intel Corporation, Alder Lake S (2023). URL {https://www.intel.la/content/www/xl/es/products/ \\platforms/details/alder-lake-s.html}
2023
-
[26]
Haseeb, N
M. Haseeb, N. Ding, J. Deslippe, M. Awan, Evaluating perfor- mance and portability of a core bioinformatics kernel on multi- ple vendor gpus, in: 2021 International Workshop on Performance, Portability and Productivity in HPC (P3HPC), 2021, pp. 68–78. doi:10.1109/P3HPC54578.2021.00010
2021
-
[27]
Vasileska, P
I. Vasileska, P. Tom ˇsiˇc, L. Kos, L. Bogdanovi ´c, Unveiling perfor- mance insights and portability achievements between cuda and sycl for particle-in-cell codes on di fferent gpu architectures, in: 2024 47th MIPRO ICT and Electronics Convention (MIPRO), 2024, pp. 1115–1120....
2024
-
[28]
Z. Jin, J. S. Vetter, Performance portability study of epistasis detection using sycl on nvidia gpu, in: Proceedings of the 13th ACM International Conference on Bioinformatics, Computational Biology and Health Infor- matics, BCB ’22, Association for Computing Machinery, New Yo...
2022
-
[29]
Z. Jin, J. S. Vetter, Understanding performance portability of bioinfor- matics applications in sycl on an nvidia gpu, in: 2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), 2022, pp. 2190–
2022
-
[30]
Solis-Vasquez, E
L. Solis-Vasquez, E. Mascarenhas, A. Koch, Experiences migrat- ing cuda to sycl: A molecular docking case study, in: Proceed- ings of the 2023 International Workshop on OpenCL, IWOCL ’23, Association for Computing Machinery, New York, NY , USA, 2023. doi:10.1145/3585341.3585372
2023
-
[31]
Casta ˜no, Y
G. Casta ˜no, Y . Faqir-Rhazoui, C. Garc´ıa, M. Prieto-Mat´ıas, Evaluation of intel’s dpc++ compatibility tool in heterogeneous computing, Journal of Parallel and Distributed Computing 165 (2022) 120–129
2022
-
[32]
Homerding, J
B. Homerding, J. Tramm, Evaluating the performance of the hipsycl toolchain for hpc kernels on nvidia v100 gpus, in: Proceed- ings of the International Workshop on OpenCL, IWOCL ’20, As- sociation for Computing Machinery, New York, NY , USA, 2020. doi:10.1145/3388333.3388660
2020
-
[33]
Faqir-Rhazoui, C
Y . Faqir-Rhazoui, C. Garc ´ıa, Exploring the performance and portability of the k-means algorithm on sycl across cpu and gpu architectures, J. Supercomput. 79 (16) (2023) 18480–18506. doi:10.1007 /s11227-023- 05373-2
2023
-
[34]
Z. Jin, J. S. Vetter, Understanding performance portability of sycl ker- nels: A case study with the all-pairs distance calculation in bioinformat- ics on gpus, in: 2023 IEEE International Parallel and Distributed Pro- cessing Symposium Workshops (IPDPSW), IEEE, 2023, pp. 366–...
2023
- [35]
-
[36]
L. A. Torres, C. J. Barrios, Y . Denneulin, C. Cient ´ıfico, G. de Investi- gaci´on, C. A. y a Gran, Evaluation of computational and energy perfor- mance in matrix multiplication algorithms on cpu and gpu using mkl, cublas and sycl, 2024. URL {https://api.semanticscholar.org/C...
2024
-
[37]
Faqir-Rhazoui, C
Y . Faqir-Rhazoui, C. Garcia, Sycl in the edge: performance and energy evaluation for heterogeneous acceleration, The Journal of Supercomput- ing 80 (2024) 1–21. doi:10.1007/s11227-024-05957-6
2024 doi
-
[38]
Schmidt, F
B. Schmidt, F. Kallenborn, A. Chacon, C. Hundt, Cudasw ++ 4.0: ultra- fast gpu-based smith–waterman protein sequence database search, BMC bioinformatics 25 (1) (2024) 342
2024
-
[39]
I. Z. Reguly, Evaluating the performance portability of sycl across cpus and gpus on bandwidth-bound applications, in: Proceedings of the SC’23 Workshops of The International Conference on High Perfor- mance Computing, Network, Storage, and Analysis, 2023, pp. 1038–
2023
-
[41]
Franquinet, Performance portability analysis of sycl with a classical cg on cpu, gpu, and fpga, Master’s thesis (2023)
J. Franquinet, Performance portability analysis of sycl with a classical cg on cpu, gpu, and fpga, Master’s thesis (2023)
2023
- [42]
-
[43]
Nguyen, P
P. Nguyen, P. Nayak, H. Anzt, Porting batched iterative solvers onto intel gpus with sycl, SC-W ’23, Association for Computing Machinery, New York, NY , USA, 2023, p. 1048–1058. doi:10.1145/3624062.3624181
2023
-
[44]
Carratal ´a-S´aez, Y
R. Carratal ´a-S´aez, Y . Torres, A. Gonzalez-Escribano, D. R. Llanos, et al., Open sycl on heterogeneous gpu systems: A case of study, arXiv preprint arXiv:2310.06947 (2023). doi:10.13140/RG.2.2.32379.73769
2023 arXiv
-
[45]
Z. Jin, J. S. Vetter, Understanding sycl portability for pseudorandom num- ber generation: a case study with gene-expression connectivity mapping, in: 2023 IEEE International Parallel and Distributed Processing Sympo- sium Workshops (IPDPSW), IEEE, 2023, pp. 295–298
2023
-
[46]
Weckert, L
C. Weckert, L. Solis-Vasquez, J. Oppermann, A. Koch, O. Sinnen, Altis- sycl: Migrating altis benchmarking suite from cuda to sycl for gpus and fpgas, in: Proceedings of the SC’23 Workshops of The International Con- ference on High Performance Computing, Network, Storage, and A...
2023
-
[47]
R. Mueller-Albrecht, Syclomatic: Sycl adoption for everyone - moving from cuda to sycl gets progressively easier: Advanced migration consid- 13 erations, in: Proceedings of the 12th International Workshop on OpenCL and SYCL, IWOCL ’24, Association for Computing Machinery, New ...
2024
-
[223]
URL https://doi.org/10.1007/978-3-319-41279-5\_6
doi:10.1007/978-3-319-41279-5 6. URL https://doi.org/10.1007/978-3-319-41279-5\_6
-
[2195]
doi:10.1109/BIBM55620.2022.9995222
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.