Pith. sign in

REVIEW 4 major objections 5 minor 57 references

Optimizing the Variant Calling Pipeline Execution on Human Genomes Using GPU-Enabled Machines

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read This paper claims that scheduling variant-calling pipeline stages with machine-learned runtimes and a flexible job-shop planner cuts workload makespan by 2x over greedy assignment.

desk verdict A solid applied result — ML-predicted stage times plus FJSP scheduling gives real ~2x makespan gains on GPU cloud variant calling — but the evaluation lacks artifacts, variance, and a working greedy baseline, so treat the speedups as provisional. read the letter →

arxiv 2509.09058 v1 pith:DVJG6HUY submitted 2025-09-10 cs.DC

classification cs.DC
keywords variantcallingGPUcloudcomputingflexiblejobshopschedulingmachinelearningmakespanworkflowgenomesequences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that executing a variant calling pipeline over a batch of human genomes on heterogeneous GPU cloud VMs can be made dramatically faster by treating it as a flexible job shop scheduling problem, using ML-predicted stage times as the schedule inputs. It claims that stage times are predictable from genome characteristics beyond size, such as read quality, duplicate rate, average read length, and GC content, and that the resulting static plans outperform both a greedy ML-based assignment and a dynamic master-worker assignment. If correct, this offers a way to cut cloud cost for large-cohort genome processing without changing the pipeline's accuracy.

What carries the argument

The FJSP model, in which each genome is a job with ordered operations (pipeline stages), each operation can run on a chosen VM, one operation per VM runs at a time, and no operation is preempted, with the goal of minimizing makespan. Predicted stage times come from per-VM, per-stage regression models trained on sequence features like size, average read length, duplicate fraction, and quality scores. The plan executor then enforces the schedule with WAIT/SIGNAL file-lock statements, so that stage ordering holds even when actual times drift from predictions.

What would settle it

Run a 10-genome batch with the same five VMs and the same pipeline but with each VM writing to local scratch instead of shared network storage; if the makespan drops sharply or the FJSP plan's order is disrupted, shared-storage I/O is a first-order effect the model ignores. Separately, inflate one predicted stage by 30% for all genomes and check whether the plan re-optimizes or simply stalls.

Watch

Extended reading notes

Core claim

The central claim is that the makespan of a batch of genome sequences on heterogeneous GPU machines is minimized by decomposing each genome's variant calling pipeline into stages, predicting each stage's execution time on each machine type with regression models trained on sequence features, solving the flexible job shop scheduling problem (FJSP) on those predicted times, and executing the resulting plan with lightweight file-lock synchronization. The paper reports that this FJSP-based plan reduced average makespan from 10,411 seconds (greedy) and 8,428 seconds (dynamic) to 5,270 seconds on nine overlapping 10-genome test batches, a 2.0x and 1.6x average speedup respectively. Random forest r

Load-bearing premise

The paper treats the ML-predicted stage durations as deterministic during a run; if concurrent stages slow each other down through shared storage I/O or GPU contention more than the predictions capture, the FJSP plan loses its optimality and the speedups shrink.

Editorial extensions

If this is right

  • Batch variant-calling workflows on GPU clouds can finish in about half the wall-clock time of a greedy assignment, which translates to roughly halved cloud cost at pay-as-you-go prices.
  • Predictive features beyond sequence size carry real signal; models using them beat size-only models by an R2 margin of roughly 0.18 (0.894 vs 0.718 for the one-stage pipeline).
  • Splitting a pipeline into more, shorter stages yields better schedules; the two-stage FJSP plan beat the one-stage FJSP plan on every test subset.
  • Static optimized schedules can beat a dynamic master-worker assignment even when the runtime predictions carry 13–15% average error.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The FJSP scheduling strategy should transfer to other multi-stage bioinformatics pipelines (for example RNA-seq or ChIP-seq) that run on heterogeneous accelerators, since it only needs per-stage time predictions and a stage graph.
  • A natural extension is closed-loop scheduling: re-solve the FJSP periodically with updated predicted times from finished stages, which would soften the deterministic-time assumption.
  • The reported speedups are on low-coverage public genomes; clinical-grade 30x coverage sequences have longer stages and may shift the balance between planning granularity and prediction error.
  • Comparing against an online list scheduler that also uses runtime estimates would isolate the value of the global plan from the value of the predictions themselves.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper addresses the problem of minimizing the makespan of a multi-stage variant calling pipeline (FASTQ->BAM->VCF, using Parabricks) executed on a heterogeneous set of GPU-enabled VMs. The authors train ML models (RF, XGBoost, LR, etc.) on sequence features to predict per-stage execution times on each VM type, then use these predictions as constants in a flexible job shop scheduling (FJSP) formulation solved with OR-Tools CP-SAT. Planned schedules are executed with file-lock based WAIT/SIGNAL synchronization. Experiments on FABRIC with 5 VMs and 9 subsets of 10 held-out low-coverage genomes show that RF predictions improve when using all features (R2 up to 0.904 for FASTQ->BAM, 0.804 for BAM->VCF, 0.894 for 1-stage), and that the FJSP 2-stage strategy achieves an average 2.00x speedup over a greedy ML-based strategy and 1.61x over a dynamic master-worker strategy.

Significance. If the claims are robust, the paper makes a practical contribution: it is one of the first attempts to formulate whole-workload variant calling execution on heterogeneous GPU VMs as an FJSP, and it validates the approach by actually running the generated plans on a real testbed rather than by simulation. The demonstration that sequence characteristics beyond file size improve stage-time predictions is useful. The strong points are the real execution, the use of a held-out set for the scheduling experiments, the comparison against a dynamic scheduler, and the resource-utilization plots. However, the evaluation has important statistical and modeling limitations that currently prevent the headline speedup numbers from being considered reliable.

major comments (4)
  1. [Section 4.2, Table 6] The 9 subsets are generated from only 18 held-out sequences, each subset containing 10 sequences. Thus the same sequence appears in multiple subsets and the rows of Table 6 are not independent. No standard deviation, confidence interval, or significance test is reported. The average speedups of 2.00x and 1.61x could be driven by a few favorable subsets or by the particular overlapping split. Please report per-sequence makespans, use disjoint batches, or at least provide a bootstrap/paired analysis that accounts for the overlap. This is load-bearing for the central speedup claim.
  2. [Algorithm 2 (Greedy strategy)] The pseudocode as written cannot correctly schedule N>M jobs. In the outer loop over i, line 9 resets \hat M <- M at every iteration, so the machine removed at line 20 is available again in the next iteration. The same VM can be selected repeatedly, and the algorithm does not implement the described assignment of remaining jobs to remaining machines. Example 3.3 uses N=M=3 and therefore does not expose this flaw, but Table 6 uses N=10, M=5. As written, the greedy baseline is not reproducible and could be far worse than intended. Please correct the pseudocode (e.g., maintain a persistent set of available machines) and confirm that the experiments use the corrected version.
  3. [Sections 3.1, 3.4, and Algorithm 3] The FJSP model treats the predicted stage times T(o^k_ij) as constants that are independent of the schedule and of other concurrently executing stages. In the 2-stage FJSP plan, FASTQ->BAM and BAM->VCF for the same job can execute on different VMs, requiring the BAM file to be transferred over the shared NFS. No transfer-time or bandwidth-contention term appears in the model. Since training measurements were likely made under lower concurrency, the CP-SAT 'optimal' plan is only optimal with respect to an approximate model. Table 7 shows average RE of 13.2% for FJSP 2-stage, with subset 7 at 39.1%, so prediction errors are not negligible. The paper should quantify NFS transfer times and contention, or provide an argument that the omitted costs do not systematically favor the FJSP 2-stage strategy over greedy/dynamic.
  4. [Section 3.4 / OR-Tools formulation] The FJSP model is only described verbally and via Algorithm 1; the actual CP-SAT constraints (decision variables, routing constraints, no-preemption constraints, makespan objective) are not given. Without the model, the optimality claim cannot be checked or reproduced. Please specify the formulation explicitly or provide a link to the solver model/code.
minor comments (5)
  1. [Algorithm 2] The input line says 'J - Set of M VMs'; this should be M, not J.
  2. [Algorithm 4] The pseudocode does not mark a VM as busy after assigning a job to it. While the 'free' check in the while loop implicitly assumes workers become busy, an explicit update would make the master-worker logic unambiguous.
  3. [Tables 4 and 5] MSE is reported in units labeled 'in Seconds', but mean squared error is in seconds squared. Please correct the units.
  4. [Section 4.3] SVM and NN results are omitted with the comment that they 'performed worse'. Since ML model comparison is a contribution, report their R2 values at least in a supplementary table.
  5. [Abstract and body] The notation '2X speedup' should be typeset consistently as '2x' or '2x'; also 'on an average' is nonstandard and should be 'on average'.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the makespan speedups are measured outcomes of executing generated plans, not quantities forced by the fitted ML inputs.

full rationale

The paper's derivation chain is: (1) measure stage runtimes on the five GPU VMs; (2) train RF models on 80 public sequences using sequence features; (3) predict stage times for 18 held-out sequences; (4) feed those predictions as fixed constants into a CP-SAT FJSP solver; (5) execute the generated plans and measure the actual makespan (Table 6). The reported 2x/1.6x speedups are measured outcomes of actually running the plans, not values derived from the fitted models by construction. The ML predictions are validated on held-out sequences, and the predicted-vs-actual makespan errors (average RE 4.96%–14.69% in Table 7) show the schedules are not forced to match the predictions. The comparison against Greedy uses the same RF predictions, and Dynamic uses no predictions, so the speedup reflects scheduling quality rather than an identity. The only self-citations ([11], [38]) appear in the related-work survey as examples of prior GPU/commodity-cluster variant calling; they are not used to justify the FJSP formulation, the ML feature set, or the speedup claim. The unmodeled NFS/GPU contention flagged in Sections 3.1 and 3.3 is a correctness and generalizability risk, not circularity: it could weaken the speedup under contention, but it does not make the measured speedup equal to the model inputs. No circular step is present.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on empirical calibration of ML models to a specific testbed, on the correctness of the FJSP model, and on the assumption that measured stage times remain valid during concurrent scheduled execution. No new physical or theoretical entities are introduced; BEGIN, EXEC, WAIT, SIGNAL, and END are software plan statements.

free parameters (4)
  • ML model parameters (RF, XGBoost, LR, etc.) = not reported
    Execution time predictions for each VM type and pipeline stage are fitted to 80 measured sequences; these fitted parameters drive all schedules. Coefficients and hyperparameters are not disclosed.
  • Polling interval s in Dynamic strategy = 30 s
    Algorithm 4 sets s=30 s; the dynamic baseline's idle time depends on this value and it is not swept in the evaluation.
  • Pipeline stage split K = 1 or 2
    The 2-stage split (FASTQ to BAM, BAM to VCF) is chosen and yields the best speedup; the paper does not explore finer granularities or justify K=2 as optimal.
  • Feature set in Table 1 = 12 features including size
    The feature set is hand-selected. The claim that additional sequence characteristics improve prediction rests on a single comparison between size-only and all-features models, with no ablation.
assumptions (5)
  • domain assumption Execution times of variant calling stages are predictable functions of the Table 1 sequence features and VM type.
    Section 3.3 formulates y = Phi(f1..fn) and trains models on measured runs. The R2 values (0.80 to 0.90) support this partially, but it is not derived.
  • domain assumption FJSP instances for 10 jobs, 5 machines, and K<=2 operations can be solved to proven optimality by OR-Tools CP-SAT within practical time.
    Algorithm 1, Line 4 calls the FJSP solver without reporting time limits, optimality gaps, or solver configuration.
  • domain assumption Intermediate BAM files can be passed between VMs over shared NFS at a cost that does not change the ranking of schedules.
    Section 3.1 describes shared storage; 2-stage FJSP hands off BAMs across VMs, but no I/O time model appears in the features.
  • domain assumption Per-stage execution times are independent across VMs, with no interference from co-scheduled jobs or storage traffic.
    The FJSP model assumes deterministic operation times; contention is not measured or modeled.
  • domain assumption The 18 held-out genomes and the 9 overlapping 10-sequence subsets represent a production WGS workload.
    Section 4.3 generates subsets from 18 sequences; the overlap means the samples are not independent and no population-level claim is established.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimizing the Variant Calling Pipeline Execution on Human Genomes Using GPU-Enabled Machines." pith.science (2026). https://pith.science/paper/DVJG6HUY

@misc{pith2026250909058,
  author       = {Pith},
  title        = {Pith review of: Optimizing the Variant Calling Pipeline Execution on Human Genomes Using GPU-Enabled Machines},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DVJG6HUY}},
  note         = {Machine review of arXiv:2509.09058}
}
read the original abstract

Variant calling is the first step in analyzing a human genome and aims to detect variants in an individual's genome compared to a reference genome. Due to the computationally-intensive nature of variant calling, genomic data are increasingly processed in cloud environments as large amounts of compute and storage resources can be acquired with the pay-as-you-go pricing model. In this paper, we address the problem of efficiently executing a variant calling pipeline for a workload of human genomes on graphics processing unit (GPU)-enabled machines. We propose a novel machine learning (ML)-based approach for optimizing the workload execution to minimize the total execution time. Our approach encompasses two key techniques: The first technique employs ML to predict the execution times of different stages in a variant calling pipeline based on the characteristics of a genome sequence. Using the predicted times, the second technique generates optimal execution plans for the machines by drawing inspiration from the flexible job shop scheduling problem. The plans are executed via careful synchronization across different machines. We evaluated our approach on a workload of publicly available genome sequences using a testbed with different types of GPU hardware. We observed that our approach was effective in predicting the execution times of variant calling pipeline stages using ML on features such as sequence size, read quality, percentage of duplicate reads, and average read length. In addition, our approach achieved 2X speedup (on an average) over a greedy approach that also used ML for predicting the execution times on the tested workload of sequences. Finally, our approach achieved 1.6X speedup (on an average) over a dynamic approach that executed the workload based on availability of resources without using any ML-based time predictions.

Figures

Figures reproduced from arXiv: 2509.09058 by the authors.

Figure 1
Figure 1. (a) Our computing model (b) Overview of our ap [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Schedules/Execution Plans: FJSP-based strategy vs. [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. CPU utilization plots (for the 5 VMs) for different strategies on a representative subset [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: GPU utilization plots (for the 5 VMs) for different strategies on a representative subset [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 1 linked inside Pith

  1. [1]

    José M Abuín, Juan C Pichel, Tomás F Pena, and Jorge Amigo. 2015. BigBWA: Ap- proaching the Burrows-Wheeler Aligner to Big Data Technologies.Bioinformatics 31, 24 (2015), 4003–4005

  2. [2]

    José M Abuín, Juan C Pichel, Tomás F Pena, and Jorge Amigo. 2016. SparkBWA: Speeding up the Alignment of High-Throughput DNA Sequencing Data.PLoS ONE11, 5 (2016)

  3. [3]

    Ahmed, Joshua M

    Azza E. Ahmed, Joshua M. Allen, Tajesvi Bhat, Prakruthi Burra, Christina E. Fliege, Steven N. Hart, Jacob R. Heldenbrand, Matthew E. Hudson, Dave De- andre Istanto, Michael T. Kalmbach, Gregory D. Kapraun, Katherine I. Kendig, Matthew Charles Kendzior, Eric W. Klee, Nate Mattson, Christian A. Ross, Sami M. Sharif, Ramshankar Venkatakrishnan, Faisal M. Fad...

  4. [4]

    Nauman Ahmed, Vlad-Mihai Sima, Ernst Houtgast, Koen Bertels, and Zaid Al-Ars

  5. [5]

    1988.Mapping and Sequencing the Human Genome

    Bruce Alberts and et.al. 1988.Mapping and Sequencing the Human Genome. National Academies Press

  6. [6]

    Frederik Otzen Bagger, Line Borgwardt, Andreas Sand Jespersen, Anna Reimer Hansen, Birgitte Bertelsen, Miyako Kodama, and Finn Cilius Nielsen. 2024. Whole Genome Sequencing in Clinical Practice.BMC Medical Genomics17, 1 (2024), 39

  7. [7]

    Monga, Kuang- Ching Wang, Tom Lehman, and Paul Ruth

    Ilya Baldin, Anita Nikolich, James Griffioen, Indermohan Inder S. Monga, Kuang- Ching Wang, Tom Lehman, and Paul Ruth. 2019. FABRIC: A National-Scale Programmable Experimental Network Infrastructure.IEEE Internet Computing 23, 6 (2019), 38–47

  8. [8]

    Peter Brucker and Rainer Schlie. 1990. Job-Shop Scheduling With Multi-Purpose Machines.Computing45, 4 (1990), 369–375

Show all 57 references
  1. [9]

    Yu-Ting Chen, Jason Cong, Zhenman Fang, Jie Lei, and Peng Wei. 2016. When Apache Spark Meets FPGAs: A Case Study for Next-Generation DNA Sequenc- ing Acceleration. InProc. of the 8th USENIX Conference on Hot Topics in Cloud Computing(Denver, CO). 64–70

  2. [10]

    Spellman, Peng Wei, and Peipei Zhou

    Jason Cong, Jie Lei, Sen Li, Myron Peto, P. Spellman, Peng Wei, and Peipei Zhou

  3. [11]

    Manas Das, Khawar Shehzad, and Praveen Rao. 2023. Efficient Variant Calling on Human Genome Sequences Using a GPU-Enabled Commodity Cluster. InProc. of 32nd ACM International Conference on Information and Knowledge Management (CIKM). 3843–3848

  4. [12]

    InHigh Throughput Sequencing Algorithms and Applications (HITSEQ)

    CS-BWAMEM: A Fast and Scalable Read Aligner at the Cloud Scale for Whole Genome Sequencing. InHigh Throughput Sequencing Algorithms and Applications (HITSEQ)

  5. [13]

    Decap, J

    D. Decap, J. Reumers, C. Herzeel, P. Costanza, and J. Fostier. 2015. Halvade: Scalable Sequence Analysis with MapReduce.Bioinformatics31, 15 (2015), 2482– 2488

  6. [14]

    Stéphane Dauzère-Pérès, Junwen Ding, Liji Shen, and Karim Tamssaouet. 2024. The Flexible Job Shop Scheduling Problem: A Review.European Journal of Operational Research314, 2 (2024), 409–432

  7. [15]

    Donald Freed, Renke Pan, Haodong Chen, Zhipan Li, Jinnan Hu, and Rafael Aldana. 2022. DNAscope: High Accuracy Small Variant Calling Using Machine Learning.bioRxiv 2022.05.20.492556(2022)

  8. [16]

    Juan J Durillo and Radu Prodan. 2014. Multi-Objective Workflow Scheduling in Amazon EC2.Cluster computing17 (2014), 169–189

  9. [17]

    Po-Jung Huang, Jui-Huan Chang, Hou-Hsien Lin, Yu-Xuan Li, Chi-Ching Lee, Chung-Tsai Su, Yun-Lung Li, Ming-Tai Chang, Sid Weng, Wei-Hung Cheng, et al

  10. [18]

    Google. 2021. DeepVariant. https://github.com/google/deepvariant

  11. [19]

    Broad Institute. 2023. GATK4. https://github.com/broadinstitute/gatk

  12. [20]

    Kenneth Katz, Oleg Shutov, Richard Lapoint, Michael Kimelman, J Rodney Brister, and Christopher O’Sullivan. 2021. The Sequence Read Archive: A Decade More of Explosive Growth.Nucleic Acids Research50, D1 (11 2021), D387–D390

  13. [21]

    Broad Institute. 2020. HaplotypeCaller in a Nutshell. https://gatk.broadinstitute. org/hc/en-us/articles/360035531412-HaplotypeCaller-in-a-nutshell

  14. [22]

    Ben Langmead and Abhinav Nellore. 2018. Cloud Computing for Genomic Data Analysis and Collaboration.Nature Reviews Genetics19, 4 (2018), 208–219

  15. [23]

    Heng Li. 2013. Aligning Sequence Reads, Clone Sequences and Assembly Contigs With BWA-MEM.arXiv preprint arXiv:1303.3997(March 2013)

  16. [24]

    Daniel C. Koboldt. 2020. Best Practices for Variant Calling in Clinical Sequencing. Genome Medicine12, 1 (2020), 91

  17. [25]

    Michael Lo, Zhenman Fang, Jie Wang, Peipei Zhou, Mau-Chung Frank Chang, and Jason Cong. 2020. Algorithm-Hardware Co-design for BQSR Acceleration in Genome Analysis ToolKit. In2020 IEEE 28th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 157–166

  18. [26]

    Alberto Mulone, Sherine Awad, Davide Chiarugi, and Marco Aldinucci. 2023. Porting the Variant Calling Pipeline for NGS Data in Cloud-HPC Environment. In2023 IEEE 47th Annual Computers, Software, and Applications Conference. 1858– 1863

  19. [27]

    Wen-Wei Liao, Mobin Asri, Jana Ebler, Daniel Doerr, Marina Haukness, Glenn Hickey, Shuangjia Lu, Julian K Lucas, Jean Monlong, Haley J Abel, et al. 2023. A Draft Human Pangenome Reference.Nature617, 7960 (2023), 312–324

  20. [28]

    Nguyen, W

    T. Nguyen, W. Shi, and D Ruden. 2011. CloudAligner: A Fast and Full-Featured MapReduce Based Tool for Sequence Mapping.BMC Research Notes4, 1 (2011), 171

  21. [29]

    Niemenmaa, A

    M. Niemenmaa, A. Kallio, A. Schumacher, P. Klemela, E. Korpelainen, and K. Hel- janko. 2012. Hadoop-BAM: Directly Manipulating Next Generation Sequencing Data in the Cloud.Bioinformatics28, 6 (2012), 876–877

  22. [30]

    NCBI. 2013. Genome Reference Consortium Human Build 38. https://www.ncbi. nlm.nih.gov/datasets/genome

  23. [31]

    Lin- derman, Michael J

    Frank Austin Nothaft, Matt Massie, Timothy Danford, Zhao Zhang, Uri Laserson, Carl Yeksigian, Jey Kottalam, Arun Ahuja, Jeff Hammerbacher, Michael D. Lin- derman, Michael J. Franklin, Anthony D. Joseph, and David A. Patterson. 2015. Rethinking Data-Intensive Science Using Scal...

  24. [32]

    NVIDIA. 2020. NVIDIA Clara Parabricks. https://developer.nvidia.com/clara- parabricks

  25. [33]

    Frank A. Nothaft. 2017.Scalable Systems and Algorithms for Genomic Variant Analysis. Ph. D. Dissertation. UC Berkeley, ProQuest

  26. [34]

    Laurent Perron, Frédéric Didier, and Steven Gay. 2023. The CP-SAT-LP Solver. In 29th International Conference on Principles and Practice of Constraint Programming (CP 2023), Vol. 280. 3:1–3:2

  27. [35]

    Luca Pireddu, Simone Leo, and Gianluigi Zanetti. 2011. SEAL: A Distributed Short Read Mapping and Duplicate Removal Tool.Bioinformatics27, 15 (2011), 2159–2160

  28. [36]

    O’Connell, Zelaikha B

    Kyle A. O’Connell, Zelaikha B. Yosufzai, Ross A. Campbell, Collin J. Lobb, Haley T. Engelken, Laura M. Gorrell, Thad B. Carlson, Josh J. Catana, Dina Mikdadi, Vivien R. Bonazzi, and Juergen A. Klenk. 2023. Accelerating Genomic Workflows Using NVIDIA Parabricks.BMC Bioinformati...

  29. [37]

    Ryan Poplin, Pi-Chuan Chang, David Alexander, Scott Schwartz, Thomas Colthurst, Alexander Ku, Dan Newburger, Jojo Dijamco, Nam Nguyen, Pegah T Afshar, Sam S Gross, Lizzie Dorfman, Cory Y McLean, and Mark A DePristo. 2018. A universal SNP and Small-Indel Variant Caller Using De...

  30. [38]

    Praveen Rao, Arun Zachariah, Deepthi Rao, Peter Tonellato, Wesley Warren, and Eduardo Simoes. 2021. Accelerating Variant Calling on Human Genomes Using a Commodity Cluster. InProc. of 30th ACM International Conference on Information and Knowledge Management (CIKM). 3388–3392

  31. [39]

    Ryan Poplin, Pi-Chuan Chang, David Alexander, Scott Schwartz, Thomas Colthurst, Alexander Ku, Dan Newburger, Jojo Dijamco, Nam Nguyen, Pegah T Afshar, Sam Gross, Lizzie Dorfman, Cory McLean, and DePristo Mark. 2018. A Universal SNP and Small-Indel Variant Caller Using Deep Neu...

  32. [40]

    Michael C. Schatz. 2009. CloudBurst: Highly Sensitive Read Mapping with MapReduce.Bioinformatics25, 11 (2009), 1363–1369

  33. [41]

    Konrad Scheffler, Severine Catreux, Taylor O’Connell, Heejoon Jo, Varun Jain, Theo Heyns, Jeffrey Yuan, Lisa Murray, James Han, and Rami Mehio. 2023. So- matic Small-Variant Calling Methods in Illumina DRAGEN™Secondary Analysis. bioRxiv 2023.03.23.534011(2023)

  34. [42]

    Maria A Rodriguez and Rajkumar Buyya. 2017. Budget-Driven Scheduling of Scientific Workflows in IaaS Clouds With Fine-Grained Billing Periods.ACM Transactions on Autonomous and Adaptive Systems (TAAS)12, 2 (2017), 1–22

  35. [43]

    Helena S. I. L. Silva, Maria C. S. Castro, Fabricio A. B. Silva, and Alba C. M. A. Melo

  36. [44]

    Viktória Spišaková, Lukáš Hejtmánek, and Jakub Hynšt. 2023. Nextflow in Bioinformatics: Executors Performance Comparison Using Genomics Data.Future Generation Computer Systems142 (2023), 328–339

  37. [45]

    Shringarpure, Andrew Carroll, Francisco M

    Suyash S. Shringarpure, Andrew Carroll, Francisco M. De La Vega, and Carlos D. Bustamante. 2015. Inexpensive and Highly Reproducible Cloud-Based Variant Calling of 2,535 Human Genomes.PLOS ONE10, 6 (06 2015), 1–10

  38. [46]

    Ahmad Taghinezhad-Niar, Saeid Pashazadeh, and Javid Taheri. 2022. QoS-aware Online Scheduling of Multiple Workflows Under Task Execution Time Uncer- tainty in Clouds.Cluster Computing25, 6 (2022), 3767–3784

  39. [47]

    2009.Hadoop: The Definitive Guide(1st ed.)

    Tom White. 2009.Hadoop: The Definitive Guide(1st ed.). O’Reilly Media, Inc

  40. [48]

    Yuanqing Xia, Yufeng Zhan, Li Dai, and Yuehong Chen. 2023. A Cost and Makespan Aware Scheduling Algorithm for Dynamic Multi-Workflow in Cloud Environment.The Journal of Supercomputing79, 2 (2023), 1814–1833

  41. [49]

    Stephens, Skylar Y

    Zachary D. Stephens, Skylar Y. Lee, Faraz Faghri, Roy H. Campbell, Chengxiang Zhai, Miles J. Efron, Ravishankar Iyer, Michael C. Schatz, Saurabh Sinha, and Gene E. Robinson. 2015. Big Data: Astronomical or Genomical?PLOS Biology13, 7 (2015), 1–11

  42. [50]

    Chih-Han Yang, Jhih-Wun Zeng, Cheng-Yueh Liu, and Shih-Hao Hung. 2020. Accelerating Variant Calling with Parallelized DeepVariant. InProceedings of the International Conference on Research in Adaptive and Convergent Systems (Gwangju, Republic of Korea). 13–18

  43. [51]

    Taedong Yun, Helen Li, Pi-Chuan Chang, Michael F Lin, Andrew Carroll, and Cory Y McLean. 2021. Accurate, Scalable Cohort Variant Calls Using DeepVariant and GLnexus.Bioinformatics36, 24 (2021), 5582–5589

  44. [52]

    Franklin, Scott Shenker, and Ion Stoica

    Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2010. Spark: Cluster Computing with Working Sets. InProc. of the 2nd USENIX Conference on Hot Topics in Cloud Computing. Boston

  45. [53]

    Tiancheng Xu, Scott Rixner, and Alan L. Cox. 2023. An FPGA Accelerator for Genome Variant Calling.ACM Transactions on Reconfigurable Technology and Systems(May 2023), 1–20

  46. [57]

    Lingqi Zhang, Cheng Liu, and Shoubin Dong. 2019. PipeMEM: A Framework to Speed Up BWA-MEM in Spark with Low Overhead.Genes10, 11 (2019). 14th International ParBio Workshop ’25, October 12, 2025, Philadelphia, PA Ajay Kumar, Praveen Rao, and Peter Sanders (a) Greedy Strategy (b...

  47. [2015]

    InIn Proc

    Heterogeneous Hardware/Software Acceleration of the BWA-MEM DNA Alignment Algorithm. InIn Proc. of 2015 IEEE/ACM International Conference on Computer-Aided Design (ICCAD). 240–246

  48. [2020]

    DeepVariant-on-Spark: Small-Scale Genome Analysis Using a Cloud-Based Computing Framework.Computational and Mathematical Methods in Medicine 2020, 1 (2020), 7231205

  49. [2024]

    In30th European Conference on Parallel and Distributed Processing (Euro-Par)(Madrid, Spain)

    A Framework for Automated Parallel Execution of Scientific Multi-workflow Applications in the Cloud with Work Stealing. In30th European Conference on Parallel and Distributed Processing (Euro-Par)(Madrid, Spain). 298–311

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.