REVIEW 3 major objections 4 minor 164 references
An Empirical Investigation on the Challenges in Scientific Workflow Systems Development
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Scientific workflow systems' hardest topic is workflow execution, with 68.5% of Stack Overflow posts unanswered.
desk verdict Useful first topic map for SWS developer discussions, but GitHub bot traffic contaminates the difficulty rankings and needs fixing before the GitHub conclusions can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a BERTopic pipeline—a transformer-based topic-modeling method that embeds documents with a sentence-transformer model (multi-qa-MiniLM-L6-dot-v1), reduces dimensions with UMAP, clusters with HDBSCAN, and labels topics from count-vectorized keywords—combined with two difficulty metrics borrowed from prior Stack Overflow studies: the percentage of posts or issues without an accepted answer or resolution, and the median time to answer or resolve. The same pipeline is applied to GitHub issue and pull-request text. A complementary manual step classifies a statistically sampled set of 2,933 Stack Overflow posts into How, Why, What, and Other question types (Cohen's kappa 0.82), and fine-grained second-pass BERTopic runs decompose the largest topics into subtopics, which is how the paper operationalizes 'challenge' as something measurable and rankable.
What would settle it
Re-run the GitHub analysis after removing activity from automated bot accounts (e.g., 'precommitci autoupdate' and 'Auto Compress Images by Calibre' PRs) and check whether the near-zero unresolved rates and sub-hour resolution times for those topics persist, and whether System Redesign and API Migration still has the longest median resolution time.
Extended reading notes
Core claim
The paper claims that mining Stack Overflow and GitHub with BERTopic reveals the principal developer challenges in scientific workflow system development: ten topics on Stack Overflow (workflow creation and scheduling, distributed task management, workflow execution, data structures and operations, and others) and thirteen on GitHub (errors and bug fixing, documentation, dependency management, system redesign and API migration, and others). Using two difficulty metrics—the percentage of posts or issues without an accepted answer or resolution, and the median time to answer or resolve—the paper finds that workflow execution is the most challenging topic on Stack Overflow (68.50% unanswered, 23.58 hours median) and that system redesign and API migration has the longest median resolution time on GitHub (117.95 hours), while browser compatibility and HDFS integration has the highest unresolved rate (25.0%). The paper also reports that How-type questions dominate across all topics (60.97% on average), indicating a need for procedural guidance, and that several topics—data structures and operations, task management, and workflow scheduling—recur on both platforms.
Load-bearing premise
The GitHub difficulty rankings assume the issues and pull requests in the dataset are genuine human developer challenges; if bot-generated updates (like 'precommitci autoupdate' or 'Auto Compress Images by Calibre' pull requests) make up a large share of several topics, those rankings reflect automation, not human difficulty.
Editorial extensions
If this is right
- Developer-support investment in SWSs should target workflow execution tooling and debugging aids, since that topic has the highest unanswered rate and slowest answer time on Stack Overflow.
- SWS maintainers should expect and plan for long-running refactoring and API migration work, as this topic shows the highest median resolution time on GitHub.
- The dominance of How-type questions implies that step-by-step tutorials and worked examples would address a community-wide need more directly than reference-style documentation.
- The unusually low duplicate-question rate (0.37%) suggests many SWS questions are novel; building a structured, searchable knowledge base could materially reduce the 60% unanswered-post rate.
- Cross-platform topics such as data structures and operations, task management, and workflow scheduling are shared pain points, so improvements there benefit both Q&A users and issue-tracker communities.
Reading between the lines
- The near-zero unresolved rates and sub-hour resolution times reported for automated GitHub topics (e.g., automated tool updates, image compression) are likely an artifact of bot-generated pull requests; excluding bot activity would probably raise the measured difficulty of those topics.
- The same two-platform mining approach could be applied to other emerging engineering fields to locate support gaps before their knowledge bases mature; the low duplicate rate suggests SWS is still early in that maturation.
- A testable extension is to correlate the How-type dominance with documentation coverage: if official SWS docs already describe a topic, the share of How questions on that topic should be lower.
- The paper's difficulty metrics are proxy signals; connecting them to actual user frustration (e.g., through surveys or issue-closing comments) would test whether median hours genuinely capture perceived difficulty.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports an empirical study of developer challenges in Scientific Workflow Systems (SWSs) by mining 35,619 Stack Overflow posts and 163,118 GitHub issues and pull requests related to eleven SWSs. The authors apply BERTopic after manual filtering and preprocessing, identify 10 SO topics and 13 GitHub topics, classify SO questions into How/Why/What/Other types, and measure topic difficulty using the percentage of unanswered/unresolved items and median times to answer/resolution. The main reported findings are that Workflow Execution is the most challenging SO topic (68.50% unanswered, 23.58 hours), while System Redesign and API Migration has the highest median GitHub resolution time (117.95 hours); the paper also compares SWSs with other software engineering domains and draws implications for practitioners, educators, and researchers.
Significance. If the results withstand scrutiny, the paper provides a useful, community-grounded map of SWS development pain points and a large, reusable dataset plus replication package. The authors give credit for the careful manual filtering of ambiguous SO tags (e.g., the Galaxy case), the inter-annotator agreement (Cohen's kappa 0.82) on question-type classification, and the explicit reporting of BERTopic hyperparameters and coherence scores. The cross-platform comparison between Stack Overflow and GitHub is a methodological strength that goes beyond many single-platform topic-mining studies. However, the GitHub-side analysis currently contains a potentially load-bearing confound because bot-generated pull requests appear to be included in the topic model and difficulty metrics.
major comments (3)
- [§3.2, §4.1, Table 10, Table 17] The GitHub analysis does not filter bot-authored issues and pull requests, and this appears to distort the topic set and difficulty rankings. Topics such as 'Automated Performance and Tool Update Integration', 'Image Compression & Optimization', and 'Automating Code Quality Checks' in Table 10 correspond to recurring automated PRs (e.g., 'precommitci autoupdate' and 'Auto Compress Images by Calibre's image-actions'), which the paper itself mentions in §4.1 without treating them as a confound. Table 17 shows that these topics have near-zero unresolved rates and very short median resolution times (0.1 h, 0.97 h, and 16.63 h), which is exactly what bot-generated PRs would produce. Since no bot-account filtering is described in §3.2, the GitHub topic model and the Table 17 difficulty metrics are computed over a mixture of human and automated activity; the claim that 'System Redesign and API Migration' is the most challenging GitHub topic (117.95 h) is therefore not yet established. The authors should rerun the GitHub pipeline after filtering bot accounts, or at least quantify the proportion of bot traffic and show that the topic structure and difficulty rankings are robust to its removal.
- [Abstract and §4.3/Table 17] The abstract states that the GitHub analysis 'discovered that data structures and operations is the most difficult', but this is not supported by any GitHub result in the paper: Table 10 does not contain a GitHub topic named 'Data Structures and Operations', and Table 17 reports System Redesign and API Migration as having the highest median resolution time (117.95 hours) while Browser Compatibility and HDFS Integration Issues have the highest unresolved rate (25.0%). This inconsistency concerns the paper's central summary of its own findings and must be resolved by aligning the abstract with the reported results.
- [§4.3, Table 17] The use of 'open state' as the sole indicator of an unresolved GitHub issue or pull request, combined with the bot contamination, may also conflate maintenance workflow with difficulty. For example, topics with very low unresolved rates (Dependencies 0.02%, Managing Releases 0.38%, Automating Code Quality Checks 0.01%) could reflect PRs that are automatically closed or merged without representing genuine developer challenges. The paper should report how many items in each topic are PRs versus issues, how many are authored by known bot accounts, and how the difficulty metrics change when only human-authored, issue-type items are considered.
minor comments (4)
- [Table 10] The description for topic 9, 'Chord Execution and Task Coordination Issues', says 'Improving code quality through refactoring, and error fixes', which does not match the topic label or its keywords; this appears to be a copy-paste error and should be corrected.
- [§5.1] The evolution claims in Figures 6-8 are based on raw counts over time, which can be influenced by the overall growth of Stack Overflow and GitHub; the paper should either normalize by platform-wide activity or explicitly acknowledge this threat when interpreting the 'divergence' between SO and GitHub.
- [Table 18] There is a typo in the column header 'Agv Score'; it should read 'Avg Score'.
- [§3.1] The description of the celery filtering process is confusing: the text reports 9,499 tagged posts, then mentions 139 remaining posts and manual scrutiny, but the arithmetic leading from the initial 9,628 posts to these intermediate numbers is not clearly explained; adding a small summary table or explicit counts for each filtering step would improve reproducibility.
Circularity Check
No circularity: the study's topic labels and difficulty rankings are direct empirical summaries of the Stack Overflow and GitHub data, not quantities derived from the paper's own assumptions or prior self-citations.
full rationale
This paper is an observational mining study rather than a derivation. The central outputs—BERTopic topic clusters, manually assigned topic labels, percentages of unanswered/unresolved items, and median resolution times (Tables 2, 10, 15, 17)—are computed directly from the collected Stack Overflow and GitHub data using an external tool (BERTopic) and standard summary statistics. No fitted parameter is reused as a predicted outcome, no self-definitional quantity is present, and no uniqueness theorem or author-imported ansatz is invoked to force the conclusions. The only self-citations are background references, e.g., Alam et al. [1] on reusability barriers and the authors' replication package [53], and neither is load-bearing for the empirical findings. The concern that bot-generated GitHub pull requests may distort the GitHub topic set and difficulty rankings is a threat to construct validity, not a circularity of argument, because the paper never defines its target using the same fitted quantities it reports. The derivation chain, such as it is, is therefore self-contained with respect to the data. Score 0.
Assumptions & free parameters
free parameters (4)
- UMAP n_neighbors =
SO: 30, GitHub: 20
- UMAP n_components =
SO: 3, GitHub: 4
- HDBSCAN min_cluster_size =
SO: 210, GitHub: 130
- GitHub text length threshold =
8 characters
assumptions (4)
- domain assumption Stack Overflow posts and GitHub issues/PRs related to the selected SWSs are a representative and accurate reflection of SWS developer challenges.
- domain assumption Automated bot activity on GitHub does not materially distort the topic and difficulty analysis.
- domain assumption The percentage of unanswered posts and open issues, and median time to resolution, are valid proxies for topic difficulty.
- domain assumption The 11 SWSs with more than 100 SO posts are representative of the broader SWS ecosystem (352 systems collected).
Cite this review
Pith. "Pith review of An Empirical Investigation on the Challenges in Scientific Workflow Systems Development." pith.science (2026). https://pith.science/paper/3LQGQ2MT
@misc{pith2026241110890,
author = {Pith},
title = {Pith review of: An Empirical Investigation on the Challenges in Scientific Workflow Systems Development},
year = {2026},
howpublished = {\url{https://pith.science/paper/3LQGQ2MT}},
note = {Machine review of arXiv:2411.10890}
}
read the original abstract
Scientific Workflow Systems (SWSs) are advanced software frameworks that drive modern research by orchestrating complex computational tasks and managing extensive data pipelines. These systems offer a range of essential features, including modularity, abstraction, interoperability, workflow composition tools, resource management, error handling, and comprehensive documentation. Utilizing these frameworks accelerates the development of scientific computing, resulting in more efficient and reproducible research outcomes. However, developing a user-friendly, efficient, and adaptable SWS poses several challenges. This study explores these challenges through an in-depth analysis of interactions on Stack Overflow (SO) and GitHub, key platforms where developers and researchers discuss and resolve issues. In particular, we leverage topic modeling (BERTopic) to understand the topics SWSs developers discuss on these platforms. We identified 10 topics developers discuss on SO (e.g., Workflow Creation and Scheduling, Data Structures and Operations, Workflow Execution) and found that workflow execution is the most challenging. By analyzing GitHub issues, we identified 13 topics (e.g., Errors and Bug Fixing, Documentation, Dependencies) and discovered that data structures and operations is the most difficult. We also found common topics between SO and GitHub, such as data structures and operations, task management, and workflow scheduling. Additionally, we categorized each topic by type (How, Why, What, and Others). We observed that the How type consistently dominates across all topics, indicating a need for procedural guidance among developers. The dominance of the How type is also evident in domains like Chatbots and Mobile development. Our study will guide future research in proposing tools and techniques to help the community overcome the challenges developers face when developing SWSs.
Reference graph
Works this paper leans on
-
[1]
In: 2023 30th Asia-Pacific Software Engineering Conference (APSEC), pp
Alam, K., Roy, B., Serebrenik, A.: Reusability challenges of scientific workflows: A case study for galaxy. In: 2023 30th Asia-Pacific Software Engineering Conference (APSEC), pp. 289–298 (2023). https://doi.org/10.1109/APSEC60848.2023.00039
arXiv 2023
-
[2]
Journal of Grid Computing13, 457–493 (2015)
Liu, J., Pacitti, E., Valduriez, P., Mattoso, M.: A survey of data-intensive scientific workflow management. Journal of Grid Computing13, 457–493 (2015)
2015
-
[3]
In: HEALTHINF, pp
Almeida, J.R., Ribeiro, R.F., Oliveira, J.L.: A modular workflow management framework. In: HEALTHINF, pp. 414–421 (2018)
2018
-
[4]
Olabarriaga, S., Pierantoni, G., Taffoni, G., Sciacca, E., Jaghoori, M., Korkhov, V., Castelli, G., Vuerli, C., Becciani, U., Carley, E.,et al.: Scientific workflow management–for whom? In: 2014 IEEE 10th International Conference on e-Science, vol. 1, pp. 298–305 (2014). IEEE
2014
-
[5]
Genome research15(10), 1451–1455 (2005)
Giardine, B., Riemer, C., Hardison, R.C., Burhans, R., Elnitski, L., Shah, P., Zhang, Y., Blankenberg, D., Albert, I., Taylor, J.,et al.: Galaxy: a platform for interactive large-scale genome analysis. Genome research15(10), 1451–1455 (2005)
2005
-
[6]
AcM SIGKDD explorations Newsletter11(1), 26–31 (2009) 48
Berthold, M.R., Cebron, N., Dill, F., Gabriel, T.R., K¨ otter, T., Meinl, T., Ohl, P., Thiel, K., Wiswedel, B.: Knime-the konstanz information miner: version 2.0 and beyond. AcM SIGKDD explorations Newsletter11(1), 26–31 (2009) 48
2009
-
[7]
Bioinformat- ics28(19), 2520–2522 (2012)
K¨ oster, J., Rahmann, S.: Snakemake—a scalable bioinformatics workflow engine. Bioinformat- ics28(19), 2520–2522 (2012)
2012
-
[8]
Nature biotechnology35(4), 316–319 (2017)
Di Tommaso, P., Chatzou, M., Floden, E.W., Barja, P.P., Palumbo, E., Notredame, C.: Nextflow enables reproducible computational workflows. Nature biotechnology35(4), 316–319 (2017)
2017
Show all 164 references
-
[9]
FGCS25, 541–551 (2009)
McPhillips, T., Bowers, S., Zinn, D., Lud¨ ascher, B.: Scientific workflow design for mere mortals. FGCS25, 541–551 (2009)
2009
-
[10]
JoDS1, 19–30 (2012)
Bowers, S.: Scientific workflow, provenance, and data modeling challenges and approaches. JoDS1, 19–30 (2012)
2012
-
[11]
IEEE TSC 2, 79–92 (2009)
Lin, C., Lu, S., Fei, X., Chebotko, A., Pai, D., Lai, Z., Fotouhi, F., Hua, J.: A reference architecture for scientific workflow management systems and the view soa solution. IEEE TSC 2, 79–92 (2009)
2009
-
[12]
The International Journal of High Performance Computing Applications33(6), 1128–1139 (2019)
Deelman, E., Mandal, A., Jiang, M., Sakellariou, R.: The role of machine learning in scientific workflows. The International Journal of High Performance Computing Applications33(6), 1128–1139 (2019)
2019
-
[13]
IEEE Transactions on Services Computing11(3), 480–492 (2016)
Marozzo, F., Talia, D., Trunfio, P.: A workflow management system for scalable data mining on clouds. IEEE Transactions on Services Computing11(3), 480–492 (2016)
2016
-
[14]
Bioinformatics20(17), 3045–3054 (2004)
Oinn, T., Addis, M., Ferris, J., Marvin, D., Senger, M., Greenwood, M., Carver, T., Glover, K., Pocock, M.R., Wipat, A.,et al.: Taverna: a tool for the composition and enactment of bioinformatics workflows. Bioinformatics20(17), 3045–3054 (2004)
2004
-
[15]
IEEE Transactions on Services Computing9(2), 213–226 (2013)
Fern´ andez, H., Tedeschi, C., Priol, T.: A chemistry-inspired workflow management system for decentralizing workflow execution. IEEE Transactions on Services Computing9(2), 213–226 (2013)
2013
-
[16]
PloS one10(3), 0116781 (2015)
Li, Z., Yang, C., Jin, B., Yu, M., Liu, K., Sun, M., Zhan, M.: Enabling big geoscience data analytics with a cloud-based, mapreduce-enabled and service-oriented workflow framework. PloS one10(3), 0116781 (2015)
2015
-
[17]
In: 2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid), pp
Hendrix, V., Fox, J., Ghoshal, D., Ramakrishnan, L.: Tigres workflow library: Supporting scientific pipelines on hpc systems. In: 2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid), pp. 146–155 (2016). IEEE
2016
-
[18]
In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pp
Mamykina, L., Manoim, B., Mittal, M., Hripcsak, G., Hartmann, B.: Design lessons from the fastest q&a site in the west. In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, pp. 2857–2866 (2011)
2011
-
[19]
In: Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work, pp
Dabbish, L., Stuart, C., Tsay, J., Herbsleb, J.: Social coding in github: transparency and collaboration in an open software repository. In: Proceedings of the ACM 2012 Conference on Computer Supported Cooperative Work, pp. 1277–1286 (2012)
2012
-
[20]
In: Proceedings of the 10th Workshop on Workflows in Support of Large- Scale Science, pp
Mork, R., Martin, P., Zhao, Z.: Contemporary challenges for data-intensive scientific workflow management systems. In: Proceedings of the 10th Workshop on Workflows in Support of Large- Scale Science, pp. 1–11 (2015)
2015
-
[21]
ESE21(2016)
Rosen, C., Shihab, E.: What are mobile developers asking about? a large scale study using stack overflow. ESE21(2016)
2016
-
[22]
In: MSR, pp
Abdellatif, A., Costa, D., Badran, K., Abdalkareem, R., Shihab, E.: Challenges in chatbot development: A study of stack overflow posts. In: MSR, pp. 174–185 (2020)
2020
-
[23]
In: Proceedings of the 33rd International Conference on Software Engineering, 49 pp
Treude, C., Barzilay, O., Storey, M.-A.: How do programmers ask and answer questions on the web?(nier track). In: Proceedings of the 33rd International Conference on Software Engineering, 49 pp. 804–807 (2011)
2011
-
[24]
In: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp
Bagherzadeh, M., Khatchadourian, R.: Going big: a large-scale study on what big data develop- ers ask. In: Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp. 432–442 (2019)
2019
-
[25]
JCST31, 910–924 (2016)
Yang, X.-L., Lo, D., Xia, X., Wan, Z.-Y., Sun, J.-L.: What security questions do developers ask? a large-scale study of stack overflow posts. JCST31, 910–924 (2016)
2016
-
[26]
In: ICSME, pp
Li, H., Khomh, F., Openja, M.,et al.: Understanding quantum software engineering challenges an empirical study on stack exchange forums and github issues. In: ICSME, pp. 343–354. IEEE, ??? (2021). IEEE
2021
-
[27]
In: 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR), pp
Scoccia, G.L., Migliarini, P., Autili, M.: Challenges in developing desktop web apps: a study of stack overflow and github. In: 2021 IEEE/ACM 18th International Conference on Mining Software Repositories (MSR), pp. 271–282 (2021). IEEE
2021
-
[28]
Elsevier (2017)
Atkinson, M., Gesing, S., Montagnat, J., Taylor, I.: Scientific workflows: Past, present and future. Elsevier (2017)
2017
-
[29]
F1000Research10(2021)
M¨ older, F., Jablonski, K.P., Letcher, B., Hall, M.B., Tomkins-Tinch, C.H., Sochat, V., Forster, J., Lee, S., Twardziok, S.O., Kanitz, A., et al.: Sustainable data analysis with snakemake. F1000Research10(2021)
2021
-
[30]
GigaScience9(6), 068 (2020)
Kluge, M., Friedl, M.-S., Menzel, A.L., Friedel, C.C.: Watchdog 2.0: New developments for reusability, reproducibility, and workflow execution. GigaScience9(6), 068 (2020)
2020
-
[31]
Bioinformatics35(19), 3815–3817 (2019)
Cervera, A., Rantanen, V., Ovaska, K., Laakso, M., Nunez-Fontarnau, J., Alkodsi, A., Casado, J., Facciotto, C., H¨ akkinen, A., Louhimo, R.,et al.: Anduril 2: upgraded large-scale data integration framework. Bioinformatics35(19), 3815–3817 (2019)
2019
-
[32]
arXiv preprint arXiv:1909.08704 (2019)
Salim, M.A., Uram, T.D., Childers, J.T., Balaprakash, P., Vishwanath, V., Papka, M.E.: Bal- sam: Automated scheduling and execution of dynamic, data-intensive hpc workflows. arXiv preprint arXiv:1909.08704 (2019)
2019 arXiv
-
[33]
GigaScience8(5), 044 (2019)
Lampa, S., Dahl¨ o, M., Alvarsson, J., Spjuth, O.: Scipipe: A workflow library for agile development of complex and dynamic bioinformatics pipelines. GigaScience8(5), 044 (2019)
2019
-
[34]
Pal, S., Przytycka, T.M.: Bioinformatics pipeline using judi: just do it! Bioinformatics36(8), 2572–2574 (2020)
2020
-
[35]
Bioinformatics28(11), 1525–1526 (2012)
Sadedin, S.P., Pope, B., Oshlack, A.: Bpipe: a tool for running and managing bioinformatics pipelines. Bioinformatics28(11), 1525–1526 (2012)
2012
-
[36]
F1000Research5(2016)
Ewels, P., Krueger, F., K¨ aller, M., Andrews, S.: Cluster flow: A user-friendly bioinformatics workflow tool. F1000Research5(2016)
2016
-
[37]
Bioinformatics31(1), 10–16 (2015)
Cingolani, P., Sladek, R., Blanchette, M.: Bigdatascript: a scripting language for data pipelines. Bioinformatics31(1), 10–16 (2015)
2015
-
[38]
In: 2017 Ieee International Parallel and Distributed Processing Symposium Workshops (ipdpsw), pp
Jimenez, I., Sevilla, M., Watkins, N., Maltzahn, C., Lofstead, J., Mohror, K., Arpaci-Dusseau, A., Arpaci-Dusseau, R.: The popper convention: Making reproducible systems evaluation prac- tical. In: 2017 Ieee International Parallel and Distributed Processing Symposium Workshops...
2017
-
[39]
Working Draft 20085(11) (2009)
Ben-Kiki, O., Evans, C., Ingerson, B.: Yaml ain’t markup language (yaml™) version 1.1. Working Draft 20085(11) (2009)
2009
-
[40]
0 (2016) 50
Amstutz, P., Crusoe, M.R., Tijani´ c, N., Chapman, B., Chilton, J., Heuer, M., Kartashov, A., Leehr, D., M´ enager, H., Nedeljkovich, M., et al.: Common workflow language, v1. 0 (2016) 50
2016
-
[41]
F1000Research 2017
Voss, K., Gentry, J., Auwera, G.: Full-stack genomics pipelining with GATK4+ WDL+ Cromwell. F1000Research 2017
2017
-
[42]
Nature biotechnology35(4), 314–316 (2017)
Vivian, J., Rao, A.A., Nothaft, F.A., Ketchum, C., Armstrong, J., Novak, A., Pfeil, J., Nark- izian, J., Deran, A.D., Musselman-Brown, A.,et al.: Toil enables reproducible, open source, big biomedical data analyses. Nature biotechnology35(4), 314–316 (2017)
2017
-
[43]
Scientific Programming2015(1), 243180 (2015)
Santana-Perez, I., P´ erez-Hern´ andez, M.S.: Towards reproducibility in scientific workflows: An infrastructure-based approach. Scientific Programming2015(1), 243180 (2015)
2015
-
[44]
In: Proceedings of the 2016 International Conference on Management of Data, pp
Chirigati, F., Rampin, R., Shasha, D., Freire, J.: Reprozip: Computational reproducibility with ease. In: Proceedings of the 2016 International Conference on Management of Data, pp. 2085–2088 (2016)
2016
-
[45]
Future Generation Computer Systems75, 271–283 (2017)
Garijo, D., Gil, Y., Corcho, O.: Abstract, link, publish, exploit: An end to end framework for workflow sharing. Future Generation Computer Systems75, 271–283 (2017)
2017
-
[46]
Nucleic Acids Research50(W1), 345–351 (2022)
The galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2022 update. Nucleic Acids Research50(W1), 345–351 (2022)
2022
-
[47]
Bioinformatics27(20), 2907–2909 (2011)
Jagla, B., Wiswedel, B., Copp´ ee, J.-Y.: Extending knime for next-generation sequencing data analysis. Bioinformatics27(20), 2907–2909 (2011)
2011
-
[48]
Journal of Bioinformatics and Systems Biology6, 121–133 (2023)
Kiran, A.D., Ay, M.C., Allmer, J.: Criteria for the evaluation of workflow management systems for scientific data analysis. Journal of Bioinformatics and Systems Biology6, 121–133 (2023)
2023
-
[49]
Future Generation Computer Systems75, 256–270 (2017)
Sethi, R.J., Gil, Y.: Scientific workflows in data analysis: Bridging expertise across multiple domains. Future Generation Computer Systems75, 256–270 (2017)
2017
-
[50]
GigaScience8(7), 084 (2019)
Kotliar, M., Kartashov, A.V., Barski, A.: Cwl-airflow: a lightweight pipeline manager support- ing common workflow language. GigaScience8(7), 084 (2019)
2019
-
[51]
Concurrency and Computation: Practice and Experience27(17), 5037–5059 (2015)
Jain, A., Ong, S.P., Chen, W., Medasani, B., Qu, X., Kocher, M., Brafman, M., Petretto, G., Rignanese, G.-M., Hautier, G.,et al.: Fireworks: a dynamic workflow system designed for high-throughput applications. Concurrency and Computation: Practice and Experience27(17), 5037–50...
2015
-
[52]
IEEE Access9, 53491–53508 (2021)
Ahmad, Z., Jehangiri, A.I., Ala’anzy, M.A., Othman, M., Latip, R., Zaman, S.K.U., Umar, A.I.: Scientific workflows management and scheduling in cloud computing: taxonomy, prospects, and challenges. IEEE Access9, 53491–53508 (2021)
2021
-
[53]
Online; last accessed May, 2025 (2025)
Alam, K., Roy, B., Roy, C., Mittal, K.: Artifact of the paper ”An Empirical Investigation on the Challenges in Scientific Workflow Systems Development.”. Online; last accessed May, 2025 (2025). https://zenodo.org/records/15454588
2025
-
[54]
467–471 (2008)
Zhao, Y., Raicu, I., Foster, I.: Scientific workflow systems for 21st century, new bottle or new wine? In: 2008 IEEE Congress on Services-Part I, pp. 467–471 (2008). IEEE
2008
-
[55]
In: CyberC (2011)
Zhao, Y., Fei, X., Raicu, I., Lu, S.: Opportunities and challenges in running scientific workflows on the cloud. In: CyberC (2011). IEEE
2011
-
[56]
In: ACM SIGMOD, pp
Davidson, S.B., Freire, J.: Provenance and scientific workflows: challenges and opportunities. In: ACM SIGMOD, pp. 1345–1350 (2008)
2008
-
[57]
Computer40(12), 24–32 (2007)
Gil, Y., Deelman, E., Ellisman, M., Fahringer, T., Fox, G., Gannon, D., Goble, C., Livny, M., Moreau, L., Myers, J.: Examining the challenges of scientific workflows. Computer40(12), 24–32 (2007)
2007
-
[58]
In: International Conference on Service-Oriented Computing, pp
Gu, Y., Cao, J., Guo, Y., Qian, S., Guan, W.: Plan, generate and match: Scientific workflow recommendation with large language models. In: International Conference on Service-Oriented Computing, pp. 86–102 (2023). Springer 51
2023
-
[59]
GigaScience13, 030 (2024)
S¨ anger, M., De Mecquenem, N., Lewi´ nska, K.E., Bountris, V., Lehmann, F., Leser, U., Kosch, T.: A qualitative assessment of using chatgpt as large language model for scientific workflow development. GigaScience13, 030 (2024)
2024
-
[60]
Available at SSRN 4594836 (2023)
Procko, T., Davidoff, A., Elvira, T., Ochoa, O.: Towards improved scientific knowledge prolifer- ation: Leveraging large language models on the traditional scientific writing workflow. Available at SSRN 4594836 (2023)
2023
-
[61]
ACM Computing Surveys (CSUR)49(4), 1–39 (2016)
Liew, C.S., Atkinson, M.P., Galea, M., Ang, T.F., Martin, P., Hemert, J.I.V.: Scientific workflows: moving across paradigms. ACM Computing Surveys (CSUR)49(4), 1–39 (2016)
2016
-
[62]
In: PPAM, pp
Barker, A., Van Hemert, J.: Scientific workflow: a survey and research directions. In: PPAM, pp. 746–753 (2007). Springer
2007
-
[63]
Online; last accessed January, 2024 (2024)
WorkflowHub Community. Online; last accessed January, 2024 (2024). https://workflowhub. eu/workflows
2024
-
[64]
Online; last accessed January, 2024 (2024)
myExperiment Workflows. Online; last accessed January, 2024 (2024). https://www. myexperiment.org/workflows
2024
-
[65]
Online; last accessed January, 2024 (2024)
SnakeMake Workflows. Online; last accessed January, 2024 (2024). https://nf-co.re/
2024
-
[66]
Multimedia Tools and Applications 78, 15169–15211 (2019)
Jelodar, H., Wang, Y., Yuan, C., Feng, X., Jiang, X., Li, Y., Zhao, L.: Latent dirichlet allocation (lda) and topic modeling: models, applications, a survey. Multimedia Tools and Applications 78, 15169–15211 (2019)
2019
-
[67]
Scientometrics, 1–26 (2023)
Wang, Z., Chen, J., Chen, J., Chen, H.: Identifying interdisciplinary topics and their evolution based on bertopic. Scientometrics, 1–26 (2023)
2023
-
[68]
Scientometrics100, 767–786 (2014)
Yau, C.-K., Porter, A., Newman, N., Suominen, A.: Clustering scientific documents with topic modeling. Scientometrics100, 767–786 (2014)
2014
-
[69]
In: Advances in Information Retrieval: 31th European Conference on IR Research, ECIR 2009, Toulouse, France, April 6-9, 2009
Yi, X., Allan, J.: A comparative study of utilizing topic models for information retrieval. In: Advances in Information Retrieval: 31th European Conference on IR Research, ECIR 2009, Toulouse, France, April 6-9, 2009. Proceedings 31, pp. 29–41 (2009). Springer
2009
-
[70]
In: Proceedings of the 19th Nordic Conference of Computational Linguistics (NODALIDA 2013), pp
Luostarinen, T., Kohonen, O.: Using topic models in content-based news recommender systems. In: Proceedings of the 19th Nordic Conference of Computational Linguistics (NODALIDA 2013), pp. 239–251 (2013)
2013
-
[71]
In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp
Eidelman, V., Boyd-Graber, J., Resnik, P.: Topic models for dynamic translation model adap- tation. In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pp. 115–119 (2012)
2012
-
[72]
Soft Computing27(7), 3965–3982 (2023)
Belwal, R.C., Rai, S., Gupta, A.: Extractive text summarization using clustering-based topic modeling. Soft Computing27(7), 3965–3982 (2023)
2023
-
[73]
In: Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering-Volume 1, pp
Asuncion, H.U., Asuncion, A.U., Taylor, R.N.: Software traceability with topic modeling. In: Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering-Volume 1, pp. 95–104 (2010)
2010
-
[74]
Journal of Information Science40(5), 621– 636 (2014)
Bagheri, A., Saraee, M., De Jong, F.: Adm-lda: An aspect detection model based on topic modelling using the structure of review sentences. Journal of Information Science40(5), 621– 636 (2014)
2014
-
[75]
In: Pro- ceedings of the Fourth ACM International Conference on Web Search and Data Mining, pp
Jo, Y., Oh, A.H.: Aspect and sentiment unification model for online review analysis. In: Pro- ceedings of the Fourth ACM International Conference on Web Search and Data Mining, pp. 815–824 (2011)
2011
-
[76]
In: Advances in Knowledge Discovery and Data Mining: 15th Pacific-Asia Conference, 52 PAKDD 2011, Shenzhen, China, May 24-27, 2011, Proceedings, Part I 15, pp
Zhai, Z., Liu, B., Xu, H., Jia, P.: Constrained lda for grouping product features in opinion mining. In: Advances in Knowledge Discovery and Data Mining: 15th Pacific-Asia Conference, 52 PAKDD 2011, Shenzhen, China, May 24-27, 2011, Proceedings, Part I 15, pp. 448–459 (2011). Springer
2011
-
[77]
In: 2012 9th IEEE Working Conference on Mining Software Repositories (MSR), pp
Chen, T.-H., Thomas, S.W., Nagappan, M., Hassan, A.E.: Explaining software defects using topic models. In: 2012 9th IEEE Working Conference on Mining Software Repositories (MSR), pp. 189–198 (2012). IEEE
2012
-
[78]
Empirical Software Engineering21, 1843–1919 (2016)
Chen, T.-H., Thomas, S.W., Hassan, A.E.: A survey on the use of topic models when mining software repositories. Empirical Software Engineering21, 1843–1919 (2016)
2016
-
[79]
In: Proceedings of the 33rd International Conference on Software Engineering, pp
Thomas, S.W.: Mining software repositories using topic models. In: Proceedings of the 33rd International Conference on Software Engineering, pp. 1138–1139 (2011)
2011
-
[80]
In: Proceedings of the 8th Working Conference on Mining Software Repositories, pp
Thomas, S.W., Adams, B., Hassan, A.E., Blostein, D.: Modeling the evolution of topics in source code histories. In: Proceedings of the 8th Working Conference on Mining Software Repositories, pp. 173–182 (2011)
2011
-
[81]
In: 2009 6th IEEE International Working Conference on Mining Software Repositories, pp
Tian, K., Revelle, M., Poshyvanyk, D.: Using latent dirichlet allocation for automatic catego- rization of software. In: 2009 6th IEEE International Working Conference on Mining Software Repositories, pp. 163–166 (2009). IEEE
2009
-
[82]
In: 2010 IEEE International Conference on Software Maintenance, pp
Gethers, M., Poshyvanyk, D.: Using relational topic models to capture coupling among classes in object-oriented software systems. In: 2010 IEEE International Conference on Software Maintenance, pp. 1–10 (2010). IEEE
2010
-
[83]
In: Proceedings of the 22nd IEEE/ACM International Conference on Automated Software Engineering, pp
Linstead, E., Rigor, P., Bajracharya, S., Lopes, C., Baldi, P.: Mining concepts from code with probabilistic topic models. In: Proceedings of the 22nd IEEE/ACM International Conference on Automated Software Engineering, pp. 461–464 (2007)
2007
-
[84]
In: 2008 Seventh International Conference on Machine Learning and Applications, pp
Linstead, E., Lopes, C., Baldi, P.: An application of latent dirichlet allocation to analyz- ing software evolution. In: 2008 Seventh International Conference on Machine Learning and Applications, pp. 813–818 (2008). IEEE
2008
-
[85]
Information and Software Technology52(9), 972–990 (2010)
Lukins, S.K., Kraft, N.A., Etzkorn, L.H.: Bug localization using latent dirichlet allocation. Information and Software Technology52(9), 972–990 (2010)
2010
-
[86]
In: 2010 IEEE International Conference on Software Maintenance, pp
Savage, T., Dit, B., Gethers, M., Poshyvanyk, D.: Topic xp: Exploring topics in source code using latent dirichlet allocation. In: 2010 IEEE International Conference on Software Maintenance, pp. 1–6 (2010). IEEE
2010
-
[87]
In: 2024 12th International Symposium on Digital Forensics and Security (ISDFS), pp
Gokcimen, T., Das, B.: Topic modelling using bertopic for robust spam detection. In: 2024 12th International Symposium on Digital Forensics and Security (ISDFS), pp. 1–5 (2024). https://doi.org/10.1109/ISDFS60797.2024.10527342
2024
-
[88]
ACM Transactions on Information Systems (TOIS)34(2), 1–32 (2016)
Cheng, Z., Shen, J.: On effective location-aware music recommendation. ACM Transactions on Information Systems (TOIS)34(2), 1–32 (2016)
2016
-
[89]
Information Systems42, 59–77 (2014)
Kim, Y., Shim, K.: Twilite: A recommendation system for twitter using a probabilistic model based on latent dirichlet allocation. Information Systems42, 59–77 (2014)
2014
-
[90]
IEEE Intelligent Systems30(3), 18–25 (2015)
Lu, H.-M., Lee, C.-H.: A twitter hashtag recommendation model that accommodates for temporal clustering effects. IEEE Intelligent Systems30(3), 18–25 (2015)
2015
-
[91]
Future Generation Computer Systems65, 196–206 (2016)
Zhao, F., Zhu, Y., Jin, H., Yang, L.T.: A personalized hashtag recommendation approach using lda-based topic model in microblog environment. Future Generation Computer Systems65, 196–206 (2016)
2016
-
[92]
Information Sciences367, 573–599 (2016)
Zoghbi, S., Vuli´ c, I., Moens, M.-F.: Latent dirichlet allocation for linking user-generated content and e-commerce data. Information Sciences367, 573–599 (2016)
2016
-
[93]
arXiv 53 preprint arXiv:2203.05794 (2022)
Grootendorst, M.: Bertopic: Neural topic modeling with a class-based tf-idf procedure. arXiv 53 preprint arXiv:2203.05794 (2022)
2022 arXiv
-
[94]
Journal of machine Learning research3(Jan), 993–1022 (2003)
Blei, D.M., Ng, A.Y., Jordan, M.I.: Latent dirichlet allocation. Journal of machine Learning research3(Jan), 993–1022 (2003)
2003
-
[95]
Advances in neural information processing systems13(2000)
Lee, D., Seung, H.S.: Algorithms for non-negative matrix factorization. Advances in neural information processing systems13(2000)
2000
-
[96]
Discourse processes25(2-3), 259–284 (1998)
Landauer, T.K., Foltz, P.W., Laham, D.: An introduction to latent semantic analysis. Discourse processes25(2-3), 259–284 (1998)
1998
-
[97]
Advances in neural information processing systems18, 147 (2006)
Blei, D., Lafferty, J.: Correlated topic models. Advances in neural information processing systems18, 147 (2006)
2006
-
[98]
In: Proceedings of the 23rd International Conference on Machine Learning, pp
Blei, D.M., Lafferty, J.D.: Dynamic topic models. In: Proceedings of the 23rd International Conference on Machine Learning, pp. 113–120 (2006)
2006
-
[99]
In: Proceedings of the 22nd International Conference on World Wide Web, pp
Yan, X., Guo, J., Lan, Y., Cheng, X.: A biterm topic model for short texts. In: Proceedings of the 22nd International Conference on World Wide Web, pp. 1445–1456 (2013)
2013
-
[100]
In: 2023 International Conference on Computer, Control, Informatics and Its Applications (IC3INA), pp
Parlina, A., Maryati, I.: Leveraging bertopic for the analysis of scientific papers on seaweed. In: 2023 International Conference on Computer, Control, Informatics and Its Applications (IC3INA), pp. 279–283 (2023). https://doi.org/10.1109/IC3INA60834.2023.10285737
2023
-
[101]
In: 2023 Congress in Computer Science, Computer Engineering, and Applied Computing (CSCE), pp
Kang, W., Kim, Y., Kim, H., Lee, J.: An analysis of research trends on language model using bertopic. In: 2023 Congress in Computer Science, Computer Engineering, and Applied Computing (CSCE), pp. 168–172 (2023). https://doi.org/10.1109/CSCE60160.2023.00032
2023
-
[102]
In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pp
Doi, T., Isonuma, M., Yanaka, H.: Topic modeling for short texts with large language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pp. 21–33 (2024)
2024
-
[103]
In: 2022 29th International Conference on Systems, Signals and Image Processing (IWSSIP), vol
Mankolli, E.: Reducing the complexity of candidate selection using natural language processing. In: 2022 29th International Conference on Systems, Signals and Image Processing (IWSSIP), vol. CFP2255E-ART, pp. 1–4 (2022). https://doi.org/10.1109/IWSSIP55020.2022.9854488
2022
-
[104]
Sensors22(13), 4925 (2022)
Atzeni, D., Bacciu, D., Mazzei, D., Prencipe, G.: A systematic review of wi-fi and machine learning integration with topic modeling techniques. Sensors22(13), 4925 (2022)
2022
-
[105]
IEEE transactions on knowledge and data engineering29(6), 1186–1198 (2017)
Nie, L., Wei, X., Zhang, D., Wang, X., Gao, Z., Yang, Y.: Data-driven answer selection in community qa systems. IEEE transactions on knowledge and data engineering29(6), 1186–1198 (2017)
2017
-
[106]
In: Topic Detection and Tracking: Event-based Information Organization, pp
Fiscus, J.G., Doddington, G.R.: Topic detection and tracking evaluation overview. In: Topic Detection and Tracking: Event-based Information Organization, pp. 17–31. Springer, ??? (2002)
2002
-
[107]
In: Proceedings of the IEEE/ACM 42nd International Conference on Software Engineering Workshops, pp
Barbosa, L.S.: Software engineering for’quantum advantage’. In: Proceedings of the IEEE/ACM 42nd International Conference on Software Engineering Workshops, pp. 427–429 (2020)
2020
-
[108]
In: Proceedings of the 28th Annual ACM Symposium on Applied Computing, pp
Wang, S., Lo, D., Jiang, L.: An empirical study on developer interactions in stackoverflow. In: Proceedings of the 28th Annual ACM Symposium on Applied Computing, pp. 1019–1024 (2013)
2013
-
[109]
Journal of Systems and Software156, 283–299 (2019)
Chen, H., Coogle, J., Damevski, K.: Modeling stack overflow tags and topics as a hierarchy of concepts. Journal of Systems and Software156, 283–299 (2019)
2019
-
[110]
In: Mining Software Repositories (MSR), pp
Allamanis, M., Sutton, C.: Why, when, and what: analyzing stack overflow questions by topic, type, and code. In: Mining Software Repositories (MSR), pp. 53–56 (2013). IEEE
2013
-
[111]
In: International Symposium on (ESEM), pp
Alshangiti, M., Sapkota, H., Murukannaiah, P.K., Liu, X., Yu, Q.: Why is developing machine 54 learning applications challenging? a study on stack overflow posts. In: International Symposium on (ESEM), pp. 1–11 (2019). IEEE
2019
-
[112]
In: Proceedings of the 12th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp
Ahmed, S., Bagherzadeh, M.: What do concurrency developers ask about? a large-scale study using stack overflow. In: Proceedings of the 12th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, pp. 1–10 (2018)
2018
-
[113]
In: 2023 49th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), pp
Mi, Q., Bao, Q., Cui, L.: Identifying topics and trends in devops: A study of stack overflow posts. In: 2023 49th Euromicro Conference on Software Engineering and Advanced Applications (SEAA), pp. 402–409 (2023). https://doi.org/10.1109/SEAA60479.2023.00067
2023
-
[114]
In: 2021 IEEE/ACIS 19th International Conference on Software Engineering Research, Management and Applications (SERA), pp
Suwonchoochit, N., Senivongse, T.: Classification of database technology problems on stack overflow. In: 2021 IEEE/ACIS 19th International Conference on Software Engineering Research, Management and Applications (SERA), pp. 21–26 (2021). https://doi.org/10.1109/ SERA51205.2021.9509047
2021 arXiv
-
[115]
Information and Software Technology84, 19–32 (2017)
Zou, J., Xu, L., Yang, M., Zhang, X., Yang, D.: Towards comprehending the non-functional requirements through developers’ eyes: An exploration of stack overflow using topic analysis. Information and Software Technology84, 19–32 (2017)
2017
-
[116]
In: Proceedings of the 13th Innovations in Software Engineering Conference on Formerly Known as India Software Engineering Conference, pp
Dhasade, A.B., Venigalla, A.S.M., Chimalakonda, S.: Towards prioritizing github issues. In: Proceedings of the 13th Innovations in Software Engineering Conference on Formerly Known as India Software Engineering Conference, pp. 1–5 (2020)
2020
-
[117]
” (2021)
Jokhio, M.: Mining github issues for bugs, feature requests and questions. ” (2021)
2021
-
[118]
In: The Art and Science of Analyzing Software Data, pp
Campbell, J.C., Hindle, A., Stroulia, E.: Latent dirichlet allocation: extracting topics from software engineering data. In: The Art and Science of Analyzing Software Data, pp. 139–159. Elsevier, ??? (2015)
2015
-
[119]
94–97 (2019)
Wang, X., Lee, M., Pinchbeck, A., Fard, F.: Where does lda sit for github? In: 2019 34th IEEE/ACM International Conference on Automated Software Engineering Workshop (ASEW), pp. 94–97 (2019). IEEE
2019
-
[120]
In: 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR), pp
Treude, C., Wagner, M.: Predicting good configurations for github and stack overflow topic models. In: 2019 IEEE/ACM 16th International Conference on Mining Software Repositories (MSR), pp. 84–95 (2019). IEEE
2019
-
[121]
935–946 (2016)
Nadi, S., Kr¨ uger, S., Mezini, M., Bodden, E.: Jumping through hoops: Why do java develop- ers struggle with cryptography apis? In: Proceedings of the 38th International Conference on Software Engineering, pp. 935–946 (2016)
2016
-
[122]
Empirical Software Engineering25, 2694–2747 (2020)
Han, J., Shihab, E., Wan, Z., Deng, S., Xia, X.: What do programmers discuss about deep learning frameworks. Empirical Software Engineering25, 2694–2747 (2020)
2020
-
[123]
Empirical software engineering19, 619–654 (2014)
Barua, A., Thomas, S.W., Hassan, A.E.: What are developers talking about? an analysis of topics and trends in stack overflow. Empirical software engineering19, 619–654 (2014)
2014
-
[124]
Biology direct10, 1–12 (2015)
Spjuth, O., Bongcam-Rudloff, E., Hern´ andez, G.C., Forer, L., Giovacchini, M., Guimera, R.V., Kallio, A., Korpelainen, E., Ka´ ndu la, M.M., Krachunov, M.,et al.: Experiences with workflows for automating data-intensive bioinformatics. Biology direct10, 1–12 (2015)
2015
-
[125]
In: SSDBM, vol
Jaeger, E., Altintas, I., Zhang, J., Lud¨ ascher, B., Pennington, D., Michener, W.: A scientific workflow approach to distributed geospatial data processing using web services. In: SSDBM, vol. 3, pp. 87–90 (2005)
2005
-
[126]
In: 29th International Conference on Software Engineering (ICSE’07), pp
Carver, J.C., Kendall, R.P., Squires, S.E., Post, D.E.: Software development environments for scientific and engineering software: A series of case studies. In: 29th International Conference on Software Engineering (ICSE’07), pp. 550–559 (2007). Ieee 55
2007
-
[127]
arXiv preprint arXiv:2110.13999 (2021)
Nouri, A., Davis, P.E., Subedi, P., Parashar, M.: Exploring the role of machine learning in scientific workflows: Opportunities and challenges. arXiv preprint arXiv:2110.13999 (2021)
2021 arXiv
-
[128]
ACM Sigmod Record34(3), 44–49 (2005)
Yu, J., Buyya, R.: A taxonomy of scientific workflow systems for grid computing. ACM Sigmod Record34(3), 44–49 (2005)
2005
-
[129]
In: SciPy, pp
Rocklin, M.,et al.: Dask: Parallel computation with blocked algorithms and task scheduling. In: SciPy, pp. 126–132 (2015)
2015
-
[130]
Online; last accessed January, 2024 (2024)
Wikipedia: Scientific Workflow System. Online; last accessed January, 2024 (2024). https://en. wikipedia.org/wiki/Scientific workflow system
2024
-
[131]
Online; last accessed January, 2024 (2024)
pditommaso: Workflow Systems. Online; last accessed January, 2024 (2024). https://github. com/pditommaso/awesome-pipeline
2024
-
[132]
Online; last accessed January, 2024 (2024)
cwl: Existing Workflow Systems. Online; last accessed January, 2024 (2024). https://github.com/common-workflow-language/common-workflow-language/wiki/ Existing-Workflow-systems
2024
-
[133]
Online; last accessed January, 2024 (2024)
community: Workflow Systems. Online; last accessed January, 2024 (2024). https://workflows. community/systems
2024
-
[134]
https://data.stackexchange.com/
Exchange, S.: Stack Exchange Data Explorer. https://data.stackexchange.com/. accessed: January 31, 2024 (2024)
2024
-
[135]
In: Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions, pp
Bird, S.: Nltk: the natural language toolkit. In: Proceedings of the COLING/ACL 2006 Interactive Presentation Sessions, pp. 69–72 (2006)
2006
-
[136]
No Starch Press, ??? (2020)
Vasiliev, Y.: Natural Language Processing with Python and spaCy: A Practical Introduction. No Starch Press, ??? (2020)
2020
-
[137]
Online;January 2024 (2024)
github: GitHub REST API. Online;January 2024 (2024). https://docs.github.com/en/rest? apiVersion=2022-11-28
2024
-
[138]
Online; last accessed August, 2024 (2024)
Face, H.: Hugging Face Model Hub. Online; last accessed August, 2024 (2024). https:// huggingface.co/docs/hub/en/models-the-hub
2024
-
[139]
arXiv preprint arXiv:1802.03426 (2018)
McInnes, L., Healy, J., Melville, J.: Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426 (2018)
2018 arXiv
-
[140]
McInnes, L., Healy, J., Astels, S.,et al.: hdbscan: Hierarchical density based clustering. J. Open Source Softw.2(11), 205 (2017)
2017
-
[141]
In: 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), pp
Openja, M., Adams, B., Khomh, F.: Analysis of modern release engineering topics:–a large-scale study using stackoverflow–. In: 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME), pp. 104–114 (2020). IEEE
2020
-
[142]
In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp
Bajaj, K., Pattabiraman, K., Mesbah, A.: Mining questions asked by web developers. In: Proceedings of the 11th Working Conference on Mining Software Repositories, pp. 112–121 (2014)
2014
-
[143]
” O’Reilly Media, Inc.” (2017)
Hochstein, L., Moser, R.: Ansible: Up and running: Automating configuration management and deployment the easy way. ” O’Reilly Media, Inc.” (2017)
2017
-
[144]
Statistics in medicine19(5), 723–741 (2000)
Blackman, N.J.-M., Koval, J.J.: Interval estimation for cohen’s kappa as a measure of agreement. Statistics in medicine19(5), 723–741 (2000)
2000
-
[145]
Biochemia medica22(3), 276–282 (2012)
McHugh, M.L.: Interrater reliability: the kappa statistic. Biochemia medica22(3), 276–282 (2012)
2012
-
[146]
Spearman Rank Correlation Coefficient, pp. 502–505. Springer, New York, NY (2008). https: 56 //doi.org/10.1007/978-0-387-32833-1 379 . https://doi.org/10.1007/978-0-387-32833-1 379
2008 doi
-
[147]
Computing in Science and Engineering16(5), 62–74 (2014) https://doi.org/10.1109/MCSE.2014.80
Towns, J., Cockerill, T., Dahan, M., Foster, I., Gaither, K., Grimshaw, A., Hazlewood, V., Lathrop, S., Lifka, D., Peterson, G.D., Roskies, R., Scott, J.R., Wilkins-Diehr, N.: Xsede: Accelerating scientific discovery. Computing in Science and Engineering16(5), 62–74 (2014) htt...
2014 doi
-
[148]
IEEE Access7, 25138–25149 (2019)
Nadeem, F., Alghazzawi, D., Mashat, A., Faqeeh, K., Almalaise, A.: Using machine learning ensemble methods to predict execution time of e-science workflows in heterogeneous distributed systems. IEEE Access7, 25138–25149 (2019)
2019
-
[149]
Nature methods18(10), 1161–1168 (2021)
Wratten, L., Wilm, A., G¨ oke, J.: Reproducible, scalable, and shareable analysis pipelines with bioinformatics workflow managers. Nature methods18(10), 1161–1168 (2021)
2021
-
[150]
In: Proceedings of the 30th International Conference on Software Engineering, pp
Souza, C.R., Redmiles, D.F.: An empirical study of software developers’ management of dependencies and changes. In: Proceedings of the 30th International Conference on Software Engineering, pp. 241–250 (2008)
2008
-
[151]
IEEE Transactions on Software Engineering35(6), 864–878 (2009)
Cataldo, M., Mockus, A., Roberts, J.A., Herbsleb, J.D.: Software dependencies, work dependen- cies, and their impact on failures. IEEE Transactions on Software Engineering35(6), 864–878 (2009)
2009
-
[152]
IBM Systems Journal41(1), 4–12 (2002)
Hailpern, B., Santhanam, P.: Software debugging, testing, and verification. IBM Systems Journal41(1), 4–12 (2002)
2002
-
[153]
In: Proceedings of the the 6th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineering, pp
Duboc, L., Rosenblum, D., Wicks, T.: A framework for characterization and analysis of software system scalability. In: Proceedings of the the 6th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Software Engineer...
2007
-
[154]
Future Generation Computer Systems46, 17–35 (2015)
Deelman, E., Vahi, K., Juve, G., Rynge, M., Callaghan, S., Maechling, P.J., Mayani, R., Chen, W., Da Silva, R.F., Livny, M.,et al.: Pegasus, a workflow management system for science automation. Future Generation Computer Systems46, 17–35 (2015)
2015
-
[155]
XRDS: Crossroads, The ACM Magazine for Students16(3), 14–18 (2010)
Juve, G., Deelman, E.: Scientific workflows and clouds. XRDS: Crossroads, The ACM Magazine for Students16(3), 14–18 (2010)
2010
-
[156]
Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences371(1983), 20120066 (2013)
Berriman, G.B., Deelman, E., Juve, G., Rynge, M., V¨ ockler, J.-S.: The application of cloud computing to scientific workflows: a study of cost and performance. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences371(1983), 20120066 (2013)
1983
-
[157]
In: 2008 Grid Computing Environments Workshop, pp
Foster, I., Zhao, Y., Raicu, I., Lu, S.: Cloud computing and grid computing 360-degree compared. In: 2008 Grid Computing Environments Workshop, pp. 1–10 (2008). Ieee
2008
-
[158]
ACM SIGOPS Operating Systems Review49(1), 71–79 (2015)
Boettiger, C.: An introduction to docker for reproducible research. ACM SIGOPS Operating Systems Review49(1), 71–79 (2015)
2015
-
[159]
In: 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud
Zaharia, M., Chowdhury, M., Franklin, M.J., Shenker, S., Stoica, I.: Spark: Cluster computing with working sets. In: 2nd USENIX Workshop on Hot Topics in Cloud Computing (HotCloud
-
[160]
Database2012, 010 (2012)
Rak, R., Rowley, A., Black, W., Ananiadou, S.: Argo: an integrative, interactive, text mining- based workbench supporting curation. Database2012, 010 (2012)
2012
-
[161]
Research advances in cloud computing, 1–20 (2017)
Baldini, I., Castro, P., Chang, K., Cheng, P., Fink, S., Ishakian, V., Mitchell, N., Muthusamy, V., Rabbah, R., Slominski, A., et al.: Serverless computing: Current trends and open problems. Research advances in cloud computing, 1–20 (2017)
2017
-
[162]
ACM Computing Surveys (CSUR)52(4), 1–36 (2019)
Adhikari, M., Amgoth, T., Srirama, S.N.: A survey on scheduling strategies for workflows in 57 cloud environment and emerging trends. ACM Computing Surveys (CSUR)52(4), 1–36 (2019)
2019
-
[163]
Organizational Research Methods10(2), 393 (2007)
Bean, C.J.: Qualitative research design: An interactive approach. Organizational Research Methods10(2), 393 (2007)
2007
-
[164]
Springer Science & Business Media (2012) 58
Wohlin, C., Runeson, P., H¨ ost, M., Ohlsson, M.C., Regnell, B., Wessl´ en, A.: Experimentation in software engineering. Springer Science & Business Media (2012) 58
2012
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.