REVIEW 3 major objections 6 minor 65 references
Biomedical Knowledge Composition: A Software Engineering Perspective
T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Biomedical knowledge graphs are hard to build because data pipelines lack the software engineering tooling that made web development composable and reproducible.
desk verdict A genuinely useful synthesis and research agenda for biomedical KG engineering, with a central causal claim that remains a clearly-labeled hypothesis rather than an established finding. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the engineering maturity gap: the comparison between the standardized tooling of web development—versioned packages, typed interfaces, dependency resolution, continuous integration—and the bespoke, largely manual practice of biomedical data integration. The paper uses this gap as a diagnostic lens, framing each of its eight open challenges as a missing analogue of a web-engineering capability, from a data package manager and namespace type safety to a canonical intermediate representation and production-readiness engineering. The gap also does prescriptive work: it makes 'pipeline over artifact' the first-class design criterion, so that a versioned, documented build pipeline counts as more valuable than a static deployed graph. That criterion organizes the profiles of the six systems and the design principles the authors recommend for new projects.
What would settle it
A systematic comparison of biomedical knowledge-graph projects that use versioned, package-managed pipelines versus those that consume static snapshots—controlling for data scale and domain—and finds no significant difference in integration cost, error rate, or update speed would falsify the claim that tooling adoption is a root cause.
Extended reading notes
Core claim
The paper's central claim is that a contributing root cause of the difficulty in biomedical knowledge infrastructure is the limited adoption of software engineering tooling and practices that make web engineering reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. The paper argues that this engineering maturity gap is what turns the assembly of a biomedical knowledge graph into brittle, largely unrepeatable manual work. The corollary the authors draw is that the field should shift its emphasis from deployed knowledge graphs to the reproducible process of assembling them: versioned dependencies, reusable build pipelines, and engineering practices that let others compile and customize a graph from source rather than consume a static artifact. The supporting evidence is a taxonomy of five harmonization challenges, profiles of six representative systems arranged along a reproducibility spectrum, a concrete reuse attempt that exposed redeployment, discoverability, and coverage failures, and a catalogue of eight open engineering challenges, each with partial solutions but no universally adopted stack.
Load-bearing premise
The load-bearing premise is that the obstacles the authors met while reusing one large harmonized knowledge graph are structural features of the biomedical data problem rather than quirks of that particular system, so if that experience is idiosyncratic the general diagnosis loses much of its force.
Editorial extensions
If this is right
- If the diagnosis is correct, a versioned 'package manager' for biomedical data releases would eliminate a major source of silent breakage in graph assembly, much as dependency managers did for software.
- Teams that publish reproducible build pipelines instead of static exports would make schema changes, source updates, and provenance queries tractable engineering tasks, shifting maintenance burden away from downstream consumers.
- Namespace-aware type checking would turn mismatched identifier joins—which currently corrupt graphs silently—into compile-time errors that are caught before a graph is shipped.
- A standardized, validated graph interchange format with a provenance subgraph would allow continuous-integration systems to verify knowledge-graph assembly compliance automatically.
- The eight open challenges collectively define a research agenda in which software engineers, not only biomedical curators, can make direct contributions to knowledge infrastructure.
Reading between the lines
- If the maturity-gap thesis is right, a testable prediction follows: projects that adopt versioned, package-managed pipelines should show measurably lower integration cost and fewer silent errors than projects that consume static snapshots, when scale and domain are controlled for.
- The argument implies that the long tail of specialized databases will not be cured by more ontologies alone; the leverage lies in standardizing release mechanics—version numbers, machine-readable changelogs, integrity hashes—across data providers.
- The paper's web-engineering comparison is qualitative, so a quantitative gap analysis measuring what fraction of knowledge-graph construction steps are covered by versioned tooling across a sample of projects could upgrade the diagnosis into a measurement.
- The proposed stack may also be a prerequisite for LLM-based agents that assemble or update graphs: reliable agentic data integration needs exactly the typed, versioned, auditable interfaces the paper calls for.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a position/synthesis article that combines a tutorial on biomedical knowledge graphs (KGs) with a software-engineering critique of how they are assembled. The first half introduces the domain: why KGs are the central integrative data structure, five data-harmonization challenges, application areas such as drug discovery and digital twins, and six representative KG systems. The second half argues that a contributing root cause of the field's difficulty is limited adoption of software practices that make web engineering reliably composable and reproducible—specifically package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. The argument is supported by a qualitative comparison with web engineering, a first-person account of three obstacles encountered when reusing the Data Distillery KG, and references to related friction in the literature. The paper then formulates eight open engineering challenges and a call to shift from shipping static KG artifacts to shipping reproducible build pipelines. Section 11 explicitly acknowledges that the web comparison is qualitative and that the Data Distillery account reflects one team's experience with one version of one system.
Significance. If its central thesis is accepted, the paper provides a useful bridge between software engineering and biomedical knowledge infrastructure and could redirect investment toward data package management, namespace type enforcement, canonical interchange representations, service composition, and lifecycle governance. Its concrete strengths are the operationalized list of eight challenges with partial solutions, the explicit treatment of pipeline reproducibility as a first-class design criterion, the concrete Data Distillery obstacles, and an unusually candid limitations section. The paper does not provide a controlled empirical demonstration; the causal claim is a well-formed hypothesis rather than an established result. The reliance on the authors' own prior work [18] and their own Data Distillery experience [47] limits evidential independence but does not make the argument circular. As a position paper and research agenda, the contribution is valuable; as a demonstrated root-cause analysis, it is not yet supported.
major comments (3)
- [Section 6, paragraph beginning 'These obstacles are not unique...'] The inference from the three Data Distillery obstacles to structural properties is load-bearing for the paper's central claim. The three obstacles are an HPC root-privilege policy for Docker, a schema with a single node label and roughly 1600 relationship types, and incomplete harmonization coverage. These are respectively a deployment-policy constraint, a schema-design choice, and a curation-scope decision; none of them directly demonstrates the absence of package management, typed namespaces, canonical interchange formats, or lifecycle governance. The cited references [48-50] document related problems, but they provide no comparative measurement showing that projects adopting the proposed practices perform better. To make the root-cause claim defensible, the paper should either reframe the abstract and Section 7 as presenting a hypothesis with an explicit evaluation design (for example, a structured survey scoring KG projects on the six proposed practices and correlating those scores with reproducibility or integration-error outcomes), or add such evidence.
- [Section 7 and Section 8.1] The maturity-gap argument overstates the green-field character of the problem. General-purpose data versioning and packaging tools such as DataLad, DVC, Quilt, and git-annex already provide semantic versioning, integrity hashes, dependency graphs, and registry-like distribution. The real gap is the absence of a domain-specific convention for biomedical data releases, including ontology-aware schemas, identifier namespaces, and release semantics, not the total absence of package-manager concepts. The open problem in Section 8.1 should engage with these existing tools and explain why they do not transfer directly to biomedical KG assembly. Without this engagement, the claim that the biomedical landscape 'lacks' these capabilities is too strong and weakens the otherwise plausible adoption-based diagnosis.
- [Abstract, Section 7, and Section 11] The abstract and Section 7 use 'contributing root cause' and 'the engineering infrastructure gap is real' as definitive statements, while Section 11 concedes that the comparison is qualitative and that the Data Distillery case reflects one team's experience. Because the paper's contribution is a research agenda rather than a completed empirical study, the causal language should be explicitly hedged in the abstract and Section 7 (for example, 'a likely contributor' or 'a hypothesis supported by our experience and related reports'), with Section 11's caveats reflected consistently throughout. This change would make the paper more accurate without diminishing its value as a call to action.
minor comments (6)
- [Abstract] The phrase 'toward thereproducible process' is missing a space and should read 'toward the reproducible process.'
- [Section 2 and Figure 1] Figure 1 presents four challenge categories, but Section 3 introduces five core harmonization challenges; the relationship between the two partitions should be stated explicitly so that readers do not see an inconsistency.
- [Listings 1 and 2] The JSON keys in Listing 1 and the process name in Listing 2 appear with inserted spaces (for example, 'q u e r y _ g r a p h' and 'NO RM AL IS E'); if these are not rendering artifacts, the listings should be cleaned up.
- [Table 2, Reactome row] The 'Pipeline reproducibility' cell for Reactome describes the export format and the graph's usefulness for mechanistic modeling, but does not state whether the build/export pipeline can be re-run; this should be aligned with the other rows.
- [Section 10] The visualization section is interesting but only loosely connected to the eight SE challenges; a sentence linking semantic visualization to pipeline reproducibility and lifecycle governance would strengthen the integration.
- [Section 10, 'apinatomy'] The text mentions 'apinatomy panels' but does not provide a citation for the apinatomy framework; a reference should be added.
Circularity Check
No significant circularity: the central claim is an interpretive synthesis supported by case examples, not a derivation that reduces to its own inputs.
full rationale
The paper makes no quantitative predictions and fits no parameters, so none of the fitted-input or self-definitional circularity patterns apply. Its central thesis—that limited adoption of software-engineering practices such as package management, typed namespaces, canonical interchange formats, service composition, reproducible pipelines, and lifecycle governance is a contributing cause of biomedical knowledge-graph construction difficulty—is presented as an argued interpretation supported by domain examples, profiles of six KG systems, and a single case study. Section 11 explicitly concedes that the comparison to web engineering is a qualitative observation rather than a systematic empirical study and that the Data Distillery account reflects one team's experience with one version of one system. The only author self-citation is reference [18], which is used to note that semantic-mapping methods assist entity resolution; this is not load-bearing for the article's central argument. External references [48–50] independently document similar frictions, and the paper does not invoke any author-specific uniqueness theorem or disguise a fitted parameter as a prediction. The article's recommendations are a research agenda and an interpretive synthesis, not conclusions forced by construction from their premises.
Assumptions & free parameters
assumptions (3)
- domain assumption Knowledge graphs are the central integrative data structure for biomedicine.
- domain assumption Web engineering tooling is a valid template for biomedical data integration.
- ad hoc to paper The friction experienced with Data Distillery is representative of structural problems in the field.
Cite this review
Pith. "Pith review of Biomedical Knowledge Composition: A Software Engineering Perspective." pith.science (2026). https://pith.science/paper/WEHLWUEC
@misc{pith2026260808927,
author = {Pith},
title = {Pith review of: Biomedical Knowledge Composition: A Software Engineering Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/WEHLWUEC}},
note = {Machine review of arXiv:2608.08927}
}
read the original abstract
Biomedical research has accumulated vast molecular, clinical, and population data, yet translating this wealth into actionable knowledge remains constrained by technical and organizational difficulties. This article presents a unified treatment of two perspectives on biomedical knowledge infrastructure. The first introduces the biomedical domain to software engineers: it explains why knowledge graphs (KGs) are the central integrative data structure in modern biomedicine, characterizes five data harmonization challenges (identifier mapping, entity resolution, schema alignment, evidence integration, and provenance tracking), surveys application domains from drug discovery to digital twins, and profiles six representative KG systems with contrasting choices. The second perspective asks why engineering biomedical knowledge infrastructure remains so difficult. We argue that a contributing root cause is limited adoption of software tooling and practices that make development in other mature domains - particularly web engineering - reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. Against this backdrop, eight open engineering challenges for biomedical data integration are catalogued, each with partial solutions but no universally adopted stack. Crucially, the article shifts emphasis from describing deployed KG instances toward the reproducible process of assembling them: reusable build pipelines, versioned dependencies, and engineering practices that let others compile and customize a KG from source rather than consuming a static artifact. Together, the two perspectives provide domain grounding for newcomers and a research agenda for software engineers seeking to make transformative contributions to biomedical knowledge infrastructure.
Figures
Reference graph
Works this paper leans on
- [18]
-
[47]
Taylor Research Lab, CFDE Data Distillery Project,https://github.com/TaylorResearchLab/CFDE_ DataDistillery, accessed: 2026-05-21 (2024)
work page 2024
-
[1]
K. Tomczak, P. Czerwi ´nska, M. Wiznerowicz, The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge, Contemporary Oncology 19 (1A) (2015) A68–A77.doi:10.5114/wo.2014.47136
arXiv 2015
-
[2]
A. Regev, S. A. Teichmann, E. S. Lander, I. Amit, C. Benoist, E. Birney, et al., The Human Cell Atlas, eLife 6 (2017) e27041.doi:10.7554/eLife.27041. 19
-
[3]
Gene Ontology Consortium, The Gene Ontology Resource: 20 years and still GOing strong, Nucleic Acids Re- search 47 (D1) (2019) D330–D338.doi:10.1093/nar/gky1055
-
[4]
UniProt Consortium, UniProt: the Universal Protein Knowledgebase in 2023, Nucleic Acids Research 51 (D1) (2023) D523–D531.doi:10.1093/nar/gkac1052
-
[5]
M. Kanehisa, M. Furumichi, Y . Sato, M. Kawashima, M. Ishiguro-Watanabe, KEGG for taxonomy-based analysis of pathways and genomes, Nucleic Acids Research 51 (D1) (2023) D587–D592.doi:10.1093/nar/gkac963
- [6]
Show all 65 references
-
[7]
D. S. Wishart, Y . D. Feunang, A. C. Guo, E. J. Lo, A. Marcu, J. R. Grant, T. Sajed, D. Johnson, C. Li, Z. Sayeeda, et al., DrugBank 5.0: a major update to the DrugBank database for 2018, Nucleic Acids Research 46 (D1) (2018) D1074–D1082.doi:10.1093/nar/gkx1037
2018 doi
-
[8]
Mendez, A
D. Mendez, A. Gaulton, A. P. Bento, J. Chambers, M. De Veij, E. F ´elix, M. P. Magari˜nos, J. F. Mosquera, P. Mu- towo, M. Nowotka, et al., ChEMBL: towards direct deposition of bioassay data, Nucleic Acids Research 47 (D1) (2019) D930–D940.doi:10.1093/nar/gky1075
2019 doi
-
[9]
M. J. Landrum, J. M. Lee, M. Benson, G. R. Brown, C. Chao, S. Chitipiralla, B. Gu, J. Hart, D. Hoffman, W. Jang, et al., ClinVar: improving access to variant interpretations and supporting evidence, Nucleic Acids Research 46 (D1) (2018) D1062–D1067.doi:10.1093/nar/gkx1153
2018 doi
-
[10]
J. G. Tate, S. Bamford, H. C. Jubb, Z. Sondka, D. M. Beare, N. Bindal, H. Boutselakis, C. G. Cole, C. Creatore, E. Dawson, et al., COSMIC: the catalogue of somatic mutations in cancer, Nucleic Acids Research 47 (D1) (2019) D941–D947.doi:10.1093/nar/gky1015
2019 doi
-
[11]
K ¨ohler, M
S. K ¨ohler, M. Gargano, N. Matentzoglu, L. C. Carmody, D. Lewis-Smith, N. A. Vasilevsky, D. Danis, G. Balagura, G. Baynam, A. M. Brower, et al., The Human Phenotype Ontology in 2021, Nucleic Acids Research 49 (D1) (2021) D1207–D1217.doi:10.1093/nar/gkaa1043
2021 doi
-
[12]
N. A. Vasilevsky, N. A. Matentzoglu, S. Toro, J. E. Flack, H. Hegde, D. R. Unni, E. M. Spiegel, S. A. Loom´ıs, N. L. Harris, M. A. Haendel, C. J. Mungall, Mondo: Unifying diseases for the world, by the world, medRxiv (2022). doi:10.1101/2022.04.13.22273750
2022 doi
-
[13]
M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al., The FAIR guiding principles for scientific data management and stew- ardship, Scientific Data 3 (1) (2016) 160018.doi:10.10...
2016 doi
-
[14]
R. L. Seal, B. Braschi, K. Gray, T. E. M. Jones, S. Tweedie, L. Haim-Vilmovsky, E. A. Bruford, HGNC: the HUGO Gene Nomenclature Committee in 2023, Nucleic Acids Research 51 (D1) (2023) D1003–D1009.doi: 10.1093/nar/gkac888
2023 doi
-
[15]
J. A. McMurry, N. Juty, N. Blomberg, T. Burdett, T. Conlin, N. Conte, M. Courtot, J. Deck, M. Dumontier, D. K. Fellows, et al., Identifiers for the 21st century: how to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data, PLO...
2017 doi
-
[16]
Malone, E
J. Malone, E. Holloway, T. Adamusiak, M. Kapushesky, J. Zheng, N. Kolesnikov, A. Zhukova, A. Brazma, H. Parkinson, Modelling sample variables with an Experimental Factor Ontology, Bioinformatics 26 (8) (2010) 1112–1118.doi:10.1093/bioinformatics/btq099
2010 doi
-
[17]
N. F. Noy, N. H. Shah, P. L. Whetzel, B. Dai, M. Dorf, N. Griffith, C. Jonquet, D. L. Rubin, M.-A. Storey, C. G. Chute, M. A. Musen, BioPortal: ontologies and integrated data resources at the click of a button, Nucleic Acids Research 37 (Web Server issue) (2009) W170–W173.doi:...
2009 doi
-
[19]
Szklarczyk, A
D. Szklarczyk, A. L. Gable, K. C. Nastou, D. Lyon, R. Kirsch, S. Pyysalo, N. T. Doncheva, M. Leeb, F. Juhl, L. J. Jensen, P. Bork, C. von Mering, The STRING database in 2021: customizable protein-protein networks and functional characterization of user-uploaded gene/measuremen...
2021 doi
-
[20]
D. R. Unni, S. A. T. Moxon, M. Bada, M. Brush, R. Bruskiewich, J. H. Caufield, P. A. Clemons, V . Dancik, M. Dumontier, K. Fecho, G. Glusman, J. J. Hadlock, N. L. Harris, A. Joshi, T. Putman, G. Qin, S. A. Ramsey, K. A. Shefchek, H. Solbrig, K. Soman, A. E. Thessen, M. A. Haen...
2022
-
[21]
Parciak, B
M. Parciak, B. Vandevoort, F. Neven, L. M. Peeters, S. Vansummeren, Schema matching with large language models: an experimental study, in: VLDB 2024 Workshop: Tabular Data Analysis Workshop (TaDA), 2024. URLhttps://vldb.org/workshops/2024/proceedings/TaDA/TaDA.8.pdf
2024
-
[22]
Mouchel, et al., Interactive data harmonization with LLM agents, arXiv preprint arXiv:2502.07132 (2025)
P.-L. Mouchel, et al., Interactive data harmonization with LLM agents, arXiv preprint arXiv:2502.07132 (2025). 20 URLhttps://arxiv.org/abs/2502.07132
2025 arXiv
-
[23]
Ochoa, A
D. Ochoa, A. Hercules, M. Carmona, D. Suveges, A. Gonzalez-Uriarte, C. Malangone, A. Miranda, L. Fumis, D. Carvalho-Silva, M. Spitzer, et al., Open Targets Platform: supporting systematic drug–target identification and prioritisation, Nucleic Acids Research 49 (D1) (2021) D130...
2021 doi
-
[24]
Moreau, P
L. Moreau, P. Missier, PROV-O: The PROV ontology, W3c recommendation, W3C (2013). URLhttps://www.w3.org/TR/prov-o/
2013
-
[25]
O. J. Wouters, M. McKee, J. Luyten, Estimated research and development investment needed to bring a new medicine to market, 2009–2018, JAMA 323 (9) (2020) 844–853.doi:10.1001/jama.2020.1166
2020
-
[26]
C. H. Wong, K. W. Siah, A. W. Lo, Estimation of clinical trial success rates and related parameters, Biostatistics 20 (2) (2019) 273–286.doi:10.1093/biostatistics/kxx069
2019 doi
-
[27]
Pushpakom, F
S. Pushpakom, F. Iorio, P. A. Eyers, K. J. Escott, S. Hopper, A. Wells, A. Doig, T. Guilliams, J. Latimer, C. Mc- Namee, et al., Drug repurposing: progress, challenges and recommendations, Nature Reviews Drug Discovery 18 (1) (2019) 41–58.doi:10.1038/nrd.2018.168
2019 doi
-
[28]
Chakravarty, J
D. Chakravarty, J. Gao, S. M. Phillips, R. Kundra, H. Zhang, J. Wang, J. E. Rudolph, R. Yaeger, T. Soumerai, M. H. Nissan, et al., OncoKB: A precision oncology knowledge base, JCO Precision Oncology 1 (2017) 1–16. doi:10.1200/PO.17.00011
2017 doi
-
[29]
K. P. Venkatesh, M. M. Raza, J. C. Kvedar, Health digital twins as tools for precision medicine: considera- tions for computation, implementation, and regulation, npj Digital Medicine 5 (1) (2022) 150.doi:10.1038/ s41746-022-00694-7
2022
-
[30]
Laubenbacher, A
R. Laubenbacher, A. Niarakis, T. Helikar, G. Lee, R. Srivastava, M. Blinov, M. Birtwistle, S. Finley, A. Luo, H. R. Chamberlin, et al., Building digital twins of the human immune system: toward a roadmap, npj Digital Medicine 5 (1) (2022) 64.doi:10.1038/s41746-022-00610-z
2022 doi
-
[31]
Karlebach, R
G. Karlebach, R. Shamir, Modelling and analysis of gene regulatory networks, Nature Reviews Molecular Cell Biology 9 (10) (2008) 770–780.doi:10.1038/nrm2503
2008 doi
-
[32]
E. C. Wood, A. K. Glen, L. G. Kvarfordt, F. Womack, L. Acevedo, T. S. Yoon, C. Ma, V . Flores, M. Sinha, Y . Chod- pathumwan, A. Termehchy, J. C. Roach, L. Mendoza, A. S. Hoffman, E. W. Deutsch, D. Koslicki, S. A. Ramsey, RTX-KG2: a system for building a semantically standardi...
2022 doi
-
[33]
P. M. Visscher, N. R. Wray, Q. Zhang, P. Sklar, M. I. McCarthy, M. A. Brown, J. Yang, 10 years of GW AS discovery: biology, function, and translation, American Journal of Human Genetics 101 (1) (2017) 5–22.doi: 10.1016/j.ajhg.2017.06.005
2017 doi
-
[34]
Buniello, J
A. Buniello, J. A. L. MacArthur, M. Cerezo, L. W. Harris, J. Hayhurst, C. Malangone, A. McMahon, J. Morales, E. Mountjoy, E. Sollis, et al., The NHGRI-EBI GW AS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019, Nucleic Acids Res...
2019
-
[35]
D. S. Himmelstein, A. Lizee, C. Hessler, L. Brueggeman, S. L. Chen, D. Hadley, A. Green, P. Khankhanian, S. E. Baranzini, Systematic integration of biomedical knowledge prioritizes drugs for repurposing, eLife 6 (2017) e26726.doi:10.7554/eLife.26726
2017 doi
-
[36]
C. J. Mungall, J. A. McMurry, S. K ¨ohler, J. P. Balhoff, C. Borromeo, M. Brush, S. Carbon, T. Conlin, N. Dunn, M. Engelstad, et al., The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species, Nucleic Acids Research 45 ...
2017 doi
-
[37]
K. A. Shefchek, N. L. Harris, M. Gargano, N. Matentzoglu, D. Unni, M. Brush, D. Keith, T. Conlin, N. Vasilevsky, X. A. Zhang, et al., The Monarch Initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species, Nucleic Acids Rese...
2020 doi
-
[38]
Chandak, K
P. Chandak, K. Huang, M. Zitnik, Building a knowledge graph to enable precision medicine, Scientific Data 10 (2023) 67.doi:10.1038/s41597-023-01960-3. URLhttps://www.nature.com/articles/s41597-023-01960-3
2023 doi
-
[39]
Zitnik, M
M. Zitnik, M. Agrawal, J. Leskovec, Modeling polypharmacy side effects with graph convolutional networks, Bioinformatics 34 (13) (2018) i457–i466.doi:10.1093/bioinformatics/bty294
2018 doi
-
[40]
J. Zhang, et al., A comprehensive large-scale biomedical knowledge graph for AI-powered data-driven biomedical research, Nature Machine IntelligenceAlso available as bioRxiv preprint:https://doi.org/10.1101/2023. 10.13.562216(2025).doi:10.1038/s42256-025-01014-w. URLhttps://ww...
2025 doi
-
[41]
M. R. Hossain, et al., A knowledge graph approach for the secondary use of cancer registry data, in: Proceedings of the IEEE International Conference on Big Data, 2019, oSTI.GOV report; describes the Louisiana Tumor Registry KG with 25+billion triples and∼4 TB storage. URLhttp...
2019
-
[42]
M. R. Hossain, et al., Knowledge graph-enabled cancer data analytics, IEEE Journal of Biomedical and Health Informatics (2021).doi:10.1109/JBHI.2021.3077810. 21 URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC8324069/
2021
-
[43]
H. Bast, B. Buchhold, QLever: A query engine for efficient SPARQL+text search, in: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (CIKM’17), ACM, 2017, pp. 647–656.doi: 10.1145/3132847.3132921. URLhttps://doi.org/10.1145/3132847.3132921
2017
-
[44]
M. R. Ackermann, H. Bast, B. M. Beckermann, J. Kalmbach, P. Neises, S. Ollinger, The dblp knowledge graph and SPARQL endpoint, Transactions on Graph Data and Knowledge (TGDK) 2 (2) (2024) 3:1–3:23.doi:10.4230/ TGDK.2.2.3. URLhttps://doi.org/10.4230/TGDK.2.2.3
2024 doi
-
[45]
Mohseni Ahooyi, B
T. Mohseni Ahooyi, B. Stear, J. A. Simmons, V . T. Metzger, P. Kumar, J. E. Evangelista, D. J. Clarke, Z. Xie, H. Kim, S. L. Jenkins, et al., The data distillery: A graph framework for semantic integration and querying of biomedical data, bioRxivAlso published in PMC:https://p...
2025 doi
-
[46]
B. J. Stear, T. Mohseni Ahooyi, J. A. Simmons, C. Kollar, L. Hartman, K. Beigel, A. Lahiri, S. Vasisht, T. J. Callahan, C. M. Nemarich, J. C. Silverstein, D. M. Taylor, Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data...
2024
-
[48]
Hofer, D
M. Hofer, D. Obraczka, A. Saeedi, H. K ¨opcke, E. Rahm, Construction of knowledge graphs: Current state and challenges, Information 15 (8) (2024) 509.doi:10.3390/info15080509. URLhttps://www.mdpi.com/2078-2489/15/8/509
2024 doi
-
[49]
K. G. Cortes, S. Sundar, S. Gehrke, K. Manpearl, J. Lin, D. R. Korn, H. Caufield, K. Schaper, J. Reese, K. Koirala, L. E. Hunter, E. K. Carter, M. DeLuca, A. Krishnan, C. Mungall, M. Haendel, Improving biomedical knowledge graph quality: A community approach, arXiv preprint (2...
2025 arXiv
-
[50]
Wratten, A
L. Wratten, A. Wilm, J. G ¨oke, Reproducible, scalable, and shareable analysis pipelines with bioinformatics work- flow managers, Nature Methods 18 (10) (2021) 1161–1168.doi:10.1038/s41592-021-01254-9
2021 doi
-
[51]
Fecho, G
K. Fecho, G. Glusman, S. E. Baranzini, et al., Announcing the biomedical data translator: Initial public release, Clinical and Translational Science 18 (7) (2025) e70284.doi:10.1111/cts.70284. URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC12241707/
2025 doi
-
[52]
Djaffardjy, G
M. Djaffardjy, G. Marchment, C. Sebe, R. Blanchet, K. Belhajjame, A. Gaignard, F. Lemoine, S. Cohen-Boulakia, Developing and reusing bioinformatics data analysis pipelines using scientific workflow systems, Computational and Structural Biotechnology Journal 21 (2023) 2075–2085...
2023 doi
-
[53]
Lobentanzer, A
S. Lobentanzer, A. Koleti, P. Brack, J. Baumbach, J. Saez-Rodriguez, P. Anagnostopoulou, et al., BioCypher: a unifying framework for biomedical research knowledge graphs, Nature Biotechnology 41 (9) (2023) 1234–1237. doi:10.1038/s41587-023-01848-y
2023 doi
-
[54]
Sporny, D
M. Sporny, D. Longley, G. Kellogg, M. Lanthaler, N. Lindstr ¨om, JSON-LD 1.1: A JSON-based serialization for linked data, W3c recommendation, W3C (2020). URLhttps://www.w3.org/TR/json-ld11/
2020
-
[55]
Knublauch, D
H. Knublauch, D. Kontokostas, Shapes Constraint Language (SHACL), W3c recommendation, W3C (2017). URLhttps://www.w3.org/TR/shacl/
2017
-
[56]
M. L. Neal, M. K ¨onig, D. Nickerson, J. M´ıˇsek, A. Kiriˇcˇsi, I. I. Moraru, C. J. Myers, J. R. Faeder, et al., Harmonizing semantic annotations for computational models in biology, Briefings in Bioinformatics 20 (2) (2019) 540–550. doi:10.1093/bib/bby087
2019 doi
-
[57]
Di Tommaso, M
P. Di Tommaso, M. Chatzou, E. W. Floden, P. P. Barja, E. Palumbo, C. Notredame, Nextflow enables reproducible computational workflows, Nature Biotechnology 35 (4) (2017) 316–319.doi:10.1038/nbt.3820
2017 doi
-
[58]
M ¨older, K
F. M ¨older, K. P. Jablonski, B. Letcher, M. B. Hall, C. H. Tomkins-Tinch, V . Sochat, J. Forster, S. Lee, S. O. Twardziok, A. Kanitz, et al., Sustainable data analysis with Snakemake, F1000Research 10 (2021) 33.doi: 10.12688/f1000research.29032.2
2021 doi
-
[59]
A. E. Ahmed, J. M. Allen, T. Bhat, P. Burra, C. E. Fliege, S. N. Hart, J. R. Heldenbrand, M. E. Hudson, D. D. Istanto, M. T. Kalmbach, et al., Design considerations for workflow management systems use in production genomics research and the clinic, Scientific Reports 11 (1) (2...
2021 doi
-
[60]
Forer, S
L. Forer, S. Sch ¨onherr, Improving the reliability, quality, and maintainability of bioinformatics pipelines with nf- test, GigaSciencePreprint posted May 2024 at bioRxiv, doi:10.1101/2024.05.25.595877 (2025).doi:10.1093/ gigascience/giaf130
2025 doi
-
[61]
Patel, C
Y . Patel, C. Zhu, T. N. Yamaguchi, Y . Z. Bugh, M. Tian, A. Holmes, S. T. Fitz-Gibbon, P. C. Boutros, NFTest: au- tomated testing of Nextflow pipelines, Bioinformatics 40 (2) (2024) btae081.doi:10.1093/bioinformatics/ 22 btae081
2024 doi
-
[62]
H ¨ansel, S
K. H ¨ansel, S. N. Dudgeon, K.-H. Cheung, T. J. S. Durant, W. L. Schulz, From data to wisdom: Biomedical knowledge graphs for real-world data insights, Journal of Medical Systems 47 (1) (2023) 65.doi:10.1007/ s10916-023-01951-2. URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC10191934/
2023
-
[63]
Cheng, Y
H. Cheng, Y . Wu, S. Khatwani, M. Kruse, D. Dligach, T. A. Miller, M. Afshar, Y . Gao, Scaling biomedical knowl- edge graph retrieval for interpretable reasoning: Applications to clinical diagnosis prediction, medRxivPreprint (2026).doi:10.64898/2026.01.12.26343957. URLhttps:/...
2026 doi
-
[64]
P. A. Ewels, A. Peltzer, S. Fillinger, H. Patel, J. Alneberg, A. Wilm, M. U. Garcia, P. Di Tommaso, S. Nahnsen, The nf-core framework for community-curated bioinformatics pipelines, Nature Biotechnology 38 (3) (2020) 276–278. doi:10.1038/s41587-020-0439-x
2020 doi
-
[65]
rep., Anthropic, open standard for connecting AI assistants to external data sources and tools
Anthropic, Model context protocol, Tech. rep., Anthropic, open standard for connecting AI assistants to external data sources and tools. Specification and SDKs available athttps://modelcontextprotocol.io(Nov. 2024). URLhttps://www.anthropic.com/news/model-context-protocol 23
2024
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.