Pith. sign in

REVIEW 3 major objections 6 minor 65 references

Biomedical Knowledge Composition: A Software Engineering Perspective

T0 review · 3 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Biomedical knowledge graphs are hard to build because data pipelines lack the software engineering tooling that made web development composable and reproducible.

desk verdict A genuinely useful synthesis and research agenda for biomedical KG engineering, with a central causal claim that remains a clearly-labeled hypothesis rather than an established finding. read the letter →

arxiv 2608.08927 v1 pith:WEHLWUEC submitted 2026-08-09 cs.SE

classification cs.SE
keywords biomedicalknowledgegraphsdataharmonizationsoftwareengineeringreproducibilityFAIRinfrastructuregraphconstructionintegration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the difficulty of building biomedical knowledge graphs is substantially an engineering problem, not just a science problem. The field, the authors claim, has not adopted the package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance that made web engineering reliably composable. The argument is grounded in a description of five data-harmonization challenges, profiles of six deployed systems with contrasting pipeline reproducibility, and the authors' own attempt to reuse a large harmonized knowledge graph, which ran into redeployment, discoverability, and coverage obstacles. If the diagnosis is right, then investing in these practices—and in a culture of shipping reproducible build pipelines rather than static graph snapshots—would materially reduce the cost and increase the reliability of biomedical knowledge graph construction.

What carries the argument

The central object is the engineering maturity gap: the comparison between the standardized tooling of web development—versioned packages, typed interfaces, dependency resolution, continuous integration—and the bespoke, largely manual practice of biomedical data integration. The paper uses this gap as a diagnostic lens, framing each of its eight open challenges as a missing analogue of a web-engineering capability, from a data package manager and namespace type safety to a canonical intermediate representation and production-readiness engineering. The gap also does prescriptive work: it makes 'pipeline over artifact' the first-class design criterion, so that a versioned, documented build pipeline counts as more valuable than a static deployed graph. That criterion organizes the profiles of the six systems and the design principles the authors recommend for new projects.

What would settle it

A systematic comparison of biomedical knowledge-graph projects that use versioned, package-managed pipelines versus those that consume static snapshots—controlling for data scale and domain—and finds no significant difference in integration cost, error rate, or update speed would falsify the claim that tooling adoption is a root cause.

Watch

Extended reading notes

Core claim

The paper's central claim is that a contributing root cause of the difficulty in biomedical knowledge infrastructure is the limited adoption of software engineering tooling and practices that make web engineering reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. The paper argues that this engineering maturity gap is what turns the assembly of a biomedical knowledge graph into brittle, largely unrepeatable manual work. The corollary the authors draw is that the field should shift its emphasis from deployed knowledge graphs to the reproducible process of assembling them: versioned dependencies, reusable build pipelines, and engineering practices that let others compile and customize a graph from source rather than consume a static artifact. The supporting evidence is a taxonomy of five harmonization challenges, profiles of six representative systems arranged along a reproducibility spectrum, a concrete reuse attempt that exposed redeployment, discoverability, and coverage failures, and a catalogue of eight open engineering challenges, each with partial solutions but no universally adopted stack.

Load-bearing premise

The load-bearing premise is that the obstacles the authors met while reusing one large harmonized knowledge graph are structural features of the biomedical data problem rather than quirks of that particular system, so if that experience is idiosyncratic the general diagnosis loses much of its force.

Editorial extensions

If this is right

  • If the diagnosis is correct, a versioned 'package manager' for biomedical data releases would eliminate a major source of silent breakage in graph assembly, much as dependency managers did for software.
  • Teams that publish reproducible build pipelines instead of static exports would make schema changes, source updates, and provenance queries tractable engineering tasks, shifting maintenance burden away from downstream consumers.
  • Namespace-aware type checking would turn mismatched identifier joins—which currently corrupt graphs silently—into compile-time errors that are caught before a graph is shipped.
  • A standardized, validated graph interchange format with a provenance subgraph would allow continuous-integration systems to verify knowledge-graph assembly compliance automatically.
  • The eight open challenges collectively define a research agenda in which software engineers, not only biomedical curators, can make direct contributions to knowledge infrastructure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the maturity-gap thesis is right, a testable prediction follows: projects that adopt versioned, package-managed pipelines should show measurably lower integration cost and fewer silent errors than projects that consume static snapshots, when scale and domain are controlled for.
  • The argument implies that the long tail of specialized databases will not be cured by more ontologies alone; the leverage lies in standardizing release mechanics—version numbers, machine-readable changelogs, integrity hashes—across data providers.
  • The paper's web-engineering comparison is qualitative, so a quantitative gap analysis measuring what fraction of knowledge-graph construction steps are covered by versioned tooling across a sample of projects could upgrade the diagnosis into a measurement.
  • The proposed stack may also be a prerequisite for LLM-based agents that assemble or update graphs: reliable agentic data integration needs exactly the typed, versioned, auditable interfaces the paper calls for.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is a position/synthesis article that combines a tutorial on biomedical knowledge graphs (KGs) with a software-engineering critique of how they are assembled. The first half introduces the domain: why KGs are the central integrative data structure, five data-harmonization challenges, application areas such as drug discovery and digital twins, and six representative KG systems. The second half argues that a contributing root cause of the field's difficulty is limited adoption of software practices that make web engineering reliably composable and reproducible—specifically package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. The argument is supported by a qualitative comparison with web engineering, a first-person account of three obstacles encountered when reusing the Data Distillery KG, and references to related friction in the literature. The paper then formulates eight open engineering challenges and a call to shift from shipping static KG artifacts to shipping reproducible build pipelines. Section 11 explicitly acknowledges that the web comparison is qualitative and that the Data Distillery account reflects one team's experience with one version of one system.

Significance. If its central thesis is accepted, the paper provides a useful bridge between software engineering and biomedical knowledge infrastructure and could redirect investment toward data package management, namespace type enforcement, canonical interchange representations, service composition, and lifecycle governance. Its concrete strengths are the operationalized list of eight challenges with partial solutions, the explicit treatment of pipeline reproducibility as a first-class design criterion, the concrete Data Distillery obstacles, and an unusually candid limitations section. The paper does not provide a controlled empirical demonstration; the causal claim is a well-formed hypothesis rather than an established result. The reliance on the authors' own prior work [18] and their own Data Distillery experience [47] limits evidential independence but does not make the argument circular. As a position paper and research agenda, the contribution is valuable; as a demonstrated root-cause analysis, it is not yet supported.

major comments (3)
  1. [Section 6, paragraph beginning 'These obstacles are not unique...'] The inference from the three Data Distillery obstacles to structural properties is load-bearing for the paper's central claim. The three obstacles are an HPC root-privilege policy for Docker, a schema with a single node label and roughly 1600 relationship types, and incomplete harmonization coverage. These are respectively a deployment-policy constraint, a schema-design choice, and a curation-scope decision; none of them directly demonstrates the absence of package management, typed namespaces, canonical interchange formats, or lifecycle governance. The cited references [48-50] document related problems, but they provide no comparative measurement showing that projects adopting the proposed practices perform better. To make the root-cause claim defensible, the paper should either reframe the abstract and Section 7 as presenting a hypothesis with an explicit evaluation design (for example, a structured survey scoring KG projects on the six proposed practices and correlating those scores with reproducibility or integration-error outcomes), or add such evidence.
  2. [Section 7 and Section 8.1] The maturity-gap argument overstates the green-field character of the problem. General-purpose data versioning and packaging tools such as DataLad, DVC, Quilt, and git-annex already provide semantic versioning, integrity hashes, dependency graphs, and registry-like distribution. The real gap is the absence of a domain-specific convention for biomedical data releases, including ontology-aware schemas, identifier namespaces, and release semantics, not the total absence of package-manager concepts. The open problem in Section 8.1 should engage with these existing tools and explain why they do not transfer directly to biomedical KG assembly. Without this engagement, the claim that the biomedical landscape 'lacks' these capabilities is too strong and weakens the otherwise plausible adoption-based diagnosis.
  3. [Abstract, Section 7, and Section 11] The abstract and Section 7 use 'contributing root cause' and 'the engineering infrastructure gap is real' as definitive statements, while Section 11 concedes that the comparison is qualitative and that the Data Distillery case reflects one team's experience. Because the paper's contribution is a research agenda rather than a completed empirical study, the causal language should be explicitly hedged in the abstract and Section 7 (for example, 'a likely contributor' or 'a hypothesis supported by our experience and related reports'), with Section 11's caveats reflected consistently throughout. This change would make the paper more accurate without diminishing its value as a call to action.
minor comments (6)
  1. [Abstract] The phrase 'toward thereproducible process' is missing a space and should read 'toward the reproducible process.'
  2. [Section 2 and Figure 1] Figure 1 presents four challenge categories, but Section 3 introduces five core harmonization challenges; the relationship between the two partitions should be stated explicitly so that readers do not see an inconsistency.
  3. [Listings 1 and 2] The JSON keys in Listing 1 and the process name in Listing 2 appear with inserted spaces (for example, 'q u e r y _ g r a p h' and 'NO RM AL IS E'); if these are not rendering artifacts, the listings should be cleaned up.
  4. [Table 2, Reactome row] The 'Pipeline reproducibility' cell for Reactome describes the export format and the graph's usefulness for mechanistic modeling, but does not state whether the build/export pipeline can be re-run; this should be aligned with the other rows.
  5. [Section 10] The visualization section is interesting but only loosely connected to the eight SE challenges; a sentence linking semantic visualization to pipeline reproducibility and lifecycle governance would strengthen the integration.
  6. [Section 10, 'apinatomy'] The text mentions 'apinatomy panels' but does not provide a citation for the apinatomy framework; a reference should be added.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claim is an interpretive synthesis supported by case examples, not a derivation that reduces to its own inputs.

full rationale

The paper makes no quantitative predictions and fits no parameters, so none of the fitted-input or self-definitional circularity patterns apply. Its central thesis—that limited adoption of software-engineering practices such as package management, typed namespaces, canonical interchange formats, service composition, reproducible pipelines, and lifecycle governance is a contributing cause of biomedical knowledge-graph construction difficulty—is presented as an argued interpretation supported by domain examples, profiles of six KG systems, and a single case study. Section 11 explicitly concedes that the comparison to web engineering is a qualitative observation rather than a systematic empirical study and that the Data Distillery account reflects one team's experience with one version of one system. The only author self-citation is reference [18], which is used to note that semantic-mapping methods assist entity resolution; this is not load-bearing for the article's central argument. External references [48–50] independently document similar frictions, and the paper does not invoke any author-specific uniqueness theorem or disguise a fitted parameter as a prediction. The article's recommendations are a research agenda and an interpretive synthesis, not conclusions forced by construction from their premises.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a review/position piece and introduces no free parameters or invented entities. Its axioms are domain-level framing assumptions and one case-study generalization, all acknowledged qualitatively in the text.

assumptions (3)
  • domain assumption Knowledge graphs are the central integrative data structure for biomedicine.
    Stated in the abstract and Section 1 without comparing against alternative integration structures; the paper is a review of knowledge graph practice, not a comparative study.
  • domain assumption Web engineering tooling is a valid template for biomedical data integration.
    The entire engineering maturity gap argument (Section 7) depends on this analogy; Section 11 labels the comparison qualitative.
  • ad hoc to paper The friction experienced with Data Distillery is representative of structural problems in the field.
    Section 6 generalizes from one case study to a structural claim. This is the paper's own experience, so it is a self-referential support and the generalization is not independently tested.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Biomedical Knowledge Composition: A Software Engineering Perspective." pith.science (2026). https://pith.science/paper/WEHLWUEC

@misc{pith2026260808927,
  author       = {Pith},
  title        = {Pith review of: Biomedical Knowledge Composition: A Software Engineering Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WEHLWUEC}},
  note         = {Machine review of arXiv:2608.08927}
}
read the original abstract

Biomedical research has accumulated vast molecular, clinical, and population data, yet translating this wealth into actionable knowledge remains constrained by technical and organizational difficulties. This article presents a unified treatment of two perspectives on biomedical knowledge infrastructure. The first introduces the biomedical domain to software engineers: it explains why knowledge graphs (KGs) are the central integrative data structure in modern biomedicine, characterizes five data harmonization challenges (identifier mapping, entity resolution, schema alignment, evidence integration, and provenance tracking), surveys application domains from drug discovery to digital twins, and profiles six representative KG systems with contrasting choices. The second perspective asks why engineering biomedical knowledge infrastructure remains so difficult. We argue that a contributing root cause is limited adoption of software tooling and practices that make development in other mature domains - particularly web engineering - reliably composable and reproducible: package management, typed namespaces, canonical interchange formats, service composition protocols, reproducible pipelines, and lifecycle governance. Against this backdrop, eight open engineering challenges for biomedical data integration are catalogued, each with partial solutions but no universally adopted stack. Crucially, the article shifts emphasis from describing deployed KG instances toward the reproducible process of assembling them: reusable build pipelines, versioned dependencies, and engineering practices that let others compile and customize a KG from source rather than consuming a static artifact. Together, the two perspectives provide domain grounding for newcomers and a research agenda for software engineers seeking to make transformative contributions to biomedical knowledge infrastructure.

Figures

Figures reproduced from arXiv: 2608.08927 by the authors.

Figure 1
Figure 1. Taxonomy of biomedical data harmonization challenges spanning four orthogonal problem categories: [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Engineering pipelines for KG-based biomedical knowledge infrastructure. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Levels of graph-based biomedical data visualization, from no e [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

65 extracted references · 26 canonical work pages

  1. [18]

    Kokash, L

    N. Kokash, L. Wang, T. H. Gillespie, A. Belloum, P. Grosso, S. Quinney, L. Li, B. de Bono, Ontology- and LLM- based data harmonization for federated learning in healthcare, Frontiers in Digital Health 8, health Informatics section (2026).doi:10.3389/fdgth.2026.1756555

  2. [47]

    Taylor Research Lab, CFDE Data Distillery Project,https://github.com/TaylorResearchLab/CFDE_ DataDistillery, accessed: 2026-05-21 (2024)

  3. [1]

    Tomczak, P

    K. Tomczak, P. Czerwi ´nska, M. Wiznerowicz, The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge, Contemporary Oncology 19 (1A) (2015) A68–A77.doi:10.5114/wo.2014.47136

  4. [2]

    Regev, S

    A. Regev, S. A. Teichmann, E. S. Lander, I. Amit, C. Benoist, E. Birney, et al., The Human Cell Atlas, eLife 6 (2017) e27041.doi:10.7554/eLife.27041. 19

  5. [3]

    Gene Ontology Consortium, The Gene Ontology Resource: 20 years and still GOing strong, Nucleic Acids Re- search 47 (D1) (2019) D330–D338.doi:10.1093/nar/gky1055

  6. [4]

    UniProt Consortium, UniProt: the Universal Protein Knowledgebase in 2023, Nucleic Acids Research 51 (D1) (2023) D523–D531.doi:10.1093/nar/gkac1052

  7. [5]

    Kanehisa, M

    M. Kanehisa, M. Furumichi, Y . Sato, M. Kawashima, M. Ishiguro-Watanabe, KEGG for taxonomy-based analysis of pathways and genomes, Nucleic Acids Research 51 (D1) (2023) D587–D592.doi:10.1093/nar/gkac963

  8. [6]

    Jassal, L

    B. Jassal, L. Matthews, G. Viteri, C. Gong, P. Lorente, A. Fabregat, K. Sidiropoulos, J. Cook, M. Gillespie, R. Haw, et al., The Reactome pathway knowledgebase, Nucleic Acids Research 48 (D1) (2020) D498–D503.doi:10. 1093/nar/gkz1031

Show all 65 references
  1. [7]

    D. S. Wishart, Y . D. Feunang, A. C. Guo, E. J. Lo, A. Marcu, J. R. Grant, T. Sajed, D. Johnson, C. Li, Z. Sayeeda, et al., DrugBank 5.0: a major update to the DrugBank database for 2018, Nucleic Acids Research 46 (D1) (2018) D1074–D1082.doi:10.1093/nar/gkx1037

  2. [8]

    Mendez, A

    D. Mendez, A. Gaulton, A. P. Bento, J. Chambers, M. De Veij, E. F ´elix, M. P. Magari˜nos, J. F. Mosquera, P. Mu- towo, M. Nowotka, et al., ChEMBL: towards direct deposition of bioassay data, Nucleic Acids Research 47 (D1) (2019) D930–D940.doi:10.1093/nar/gky1075

  3. [9]

    M. J. Landrum, J. M. Lee, M. Benson, G. R. Brown, C. Chao, S. Chitipiralla, B. Gu, J. Hart, D. Hoffman, W. Jang, et al., ClinVar: improving access to variant interpretations and supporting evidence, Nucleic Acids Research 46 (D1) (2018) D1062–D1067.doi:10.1093/nar/gkx1153

  4. [10]

    J. G. Tate, S. Bamford, H. C. Jubb, Z. Sondka, D. M. Beare, N. Bindal, H. Boutselakis, C. G. Cole, C. Creatore, E. Dawson, et al., COSMIC: the catalogue of somatic mutations in cancer, Nucleic Acids Research 47 (D1) (2019) D941–D947.doi:10.1093/nar/gky1015

  5. [11]

    K ¨ohler, M

    S. K ¨ohler, M. Gargano, N. Matentzoglu, L. C. Carmody, D. Lewis-Smith, N. A. Vasilevsky, D. Danis, G. Balagura, G. Baynam, A. M. Brower, et al., The Human Phenotype Ontology in 2021, Nucleic Acids Research 49 (D1) (2021) D1207–D1217.doi:10.1093/nar/gkaa1043

  6. [12]

    N. A. Vasilevsky, N. A. Matentzoglu, S. Toro, J. E. Flack, H. Hegde, D. R. Unni, E. M. Spiegel, S. A. Loom´ıs, N. L. Harris, M. A. Haendel, C. J. Mungall, Mondo: Unifying diseases for the world, by the world, medRxiv (2022). doi:10.1101/2022.04.13.22273750

  7. [13]

    M. D. Wilkinson, M. Dumontier, I. J. Aalbersberg, G. Appleton, M. Axton, A. Baak, N. Blomberg, J.-W. Boiten, L. B. da Silva Santos, P. E. Bourne, et al., The FAIR guiding principles for scientific data management and stew- ardship, Scientific Data 3 (1) (2016) 160018.doi:10.10...

  8. [14]

    R. L. Seal, B. Braschi, K. Gray, T. E. M. Jones, S. Tweedie, L. Haim-Vilmovsky, E. A. Bruford, HGNC: the HUGO Gene Nomenclature Committee in 2023, Nucleic Acids Research 51 (D1) (2023) D1003–D1009.doi: 10.1093/nar/gkac888

  9. [15]

    J. A. McMurry, N. Juty, N. Blomberg, T. Burdett, T. Conlin, N. Conte, M. Courtot, J. Deck, M. Dumontier, D. K. Fellows, et al., Identifiers for the 21st century: how to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data, PLO...

  10. [16]

    Malone, E

    J. Malone, E. Holloway, T. Adamusiak, M. Kapushesky, J. Zheng, N. Kolesnikov, A. Zhukova, A. Brazma, H. Parkinson, Modelling sample variables with an Experimental Factor Ontology, Bioinformatics 26 (8) (2010) 1112–1118.doi:10.1093/bioinformatics/btq099

  11. [17]

    N. F. Noy, N. H. Shah, P. L. Whetzel, B. Dai, M. Dorf, N. Griffith, C. Jonquet, D. L. Rubin, M.-A. Storey, C. G. Chute, M. A. Musen, BioPortal: ontologies and integrated data resources at the click of a button, Nucleic Acids Research 37 (Web Server issue) (2009) W170–W173.doi:...

  12. [19]

    Szklarczyk, A

    D. Szklarczyk, A. L. Gable, K. C. Nastou, D. Lyon, R. Kirsch, S. Pyysalo, N. T. Doncheva, M. Leeb, F. Juhl, L. J. Jensen, P. Bork, C. von Mering, The STRING database in 2021: customizable protein-protein networks and functional characterization of user-uploaded gene/measuremen...

  13. [20]

    D. R. Unni, S. A. T. Moxon, M. Bada, M. Brush, R. Bruskiewich, J. H. Caufield, P. A. Clemons, V . Dancik, M. Dumontier, K. Fecho, G. Glusman, J. J. Hadlock, N. L. Harris, A. Joshi, T. Putman, G. Qin, S. A. Ramsey, K. A. Shefchek, H. Solbrig, K. Soman, A. E. Thessen, M. A. Haen...

  14. [21]

    Parciak, B

    M. Parciak, B. Vandevoort, F. Neven, L. M. Peeters, S. Vansummeren, Schema matching with large language models: an experimental study, in: VLDB 2024 Workshop: Tabular Data Analysis Workshop (TaDA), 2024. URLhttps://vldb.org/workshops/2024/proceedings/TaDA/TaDA.8.pdf

  15. [22]

    Mouchel, et al., Interactive data harmonization with LLM agents, arXiv preprint arXiv:2502.07132 (2025)

    P.-L. Mouchel, et al., Interactive data harmonization with LLM agents, arXiv preprint arXiv:2502.07132 (2025). 20 URLhttps://arxiv.org/abs/2502.07132

  16. [23]

    Ochoa, A

    D. Ochoa, A. Hercules, M. Carmona, D. Suveges, A. Gonzalez-Uriarte, C. Malangone, A. Miranda, L. Fumis, D. Carvalho-Silva, M. Spitzer, et al., Open Targets Platform: supporting systematic drug–target identification and prioritisation, Nucleic Acids Research 49 (D1) (2021) D130...

  17. [24]

    Moreau, P

    L. Moreau, P. Missier, PROV-O: The PROV ontology, W3c recommendation, W3C (2013). URLhttps://www.w3.org/TR/prov-o/

  18. [25]

    O. J. Wouters, M. McKee, J. Luyten, Estimated research and development investment needed to bring a new medicine to market, 2009–2018, JAMA 323 (9) (2020) 844–853.doi:10.1001/jama.2020.1166

  19. [26]

    C. H. Wong, K. W. Siah, A. W. Lo, Estimation of clinical trial success rates and related parameters, Biostatistics 20 (2) (2019) 273–286.doi:10.1093/biostatistics/kxx069

  20. [27]

    Pushpakom, F

    S. Pushpakom, F. Iorio, P. A. Eyers, K. J. Escott, S. Hopper, A. Wells, A. Doig, T. Guilliams, J. Latimer, C. Mc- Namee, et al., Drug repurposing: progress, challenges and recommendations, Nature Reviews Drug Discovery 18 (1) (2019) 41–58.doi:10.1038/nrd.2018.168

  21. [28]

    Chakravarty, J

    D. Chakravarty, J. Gao, S. M. Phillips, R. Kundra, H. Zhang, J. Wang, J. E. Rudolph, R. Yaeger, T. Soumerai, M. H. Nissan, et al., OncoKB: A precision oncology knowledge base, JCO Precision Oncology 1 (2017) 1–16. doi:10.1200/PO.17.00011

  22. [29]

    K. P. Venkatesh, M. M. Raza, J. C. Kvedar, Health digital twins as tools for precision medicine: considera- tions for computation, implementation, and regulation, npj Digital Medicine 5 (1) (2022) 150.doi:10.1038/ s41746-022-00694-7

  23. [30]

    Laubenbacher, A

    R. Laubenbacher, A. Niarakis, T. Helikar, G. Lee, R. Srivastava, M. Blinov, M. Birtwistle, S. Finley, A. Luo, H. R. Chamberlin, et al., Building digital twins of the human immune system: toward a roadmap, npj Digital Medicine 5 (1) (2022) 64.doi:10.1038/s41746-022-00610-z

  24. [31]

    Karlebach, R

    G. Karlebach, R. Shamir, Modelling and analysis of gene regulatory networks, Nature Reviews Molecular Cell Biology 9 (10) (2008) 770–780.doi:10.1038/nrm2503

  25. [32]

    E. C. Wood, A. K. Glen, L. G. Kvarfordt, F. Womack, L. Acevedo, T. S. Yoon, C. Ma, V . Flores, M. Sinha, Y . Chod- pathumwan, A. Termehchy, J. C. Roach, L. Mendoza, A. S. Hoffman, E. W. Deutsch, D. Koslicki, S. A. Ramsey, RTX-KG2: a system for building a semantically standardi...

  26. [33]

    P. M. Visscher, N. R. Wray, Q. Zhang, P. Sklar, M. I. McCarthy, M. A. Brown, J. Yang, 10 years of GW AS discovery: biology, function, and translation, American Journal of Human Genetics 101 (1) (2017) 5–22.doi: 10.1016/j.ajhg.2017.06.005

  27. [34]

    Buniello, J

    A. Buniello, J. A. L. MacArthur, M. Cerezo, L. W. Harris, J. Hayhurst, C. Malangone, A. McMahon, J. Morales, E. Mountjoy, E. Sollis, et al., The NHGRI-EBI GW AS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019, Nucleic Acids Res...

  28. [35]

    D. S. Himmelstein, A. Lizee, C. Hessler, L. Brueggeman, S. L. Chen, D. Hadley, A. Green, P. Khankhanian, S. E. Baranzini, Systematic integration of biomedical knowledge prioritizes drugs for repurposing, eLife 6 (2017) e26726.doi:10.7554/eLife.26726

  29. [36]

    C. J. Mungall, J. A. McMurry, S. K ¨ohler, J. P. Balhoff, C. Borromeo, M. Brush, S. Carbon, T. Conlin, N. Dunn, M. Engelstad, et al., The Monarch Initiative: an integrative data and analytic platform connecting phenotypes to genotypes across species, Nucleic Acids Research 45 ...

  30. [37]

    K. A. Shefchek, N. L. Harris, M. Gargano, N. Matentzoglu, D. Unni, M. Brush, D. Keith, T. Conlin, N. Vasilevsky, X. A. Zhang, et al., The Monarch Initiative in 2019: an integrative data and analytic platform connecting phenotypes to genotypes across species, Nucleic Acids Rese...

  31. [38]

    Chandak, K

    P. Chandak, K. Huang, M. Zitnik, Building a knowledge graph to enable precision medicine, Scientific Data 10 (2023) 67.doi:10.1038/s41597-023-01960-3. URLhttps://www.nature.com/articles/s41597-023-01960-3

  32. [39]

    Zitnik, M

    M. Zitnik, M. Agrawal, J. Leskovec, Modeling polypharmacy side effects with graph convolutional networks, Bioinformatics 34 (13) (2018) i457–i466.doi:10.1093/bioinformatics/bty294

  33. [40]

    J. Zhang, et al., A comprehensive large-scale biomedical knowledge graph for AI-powered data-driven biomedical research, Nature Machine IntelligenceAlso available as bioRxiv preprint:https://doi.org/10.1101/2023. 10.13.562216(2025).doi:10.1038/s42256-025-01014-w. URLhttps://ww...

  34. [41]

    M. R. Hossain, et al., A knowledge graph approach for the secondary use of cancer registry data, in: Proceedings of the IEEE International Conference on Big Data, 2019, oSTI.GOV report; describes the Louisiana Tumor Registry KG with 25+billion triples and∼4 TB storage. URLhttp...

  35. [42]

    M. R. Hossain, et al., Knowledge graph-enabled cancer data analytics, IEEE Journal of Biomedical and Health Informatics (2021).doi:10.1109/JBHI.2021.3077810. 21 URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC8324069/

  36. [43]

    H. Bast, B. Buchhold, QLever: A query engine for efficient SPARQL+text search, in: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management (CIKM’17), ACM, 2017, pp. 647–656.doi: 10.1145/3132847.3132921. URLhttps://doi.org/10.1145/3132847.3132921

  37. [44]

    M. R. Ackermann, H. Bast, B. M. Beckermann, J. Kalmbach, P. Neises, S. Ollinger, The dblp knowledge graph and SPARQL endpoint, Transactions on Graph Data and Knowledge (TGDK) 2 (2) (2024) 3:1–3:23.doi:10.4230/ TGDK.2.2.3. URLhttps://doi.org/10.4230/TGDK.2.2.3

  38. [45]

    Mohseni Ahooyi, B

    T. Mohseni Ahooyi, B. Stear, J. A. Simmons, V . T. Metzger, P. Kumar, J. E. Evangelista, D. J. Clarke, Z. Xie, H. Kim, S. L. Jenkins, et al., The data distillery: A graph framework for semantic integration and querying of biomedical data, bioRxivAlso published in PMC:https://p...

  39. [46]

    B. J. Stear, T. Mohseni Ahooyi, J. A. Simmons, C. Kollar, L. Hartman, K. Beigel, A. Lahiri, S. Vasisht, T. J. Callahan, C. M. Nemarich, J. C. Silverstein, D. M. Taylor, Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data...

  40. [48]

    Hofer, D

    M. Hofer, D. Obraczka, A. Saeedi, H. K ¨opcke, E. Rahm, Construction of knowledge graphs: Current state and challenges, Information 15 (8) (2024) 509.doi:10.3390/info15080509. URLhttps://www.mdpi.com/2078-2489/15/8/509

  41. [49]

    K. G. Cortes, S. Sundar, S. Gehrke, K. Manpearl, J. Lin, D. R. Korn, H. Caufield, K. Schaper, J. Reese, K. Koirala, L. E. Hunter, E. K. Carter, M. DeLuca, A. Krishnan, C. Mungall, M. Haendel, Improving biomedical knowledge graph quality: A community approach, arXiv preprint (2...

  42. [50]

    Wratten, A

    L. Wratten, A. Wilm, J. G ¨oke, Reproducible, scalable, and shareable analysis pipelines with bioinformatics work- flow managers, Nature Methods 18 (10) (2021) 1161–1168.doi:10.1038/s41592-021-01254-9

  43. [51]

    Fecho, G

    K. Fecho, G. Glusman, S. E. Baranzini, et al., Announcing the biomedical data translator: Initial public release, Clinical and Translational Science 18 (7) (2025) e70284.doi:10.1111/cts.70284. URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC12241707/

  44. [52]

    Djaffardjy, G

    M. Djaffardjy, G. Marchment, C. Sebe, R. Blanchet, K. Belhajjame, A. Gaignard, F. Lemoine, S. Cohen-Boulakia, Developing and reusing bioinformatics data analysis pipelines using scientific workflow systems, Computational and Structural Biotechnology Journal 21 (2023) 2075–2085...

  45. [53]

    Lobentanzer, A

    S. Lobentanzer, A. Koleti, P. Brack, J. Baumbach, J. Saez-Rodriguez, P. Anagnostopoulou, et al., BioCypher: a unifying framework for biomedical research knowledge graphs, Nature Biotechnology 41 (9) (2023) 1234–1237. doi:10.1038/s41587-023-01848-y

  46. [54]

    Sporny, D

    M. Sporny, D. Longley, G. Kellogg, M. Lanthaler, N. Lindstr ¨om, JSON-LD 1.1: A JSON-based serialization for linked data, W3c recommendation, W3C (2020). URLhttps://www.w3.org/TR/json-ld11/

  47. [55]

    Knublauch, D

    H. Knublauch, D. Kontokostas, Shapes Constraint Language (SHACL), W3c recommendation, W3C (2017). URLhttps://www.w3.org/TR/shacl/

  48. [56]

    M. L. Neal, M. K ¨onig, D. Nickerson, J. M´ıˇsek, A. Kiriˇcˇsi, I. I. Moraru, C. J. Myers, J. R. Faeder, et al., Harmonizing semantic annotations for computational models in biology, Briefings in Bioinformatics 20 (2) (2019) 540–550. doi:10.1093/bib/bby087

  49. [57]

    Di Tommaso, M

    P. Di Tommaso, M. Chatzou, E. W. Floden, P. P. Barja, E. Palumbo, C. Notredame, Nextflow enables reproducible computational workflows, Nature Biotechnology 35 (4) (2017) 316–319.doi:10.1038/nbt.3820

  50. [58]

    M ¨older, K

    F. M ¨older, K. P. Jablonski, B. Letcher, M. B. Hall, C. H. Tomkins-Tinch, V . Sochat, J. Forster, S. Lee, S. O. Twardziok, A. Kanitz, et al., Sustainable data analysis with Snakemake, F1000Research 10 (2021) 33.doi: 10.12688/f1000research.29032.2

  51. [59]

    A. E. Ahmed, J. M. Allen, T. Bhat, P. Burra, C. E. Fliege, S. N. Hart, J. R. Heldenbrand, M. E. Hudson, D. D. Istanto, M. T. Kalmbach, et al., Design considerations for workflow management systems use in production genomics research and the clinic, Scientific Reports 11 (1) (2...

  52. [60]

    Forer, S

    L. Forer, S. Sch ¨onherr, Improving the reliability, quality, and maintainability of bioinformatics pipelines with nf- test, GigaSciencePreprint posted May 2024 at bioRxiv, doi:10.1101/2024.05.25.595877 (2025).doi:10.1093/ gigascience/giaf130

  53. [61]

    Patel, C

    Y . Patel, C. Zhu, T. N. Yamaguchi, Y . Z. Bugh, M. Tian, A. Holmes, S. T. Fitz-Gibbon, P. C. Boutros, NFTest: au- tomated testing of Nextflow pipelines, Bioinformatics 40 (2) (2024) btae081.doi:10.1093/bioinformatics/ 22 btae081

  54. [62]

    H ¨ansel, S

    K. H ¨ansel, S. N. Dudgeon, K.-H. Cheung, T. J. S. Durant, W. L. Schulz, From data to wisdom: Biomedical knowledge graphs for real-world data insights, Journal of Medical Systems 47 (1) (2023) 65.doi:10.1007/ s10916-023-01951-2. URLhttps://pmc.ncbi.nlm.nih.gov/articles/PMC10191934/

  55. [63]

    Cheng, Y

    H. Cheng, Y . Wu, S. Khatwani, M. Kruse, D. Dligach, T. A. Miller, M. Afshar, Y . Gao, Scaling biomedical knowl- edge graph retrieval for interpretable reasoning: Applications to clinical diagnosis prediction, medRxivPreprint (2026).doi:10.64898/2026.01.12.26343957. URLhttps:/...

  56. [64]

    P. A. Ewels, A. Peltzer, S. Fillinger, H. Patel, J. Alneberg, A. Wilm, M. U. Garcia, P. Di Tommaso, S. Nahnsen, The nf-core framework for community-curated bioinformatics pipelines, Nature Biotechnology 38 (3) (2020) 276–278. doi:10.1038/s41587-020-0439-x

  57. [65]

    rep., Anthropic, open standard for connecting AI assistants to external data sources and tools

    Anthropic, Model context protocol, Tech. rep., Anthropic, open standard for connecting AI assistants to external data sources and tools. Specification and SDKs available athttps://modelcontextprotocol.io(Nov. 2024). URLhttps://www.anthropic.com/news/model-context-protocol 23

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.