REVIEW 3 cited by
Otter-Knowledge: benchmarks of multimodal knowledge graph representation learning from different sources for drug discovery
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent research on predicting the binding affinity between drug molecules and proteins use representations learned, through unsupervised learning techniques, from large databases of molecule SMILES and protein sequences. While these representations have significantly enhanced the predictions, they are usually based on a limited set of modalities, and they do not exploit available knowledge about existing relations among molecules and proteins. In this study, we demonstrate that by incorporating knowledge graphs from diverse sources and modalities into the sequences or SMILES representation, we can further enrich the representation and achieve state-of-the-art results for drug-target binding affinity prediction in the established Therapeutic Data Commons (TDC) benchmarks. We release a set of multimodal knowledge graphs, integrating data from seven public data sources, and containing over 30 million triples. Our intention is to foster additional research to explore how multimodal knowledge enhanced protein/molecule embeddings can improve prediction tasks, including prediction of binding affinity. We also release some pretrained models learned from our multimodal knowledge graphs, along with source code for running standard benchmark tasks for prediction of biding affinity.
Forward citations
Cited by 3 Pith papers
-
Does your model understand genes? A benchmark of gene properties for biological and text models
An architecture-agnostic benchmark of gene property prediction shows text and protein models lead on genomic and regulatory tasks, while expression-based models lead on localization.
-
Multimodal Contrastive Representation Learning in Augmented Biomedical Knowledge Graphs
Combining frozen biomedical language model embeddings with graph contrastive learning produces better initial node embeddings for biomedical knowledge graph link prediction.
-
Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT
SPRINT co-embeds drugs and proteins with a structure-aware language model and attention pooling, achieving leading virtual screening enrichment and billion-scale retrieval speed.
Discussion (0). Continue with ORCID to comment.