Pith. sign in

REVIEW 2 cited by

Multi-modal Transfer Learning between Biological Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14150 v1 pith:I4AH4Y6J submitted 2024-06-20 cs.LG

classification cs.LG
keywords biologicalexpressionmodalitiesmodelmodelsmulti-modalmultiplesequence
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Biological sequences encode fundamental instructions for the building blocks of life, in the form of DNA, RNA, and proteins. Modeling these sequences is key to understand disease mechanisms and is an active research area in computational biology. Recently, Large Language Models have shown great promise in solving certain biological tasks but current approaches are limited to a single sequence modality (DNA, RNA, or protein). Key problems in genomics intrinsically involve multiple modalities, but it remains unclear how to adapt general-purpose sequence models to those cases. In this work we propose a multi-modal model that connects DNA, RNA, and proteins by leveraging information from different pre-trained modality-specific encoders. We demonstrate its capabilities by applying it to the largely unsolved problem of predicting how multiple RNA transcript isoforms originate from the same gene (i.e. same DNA sequence) and map to different transcription expression levels across various human tissues. We show that our model, dubbed IsoFormer, is able to accurately predict differential transcript expression, outperforming existing methods and leveraging the use of multiple modalities. Our framework also achieves efficient transfer knowledge from the encoders pre-training as well as in between modalities. We open-source our model, paving the way for new multi-modal gene expression approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 6 citations worldwide. Full citation record

  1. DEFEND: A Large-scale 1M Dataset and Foundation Model for Tobacco Addiction Prevention

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A new 1M-image tobacco product dataset and a multimodal model that combines contrastive, coherence, and description losses, with reported gains over prior baselines.

  2. Artificial Intelligence for Central Dogma-Centric Multi-Omics: Challenges and Breakthroughs

    q-bio.GN 2024-12 conditional novelty 1.0 of 10

    A literature review that maps AI and deep learning methods for central-dogma-centric multi-omics integration and disease modeling.

Pith tools