Pith. sign in

REVIEW 8 cited by

Computational Protein Science in the Era of Large Language Models (LLMs)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.10282 v2 pith:DQZLB4BP submitted 2025-01-17 cs.CE cs.CLq-bio.BM

classification cs.CEcs.CLq-bio.BM
keywords proteincomputationalscienceplmsdesignknowledgelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Considering the significance of proteins, computational protein science has always been a critical scientific field, dedicated to revealing knowledge and developing applications within the protein sequence-structure-function paradigm. In the last few decades, Artificial Intelligence (AI) has made significant impacts in computational protein science, leading to notable successes in specific protein modeling tasks. However, those previous AI models still meet limitations, such as the difficulty in comprehending the semantics of protein sequences, and the inability to generalize across a wide range of protein modeling tasks. Recently, LLMs have emerged as a milestone in AI due to their unprecedented language processing & generalization capability. They can promote comprehensive progress in fields rather than solving individual tasks. As a result, researchers have actively introduced LLM techniques in computational protein science, developing protein Language Models (pLMs) that skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems. While witnessing prosperous developments, it's necessary to present a systematic overview of computational protein science empowered by LLM techniques. First, we summarize existing pLMs into categories based on their mastered protein knowledge, i.e., underlying sequence patterns, explicit structural and functional information, and external scientific languages. Second, we introduce the utilization and adaptation of pLMs, highlighting their remarkable achievements in promoting protein structure prediction, protein function prediction, and protein design studies. Then, we describe the practical application of pLMs in antibody design, enzyme design, and drug discovery. Finally, we specifically discuss the promising future directions in this fast-growing field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQA

    cs.CV 2025-08 unverdicted novelty 7.0 of 10

    mKG-RAG constructs multimodal KGs via MLLM-driven extraction and vision-text matching then applies dual-stage query-aware retrieval to achieve new state-of-the-art results on knowledge-based VQA.

  2. Geometric Flow Matching for Molecular Conformation Generation via Manifold Decomposition

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    GO-Flow applies manifold decomposition to flow matching for molecular conformations by separating translation, SO(3) rotation, and conformation spaces.

  3. Unlocking Biological Workflows for Robust Protein-Text Question Answering: A Dual-Dimensional RAG Framework

    cs.IR 2026-05 unverdicted novelty 6.0 of 10

    2D-ProteinRAG is a dual-dimensional RAG framework that incorporates BLAST workflows plus horizontal attribute alignment and vertical homology denoising to improve protein-text QA on both in-distribution and out-of-dis...

  4. HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous Structure Tokens

    cs.CE 2025-12 conditional novelty 6.0 of 10

    HD-Prot shows that a protein language model can jointly generate sequences and structures using continuous structure tokens instead of quantized tokens, reaching competitive performance on four protein design tasks.

  5. Enhancing Protein-Protein Interaction Prediction with Hierarchical Motif-based Multimodal Protein Embedding

    q-bio.QM 2026-05 unverdicted novelty 5.0 of 10

    MMM-PPI is a new hierarchical multimodal protein encoder that predicts protein-protein interactions by integrating features across residue, motif, and protein scales.

  6. PriHA: A RAG-Enhanced LLM Framework for Primary Healthcare Assistant in Hong Kong

    cs.IR 2026-04 unverdicted novelty 5.0 of 10

    PriHA is a tri-stage RAG framework with query optimization and dual retrieval that outperforms baselines on accuracy and clarity for Hong Kong primary healthcare queries.

  7. PriHA: A RAG-Enhanced LLM Framework for Primary Healthcare Assistant in Hong Kong

    cs.IR 2026-04 unverdicted novelty 5.0 of 10

    PriHA combines query optimization with a Dual Retrieval Augmented Generation pipeline to improve accuracy and clarity of LLM responses on fragmented Hong Kong primary care guidelines.

  8. STELLA: A Multimodal LLM for Protein Functional Annotation via Unified Sequence-Structure Encoding

    q-bio.BM 2025-06 unverdicted novelty 5.0 of 10

    STELLA aligns ESM3 bimodal sequence-structure encodings with Llama-3.1-8B text modeling to claim state-of-the-art results on protein functional description prediction and enzyme-catalyzed reaction prediction.

Pith tools