Pith. sign in

REVIEW 2 cited by

Enhancing the Protein Tertiary Structure Prediction by Multiple Sequence Alignment Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01824 v1 pith:3RH5ODIE submitted 2023-06-02 q-bio.QM cs.CEcs.LGq-bio.BM

classification q-bio.QMcs.CEcs.LGq-bio.BM
keywords proteinsequencesmsaspredictionstructureaccuracyalignmentenhancing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of protein folding research has been greatly advanced by deep learning methods, with AlphaFold2 (AF2) demonstrating exceptional performance and atomic-level precision. As co-evolution is integral to protein structure prediction, AF2's accuracy is significantly influenced by the depth of multiple sequence alignment (MSA), which requires extensive exploration of a large protein database for similar sequences. However, not all protein sequences possess abundant homologous families, and consequently, AF2's performance can degrade on such queries, at times failing to produce meaningful results. To address this, we introduce a novel generative language model, MSA-Augmenter, which leverages protein-specific attention mechanisms and large-scale MSAs to generate useful, novel protein sequences not currently found in databases. These sequences supplement shallow MSAs, enhancing the accuracy of structural property predictions. Our experiments on CASP14 demonstrate that MSA-Augmenter can generate de novo sequences that retain co-evolutionary information from inferior MSAs, thereby improving protein structure prediction quality on top of strong AF2.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Steering Protein Family Design through Profile Bayesian Flow

    q-bio.BM 2025-02 conditional novelty 6.0 of 10

    ProfileBFN adapts Bayesian flow networks to accept protein-family profiles, enabling diverse, novel, and apparently functional family protein generation from single-sequence training.

  2. A Comprehensive Review of Protein Language Models

    q-bio.BM 2025-02 conditional novelty 2.0 of 10

    A survey paper that catalogs protein language models, their architectures, training data, benchmarks, and tools, but lacks a systematic methodology and contains several factual errors.

Pith tools