REVIEW 2 cited by
Enhancing the Protein Tertiary Structure Prediction by Multiple Sequence Alignment Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The field of protein folding research has been greatly advanced by deep learning methods, with AlphaFold2 (AF2) demonstrating exceptional performance and atomic-level precision. As co-evolution is integral to protein structure prediction, AF2's accuracy is significantly influenced by the depth of multiple sequence alignment (MSA), which requires extensive exploration of a large protein database for similar sequences. However, not all protein sequences possess abundant homologous families, and consequently, AF2's performance can degrade on such queries, at times failing to produce meaningful results. To address this, we introduce a novel generative language model, MSA-Augmenter, which leverages protein-specific attention mechanisms and large-scale MSAs to generate useful, novel protein sequences not currently found in databases. These sequences supplement shallow MSAs, enhancing the accuracy of structural property predictions. Our experiments on CASP14 demonstrate that MSA-Augmenter can generate de novo sequences that retain co-evolutionary information from inferior MSAs, thereby improving protein structure prediction quality on top of strong AF2.
Forward citations
Cited by 2 Pith papers
-
Steering Protein Family Design through Profile Bayesian Flow
ProfileBFN adapts Bayesian flow networks to accept protein-family profiles, enabling diverse, novel, and apparently functional family protein generation from single-sequence training.
-
A Comprehensive Review of Protein Language Models
A survey paper that catalogs protein language models, their architectures, training data, benchmarks, and tools, but lacks a systematic methodology and contains several factual errors.
Discussion (0). Continue with ORCID to comment.