REVIEW 3 major objections 5 minor 1 cited by
State-aware protein-ligand complex prediction using AlphaFold3 with purified sequences
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read By steering AlphaFold3 with an MSA purified for one functional state, this paper produces accurate poses for allosteric inhibitors that default AlphaFold3 misplaces by 14.9–19.1 Å, cutting ligand RMSD to 1.7–2.5 Å.
desk verdict The EGFR half of this paper is a genuine semi-blind result; the IL-1β half is in-sample selection, and the conclusion overstates what the method does without a reference state. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is AF-ClaSeq's iterative sequence purification. Starting from a large MSA, the pipeline splits sequences into small groups, predicts each group with AlphaFold2, computes RMSD of the conformation-defining regions (αC helix and A-loop two-turn helices for EGFR, residues 756–769 and 857–863; β4-5 and β7-8 loops for IL-1β, residues 46–55 and 86–96) against reference structures for the desired state, keeps sequences from groups closest to that state, and repeats. After several enrichment rounds, M-fold sampling with statistical voting identifies the few sequences that most often produce that state; those purified sequences are compiled into a custom MSA and fed into AlphaFold3 as the only evolutionary input, together with the ligand SMILES and without any template.
What would settle it
Repeat the EGFR purification using the active-state reference 2ITP instead of the inactive-state 2GS7, then run AlphaFold3 on the same allosteric ligands; if the poses still land below 2.5 Å, the reference is not doing the claimed work, and if they fail, the state-specific reference is the load-bearing element.
Extended reading notes
Core claim
Default AlphaFold3 produces the wrong protein conformation and ligand pose for the allosteric EGFR inhibitors 8A2A, 8A2B, and 8A2D because the input MSA is dominated by sequences encoding the active kinase state. The authors show that when the MSA is iteratively enriched for sequences whose AlphaFold2 predictions approach the inactive-state reference 2GS7, then passed to AlphaFold3 together with the ligand SMILES and no template, the model consistently produces the inactive pocket and places the ligand with RMSD 2.5 Å, 1.7 Å, and 1.9 Å, respectively. The same protocol, targeting the β4-5 and β7-8 loop conformation of IL-1β seen in the antagonist-bound structure 8C3U, turns a complete failure into near-perfect convergence: using the top 20 purified sequences, all 40 AlphaFold3 runs match the experimental complex. The paper presents this as evidence that AlphaFold3's poor performance on rare allosteric states is not a model limitation requiring retraining, but an input-MSA problem that can be corrected by state-aware sequence selection.
Load-bearing premise
The purification pipeline needs a correct reference structure for the desired conformational state, because it chooses sequences by their RMSD to that reference; if the reference is wrong, or is the very target structure being predicted, the improvement is guided rather than blind.
Editorial extensions
If this is right
- For the two systems tested, sequence purification lifts AlphaFold3 from total ligand-pose failure (14.9–19.1 Å) to sub-2.5 Å accuracy on all three EGFR allosteric complexes.
- The protocol transfers to an unrelated system: IL-1β with a cryptic-pocket antagonist, where top-20 purified sequences give 40 of 40 predictions converged to the experimental structure.
- State-bias strength matters: the most strongly biased purified sequences produce the best AlphaFold3 results, while weaker normal-voting subsets perform worse, suggesting AlphaFold3 benefits from a sharpened state signal.
- No template of the specific ligand complex is required; only a generalized conformational reference for the target state, so the pipeline can in principle cover entire classes of allosteric modulators once such a reference exists.
Reading between the lines
- A direct generalizability test would apply the same two-reference purification to a GPCR with known inactive and active structures and an allosteric ligand; pose improvement of the same magnitude would show the mechanism is not specific to kinases or cytokines.
- The method is a state-selection tool rather than a blind conformational search: it can only produce the state it is pointed at, so its discovery value depends on having a reference structure for the biologically relevant state.
- One could also ask whether the purified sequences work as a fine-tuning prior for AlphaFold3-like models, which would remove the need to supply reference-derived RMSD metrics at inference time.
- A practical screen would dock a ligand library into pockets generated from different purified states and compare state-dependent enrichment, testing whether the pose accuracy translates into virtual-screening signal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a state-aware protein-ligand complex prediction strategy that combines AlphaFold3 (AF3) with sequence subsets purified by the authors' previously described AF-ClaSeq method. The central idea is that AF-ClaSeq isolates evolutionarily encoded conformational sub-signals from a heterogeneous MSA, and using such purified sequences as the MSA input to AF3 biases the model toward a desired functional state, such as the inactive conformation of EGFR or the cryptic-pocket-bound loop arrangement of IL-1β. The authors demonstrate the approach on two test systems: three allosteric EGFR inhibitors (PDB 8A2A, 8A2B, 8A2D) and the IL-1β allosteric antagonist complex (PDB 8C3U). They report that default AF3 predictions fail for these targets, while AF3 using purified sequences achieves ligand RMSDs near or below 2.5 Å and high ligand pLDDT scores with strong structural convergence. The paper concludes that the method is a generalizable, evolutionary-information-driven framework that does not require knowledge of individual inhibitor-bound structures.
Significance. The idea of using sequence purification to inject conformational state information into co-folding models is timely and potentially useful for drug-design applications involving allosteric and cryptic-pocket sites, where default AF3 is known to struggle. The manuscript is clearly written, presents a transparent comparison of default and purified-sequence predictions, and reports multiple random seeds per condition, which is commendable. However, the significance of the central claim is not established by the evidence as presented. For IL-1β, the sequence enrichment is guided directly by RMSD to the experimentally determined bound structure 8C3U, making the subsequent AF3 prediction an in-sample consistency check rather than an independent prediction. For EGFR, the method relies on a pre-chosen inactive-state reference (2GS7) and a posteriori knowledge that the inactive state is required. The evaluation also lacks statistical error bars and sensitivity analyses for the many free parameters.
major comments (3)
- [Results, IL-1β subsection (Figure 5C, Figure 6)] The IL-1β demonstration is circular and therefore does not support a generalizable state-aware prediction claim. The text states that the enrichment used 'RMSD relative to the β4-5 loop (residues 46-55) and β7-8 loop (residues 86-96)' and Figure 5C plots these RMSDs 'relative to the experimentally determined allosteric conformation' (PDB 8C3U). Sequences were then selected because they produced structures closest to 8C3U, and AF3 was run with those MSAs to predict 8C3U. The resulting sub-0.5 Å RMSD values are a consistency check of the selection loop, not an independent prediction. The Conclusion statement that the method 'does not require prior knowledge of individual inhibitor-bound structures (such as PDB codes 8A2A, 8A2B, 8A2D, or 8C3U)' is contradicted by this procedure for the IL-1β case.
- [Results, EGFR subsection (Figure 2A, Figure 3)] The EGFR case is only semi-blind: the desired inactive state is selected using a known reference (2GS7) after the authors observed that default AF3 predicts an active-state-like conformation. The paper does not provide a protocol for discovering the relevant conformational state without such a reference, so the claim of a broadly applicable 'state-aware' strategy overreaches. Furthermore, the residue ranges used for the RMSD metric (αC helix and activation-loop two-turn helices, residues 756–769 and 857–863) and all other thresholds (coverage 0.6, group sizes 28 and 6, lowest-15% selection, four iterations, 0.15 voting threshold) are introduced without justification or sensitivity analysis, leaving the reader uncertain whether the improvement depends on these specific arbitrary choices.
- [Results and Figure legends (Figures 2–6)] The paper reports no error bars, confidence intervals, or statistical tests for the RMSD and pLDDT improvements. Although ten seeds with five structures per seed are used, the variance across seeds is not quantified, and no comparison is made against a proper null model (e.g., randomly selected sequences matched for MSA depth) for the final AF3 predictions. Without such analysis, the reported 'astonishingly good' results cannot be distinguished from selection-induced overfitting or from stochastic variation in a small number of cases.
minor comments (5)
- [Abstract / Introduction] The abstract states that the authors 'introduced a state-aware protein-ligand prediction strategy,' but the manuscript contains no methods section or pseudocode, and the AF-ClaSeq procedure is only cited to reference 14; the paper would benefit from a self-contained description of the purification algorithm and its parameters.
- [Figure 1 and Results] The definition of ligand RMSD is not stated: it should be clarified whether the RMSD is computed after a global protein alignment, a pocket-only alignment, or a ligand-only superposition, and whether symmetry-corrected RMSD is used for the small molecule.
- [Results, EGFR subsection] The phrase 'we did not obtain many sequences for the active state but acquired very few sequences for the inactive state' is confusing; the intended contrast between the two voting modes should be clarified.
- [Figure 2D and 2E] The color-coding of panels D and E is not fully explained in the figure legend; the reader cannot determine which points correspond to which bin without referring to the main text.
- [Conclusion] The claim that default AF3 predictions fail because the model 'memorizes' training-set entries is not directly tested; the authors do not check whether the exact targets or close analogues appear in AF3's training data.
Circularity Check
IL-1β demonstration is in-sample: MSA sequences are selected by RMSD to the target structure 8C3U, so the 'corrected' AF3 prediction is a consistency check; EGFR is semi-blind but also hardwires a known inactive reference.
-
self definitional
[Results / 'Sequence purification of ligand bound interleukin-1β (IL-1β) loop conformation corrects IL-1β/ligand complex prediction'; Figure 5C legend]
"Starting with 4,129 sequences obtained from DeepMSA2, we performed a similar iterative enrichment approach as used for EGFR, where RMSD relative to the β4-5 loop (residues 46-55) and β7-8 loop (residues 86-96) was used as the metric to enrich sequences that produced structures with low RMSD values. ... a focused range of PDB structures with RMSD relative to the β4-5 loop lower than 2.5 Å and β7-8 loop lower than 3.0 Å was selected. ... Figure 5C ... plotted according to RMSD with respect to the β4-5 loop and β7-8 loop regions relative to the experimentally determined allosteric conformation."
The IL-1β purification metric is the RMSD to the experimental bound structure (8C3U) that later serves as ground truth. Sequences are kept because AF2 predicts their loops close to 8C3U, so the top-10/top-20 MSA inputs are constructed to encode the 8C3U conformation. Running AF3 with those MSAs and reporting sub-0.5 Å RMSD is a consistency check on the selection filter, not an independent prediction of the IL-1β/ligand complex. The conclusion's claim that the method 'does not require prior knowledge of individual inhibitor-bound structures' is contradicted for this case, because the allosterically bound 8C3U structure is the very reference used for sequence ranking.
-
fitted input called prediction
[Results / 'Sequence Purification of Epidermal Growth Factor Receptor (EGFR) Inactive State Improves Ligand Pose Prediction for Allosteric Inhibitors']
"we sought to purify sequences that bias toward the inactive conformation. ... At each iteration, we selected the sequences with lowest 15% of αC helix and activation loop two-turn helix regions RMSD with respect to the inactive state (PDB structure 2GS7) for the next iteration."
The EGFR sequence subset is selected by low RMSD to the known inactive reference 2GS7. The test ligands (8A2A, 8A2B, 8A2D) are allosteric inhibitors known to stabilize the inactive state, so selecting sequences for the inactive state selects for the expected protein conformation of the answer. The subsequent ligand-pose improvement is largely inherited from this enforced pocket state; it is a self-state upper-bound test rather than a blind correction. This step is less directly circular than the IL-1β case because 2GS7 is not the exact target structure, but the desired conformational state is still supplied as input and then reported as a 'corrected prediction'.
full rationale
The paper's central claim is that AF-ClaSeq-purified MSA subsets correct AlphaFold3's failed ligand-pose predictions by encoding the relevant functional state. The IL-1β demonstration is circular by construction: the purification and ranking metric is RMSD to the experimentally determined allosteric conformation (8C3U), and the top sequences are chosen precisely because they reproduce that conformation; the resulting near-perfect AF3 output is therefore in-sample. The EGFR case is semi-blind because the reference (2GS7) is a generalized inactive state rather than the target complex, but the allosteric inhibitors tested are selected because they bind that known inactive state, so the state constraint is still part of the input. These two cases do not establish a generalizable blind prediction strategy; they show that MSA subsets can be tuned to return a supplied conformation. The self-citation to AF-ClaSeq (ref 14) is not itself load-bearing beyond describing the selection procedure, and the paper uses no machine-checked or external-benchmark validation that would make the IL-1β result independent. Overall, the IL-1β 'prediction' reduces to its own input, and the EGFR result is substantially conditioned on the same prior state knowledge, warranting a score of 8 rather than a lower partial-circularity score.
Assumptions & free parameters
free parameters (10)
- EGFR RMSD residue ranges (756-769, 857-863)
- IL-1β RMSD residue ranges (46-55, 86-96)
- MSA coverage threshold =
0.6
- Group size for initial distribution analysis =
28
- Group size for iterative enrichment =
6
- Number of enrichment iterations =
4
- Selection cutoff for enrichment =
lowest 15% RMSD
- Voting threshold =
0.15
- Top-K sequence selection for IL-1β =
10 and 20
- Number of AF3 seeds / structures per seed =
10 seeds, 5 structures = 50
assumptions (3)
- domain assumption AF2 and AF3 process MSA information similarly, so purified sequences validated in AF2 transfer to AF3.
- domain assumption The chosen reference structures (2ITP, 2GS7, 8C3U) correctly represent the relevant conformational states.
- domain assumption RMSD over the selected loop/helix regions captures the functionally relevant conformational difference.
Cite this review
Pith. "Pith review of State-aware protein-ligand complex prediction using AlphaFold3 with purified sequences." pith.science (2026). https://pith.science/paper/SJJBSCLN
@misc{pith2026250600147,
author = {Pith},
title = {Pith review of: State-aware protein-ligand complex prediction using AlphaFold3 with purified sequences},
year = {2026},
howpublished = {\url{https://pith.science/paper/SJJBSCLN}},
note = {Machine review of arXiv:2506.00147}
}
read the original abstract
Deep learning-based prediction of protein-ligand complexes has advanced significantly with the development of architectures such as AlphaFold3, Boltz-1, Chai-1, Protenix, and NeuralPlexer. Multiple sequence alignment (MSA) has been a key input, providing coevolutionary information critical for structural inference. However, recent benchmarks reveal a major limitation: these models often memorize ligand poses from training data and perform poorly on novel chemotypes or dynamic binding events involving substantial conformational changes in binding pockets. To overcome this, we introduced a state-aware protein-ligand prediction strategy leveraging purified sequence subsets generated by AF-ClaSeq - a method previously developed by our group. AF-ClaSeq isolates coevolutionary signals and selects sequences that preferentially encode distinct structural states as predicted by AlphaFold2. By applying MSA-derived conformational restraints, we observed significant improvements in predicting ligand poses. In cases where AlphaFold3 previously failed-producing incorrect ligand placements and associated protein conformations-we were able to correct the predictions by using sequence subsets corresponding to the relevant functional state, such as the inactive form of an enzyme bound to a negative allosteric modulator. We believe this approach represents a powerful and generalizable strategy for improving protein-ligand complex predictions, with potential applications across a broad range of molecular modeling tasks.
Figures
Forward citations
Cited by 1 Pith paper
-
OTMol: Robust Molecular Structure Comparison via Optimal Transport
OTMol formulates molecular superimposition as fused supervised Gromov-Wasserstein optimal transport, yielding atom-order-independent, chirality-preserving alignments with preserved bond connectivity across ligands, pe...
Reference graph
Works this paper leans on
-
[1]
Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 630, 493- 500 (2024)
work page 2024
-
[2]
Hong, Y . et al. Accurate prediction of protein –ligand interactions by combining physical energy functions and graph-neural networks. Journal of Cheminformatics 16, 121 (2024)
work page 2024
- [3]
-
[4]
Evans, R. et al. Protein complex prediction with AlphaFold-Multimer. biorxiv, 2021.2010. 2004.463034 (2021)
arXiv 2021
-
[5]
Baek, M. et al. Accurate prediction of protein structures and interactions using a three -track neural network. Science 373, 871-876 (2021)
work page 2021
-
[6]
Buttenschoen, M., Morris, G.M. & Deane, C.M. PoseBusters: AI -based docking methods fail to generate physically valid poses or generalise to novel sequences. Chemical Science 15, 3130-3139 (2024)
work page 2024
- [7]
-
[8]
Wohlwend, J. et al. B<span class="sc">oltz</span>-1 Democratizing Biomolecular Interaction Modeling. bioRxiv, 2024.2011.2019.624167 (2025)
arXiv 2025
Show all 28 references
-
[9]
Discovery, C. et al. Chai-1: Decoding the molecular interactions of life. bioRxiv, 2024.2010.2010.615955 (2024)
2024
-
[10]
& Schwede, T
Škrinjar, P., Eberhardt, J., Durairaj, J. & Schwede, T. Have protein -ligand co-folding methods moved beyond memorisation? bioRxiv, 2025.2002.2003.636309 (2025)
2025
-
[11]
Zheng, H. et al. AlphaFold3 in Drug Discovery: A Comprehensive Assessment of Capabilities, Limitations, and Applications. bioRxiv, 2025.2004.2007.647682 (2025)
2025
-
[12]
Research Square (2025)
Eva Nittinger, Ö.Y ., Alessandro Tibo, Gustav Olanders, Christian Tyrchan Co-folding, the Future of Docking – Prediction of Allosteric and Orthosteric Ligands. Research Square (2025)
2025
-
[13]
& Anandkumar, A
Qiao, Z., Nie, W., Vahdat, A., Miller, T.F. & Anandkumar, A. State -specific protein –ligand complex structure prediction with a multiscale deep generative model. Nature Machine Intelligence 6, 195-208 (2024)
2024
-
[14]
& Cheng, X
Xing, E., Zhang, J., Wang, S. & Cheng, X. Leveraging Sequence Purification for Accurate Prediction of Multiple Conformational States with AlphaFold2. arXiv preprint arXiv:2503.00165 (2025)
2025 arXiv
-
[15]
& Misra, A
Yewale, C., Baradia, D., Vhora, I., Patil, S. & Misra, A. Epidermal growth factor receptor targeting in cancer: a review of trends and strategies. Biomaterials 34, 8690-8707 (2013)
2013
-
[16]
Johnston, J.B. et al. Targeting the EGFR pathway for cancer therapy. Current medicinal chemistry 13, 3483-3492 (2006)
2006
-
[17]
Russo, A. et al. A decade of EGFR inhibition in EGFR -mutated non small cell lung cancer (NSCLC): Old successes and future perspectives. Oncotarget 6, 26814 (2015)
2015
-
[18]
Zhang, Y .-L. et al. The prevalence of EGFR mutation in patients with non -small cell lung cancer: a systematic review and meta-analysis. Oncotarget 7, 78985 (2016)
2016
-
[19]
Solca, F. et al. Target binding properties and cellular activity of afatinib (BIBW 2992), an irreversible ErbB family blocker. The Journal of pharmacology and experimental therapeutics 343, 342-350 (2012)
2012
-
[20]
Kashima, K. et al. CH7233163 overcomes osimertinib- resistant EGFR -Del19/T790M/C797S mutation. Molecular Cancer Therapeutics 19, 2288-2297 (2020)
2020
-
[21]
Wood, E.R. et al. A unique structure for epidermal growth factor receptor bound to GW572016 (Lapatinib) relationships among protein conformation, inhibitor off-rate, and receptor activity in tumor cells. Cancer research 64, 6652-6659 (2004)
2004
-
[22]
Zhao, P., Yao, M.-Y ., Zhu, S.-J., Chen, J. -Y . & Yun, C.-H. Crystal structure of EGFR T790M/C797S/V948R in complex with EAI045. Biochemical and biophysical research communications 502, 332-337 (2018)
2018
-
[23]
Beyett, T.S. et al. Molecular basis for cooperative binding and synergy of ATP-site and allosteric EGFR inhibitors. Nature communications 13, 2530 (2022)
2022
-
[24]
Obst-Sander, U. et al. Discovery of novel allosteric EGFR L858R inhibitors for the treatment of non- small-cell lung cancer as a single agent or in combination with osimertinib. Journal of medicinal chemistry 65, 13052-13073 (2022)
2022
-
[25]
& Shaw, D.E
Shan, Y ., Arkhipov, A., Kim, E.T., Pan, A.C. & Shaw, D.E. Transitions to catalytically inactive conformations in EGFR kinase. Proceedings of the National Academy of Sciences 110, 7270-7275 (2013)
2013
-
[26]
& Ghiringhelli, F
Rébé, C. & Ghiringhelli, F. Interleukin-1β and cancer. Cancers 12, 1791 (2020)
2020
-
[27]
& Torres, R
Ren, K. & Torres, R. Role of interleukin- 1β during pain and inflammation. Brain research reviews 60, 57 -64 (2009)
2009
-
[28]
Hommel, U. et al. Discovery of a selective and biologically active low -molecular weight antagonist of human interleukin-1β. Nature Communications 14, 5497 (2023). PDB Code: 8A2A Ligand RMSD (pred vs true): 19.1 Å A-loop/αC helix RMSD: 7.2 Å AF3 Default prediction Ground truth...
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.