Pith. sign in

REVIEW 3 major objections 6 minor 22 references

UMA-Inverse: Ligand-Conditioned Protein Inverse Folding with a Distogram-Supervised Dense Pair Encoder

T0 review · 3 major / 6 minor · reviewed 2026-07-10 · grok-4.5

Pith's one-line read A compact dense pair encoder for ligand-conditioned protein design trails LigandMPNN on recovery but carries ligand identity far past the pocket.

desk verdict Honest compact dense-pair baseline that trails matched LigandMPNN but cleanly measures distal ligand-signal persistence; missing ablations are real but not fatal to the modest claim. read the letter →

arxiv 2607.07866 v1 pith:2QRUDYSP submitted 2026-07-08 q-bio.BM

classification q-bio.BM
keywords inversefoldingligand-conditioneddesigndensepairrepresentationMixerdistogramsupervisionprotein–ligandinterfacetrianglemultiplication
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper offers UMA-Inverse, a small (~3.3 million parameter) inverse-folding model that designs protein sequences to bind a given ligand. Instead of the usual sparse neighbor graph, it maintains a dense pair tensor over every residue–residue and residue–ligand pair, refined by triangle multiplication and trained with an auxiliary distance-prediction (distogram) loss, then decoded with a position-specific attention over ligand atoms. On the standard small-molecule, metal, and nucleotide test splits it reaches 56.1%, 55.1%, and 35.3% interface recovery—behind a matched re-run of LigandMPNN, though the gap is smaller than published LigandMPNN numbers imply. Redesigns with a fixed pocket still fold and bind under cofolding checks. The distinctive measured property is representational: the dense encoder keeps a clear ligand-conditioning signal at residues tens of ångströms from the pocket, where the local-graph baseline’s signal collapses. The work is offered as a compact baseline and as a concrete characterization of how all-pairs geometry distributes ligand information through a fold.

What carries the argument

The six-block PairMixer dense pair encoder: an all-pairs residue–residue and residue–ligand tensor refined only by triangle multiplication and a transition MLP, kept structure-predictive by an auxiliary distogram head and read out by position-specific ligand attention for the autoregressive decoder.

What would settle it

Train matched on/off ablations that remove the distogram head and replace the learned ligand attention with uniform mean-pool; if interface recovery and the distal-shell KL signal stay essentially unchanged, the claimed causal role of those two additions fails.

Watch

Extended reading notes

Core claim

A distogram-supervised dense PairMixer encoder (triangle multiplication only, no triangle self-attention or sequence track) plus a learned position-specific ligand readout is a viable, compact architecture for ligand-conditioned inverse folding. It achieves 56.1/55.1/35.3% interface recovery on the small-molecule/metal/nucleotide splits and produces foldable, ligand-competent redesigns under cofolding, while measurably propagating ligand identity to distal residues at several-fold to roughly 30× the residual signal of a sparse local graph beyond about 10 Å.

Load-bearing premise

That the two added pieces—the auxiliary distance head and the position-specific ligand attention—are what make the dense encoder work and produce the distal ligand signal, even though the paper never turns either of them off in a controlled test.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. UMA-Inverse is a compact (~3.3M-parameter) ligand- and nucleic-acid-conditioned inverse-folding model that replaces LigandMPNN’s sparse k-NN GNN encoder with a six-block PairMixer (triangle multiplication only, no triangle self-attention or sequence track) over a dense residue–residue and residue–ligand pair tensor, supervised by an auxiliary distogram loss and decoded with a learned position-specific ligand attention readout. On the LigandMPNN test splits the model reports 56.1%/55.1%/35.3% interface recovery (small-molecule/metal/nucleotide), trailing a matched-protocol re-run of LigandMPNN (59.8/64.4/53.3) by less than published LigandMPNN numbers imply. Pocket-fixed redesigns are foldable and ligand-competent under Boltz-2 cofolding, again modestly behind LigandMPNN. The distinctive empirical claim is representational: shell-wise KL analysis shows the dense encoder retains ligand-conditioning signal far beyond ~10 Å, where LigandMPNN’s local graph collapses. The paper positions the model as a compact baseline plus a characterization of dense all-pairs ligand information flow, not as a SOTA design tool.

Significance. If the reported numbers and distal-signal characterization hold, the work is a useful, modest contribution to ligand-conditioned inverse folding: it documents that a stripped PairMixer encoder is viable at ProteinMPNN/LigandMPNN scale, provides a carefully matched re-evaluation of LigandMPNN that shrinks the apparent gap, and quantifies a concrete architectural difference (long-range ligand propagation) that sparse graphs do not exhibit. Strengths that should be credited include the matched-protocol LigandMPNN re-run with paired Wilcoxon tests and bootstrap CIs, Boltz-2 cofolding as an orthogonal structural check, the distance-shell KL diagnostic, public code/weights/demo, and an unusually candid discussion of the nucleotide failure and metal-split contamination. The result is not accuracy-leading and does not claim to be; its value is as a compact baseline and a representational analysis rather than a new design frontier.

major comments (3)
  1. §3.2 and Contributions 1–2 present the auxiliary distogram objective (Eq. 1, λ=0.2) and the position-specific ligand-attention readout (Eq. 2) as the two design choices that distinguish UMA-Inverse from a bare dense-encoder baseline and that target specific failure modes. §5 Limitations explicitly admits there is no on/off ablation of either component, and no recovery or distal-KL comparison against a dense encoder without them. Without those controls, the causal link from the named additions to the reported recovery and to the distal signal is unsupported; the dense-vs-sparse contrast in §4.5 can still stand, but the architectural narrative and contribution list currently over-claim. Either ablations (distogram on/off; mean-pool vs learned ligand attention) or a clear reframing that treats both as untested design rationale is needed before the methods contribution is load-bearing.
  2. §4.3 and §5 attribute the nucleotide failure (35.3%, matching ligand-agnostic ProteinMPNN) to the global ≤50-atom ligand budget under a shared dense token axis, with quantitative support (median 9% of NA heavy atoms retained; median 24% of the true interface). That diagnosis is plausible and well argued, but the paper then proposes a hybrid residue-tensor + per-residue atom-context fix as “the most promising path” without any experiment. Given that nucleotide conditioning is one of the three headline splits and is listed as a design goal (§1, nucleic-acid routing), either a minimal hybrid/local-atom-context experiment or a sharper statement that nucleotide design is out of scope for this architecture would make the central multi-class claim more defensible.
  3. §4.5 reports distal ligand signal “several-fold, up to ~30×” above LigandMPNN beyond ~10 Å, while §4.3–4.4 show LigandMPNN remains more accurate at the interface and on pocket-fixed distal recovery. The paper correctly cautions that distal persistence does not yield a design advantage on native recovery. To keep the representational claim from being over-read, the abstract and conclusion should state more explicitly that the distal KL elevation is a measured encoding property of the dense pair tensor, not evidence of improved allosteric or multi-site design competence—those tasks are left to future work and are not tested here.
minor comments (6)
  1. §3.1: the 384-node crop and M=50 ligand-atom cap are important protocol details; stating how often training/test structures hit these caps (and whether recovery correlates with cropping) would help readers judge selection effects.
  2. Figure 5: log-scale KL with ±1 SEM is informative; adding the absolute mean KL values in a small table (or in the caption) would make the “~30×” claim easier to verify without reading off the plot.
  3. §4.2 teacher-forced recovery (66.1%) and Stage-3 val accuracy (63.2%) are on different metrics/splits than interface recovery; a one-sentence reminder in §4.1 that these are not comparable to LigandMPNN’s published recovery would reduce misreading.
  4. §5 Split integrity: the discussion of adventitious Zn²⁺ in 1F35/1JOB is valuable; if feasible, a brief count of how many metal-split entries are crystallization additives (even a lower bound) would strengthen that caveat.
  5. Typos/clarity: Abstract “UMA-Inverse, which replaces…” is fine; ensure consistent hyphenation of “residue–residue” / “residue-ligand” and of “PairMixer” vs “pair mixer” throughout. Reference list formatting is generally clean.
  6. §3.4 Inference: “autoregressive context is zeroed… making the decoder a one-shot predictor” is important; cross-reference that this matches ProteinMPNN/LigandMPNN practice so readers do not treat it as a limitation unique to UMA-Inverse.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: empirical ML architecture paper with external-benchmark recovery and measured distal KL, not a derivation that reduces to its inputs by construction.

full rationale

UMA-Inverse is a methods paper that proposes a compact PairMixer encoder plus auxiliary distogram and position-specific ligand attention, trains on the public LigandMPNN splits, and reports interface recovery, pocket-fixed redesigns, Boltz-2 cofolds, and distance-shell KL (ligand-on vs. ligand-masked). None of the load-bearing claims is definitional or fitted-then-predicted: recovery is argmax match to native sequence on held-out PDBs under a matched protocol that also re-runs LigandMPNN; the distal-signal claim is an empirical KL contrast between two complete systems (dense all-pairs tensor vs. sparse k-NN graph), not a tautology of the architecture. The auxiliary distogram (Eq. 1, λ=0.2) and ligand-attention readout (Eq. 2) are training/inference design choices whose contribution is not ablated, but that is an experimental gap, not circularity. Citations (ProteinMPNN, LigandMPNN, AlphaFold2/3, PairMixer of Ouyang-Zhang et al., Boltz-2) are external; the sole author has no load-bearing self-citation chain. Training curriculum, noise, and crop sizes are ordinary hyperparameters that do not force the reported numbers. The paper is therefore self-contained against external benchmarks; score 0 with empty steps is the correct outcome.

Assumptions & free parameters 6 free parameters · 4 assumptions · 1 invented entities

Central claims rest on standard inverse-folding and structure-prediction practice plus a handful of hand-chosen training/architecture knobs. No new physical entities. The main domain assumptions are that native-sequence interface recovery and Boltz-2 cofolding are meaningful proxies for design quality, and that triangle multiplication alone carries the needed geometric bias (imported from cited PairMixer work).

free parameters (6)
  • distogram loss weight λ = 0.2
    Set to 0.2 in the combined CE+distogram objective (Eq. 1); not swept in a reported ablation.
  • PairMixer depth and width = 6 blocks, d=128
    Six blocks, d_pair=128, 38 distogram bins; architectural capacity chosen by hand following PairMixer-style designs.
  • global ligand atom budget M = 50
    Cap of 50 ligand heavy atoms nearest protein centroid; load-bearing for nucleotide failure analysis and cubic cost.
  • node crop limit = 384
    Structures cropped at 384 residue+ligand nodes for batching; affects what context the dense tensor sees.
  • coordinate noise and sidechain-as-ligand rate = σ=0.1Å, p=0.03
    Gaussian σ=0.1Å and p=0.03 sidechain injection during training; regularization choices.
  • sampling temperature T = 0.1
    T=0.1 for headline recovery and diversity snapshot; diversity comparison is temperature-sensitive.
assumptions (4)
  • domain assumption Triangle multiplication without triangle self-attention recovers nearly all geometric inductive bias needed for structure-aware pair tensors.
    Imported from Ouyang-Zhang et al. PairMixer work and used to justify the six-block encoder (§2–3).
  • domain assumption Native interface sequence recovery (5Å sidechain–nonprotein cutoff) and Boltz-2 ligand ipTM/RMSD are valid proxies for ligand-binding design quality.
    Evaluation protocol throughout §4; authors note recovery penalizes valid non-native sequences but still lead with it.
  • domain assumption LigandMPNN train/val/test splits at 30% sequence identity are an appropriate shared benchmark despite retrieval and adventitious-metal issues.
    Dataset §3.1 and split-integrity discussion in §5.
  • ad hoc to paper Routing DNA/RNA ATOM records into a shared ligand atom pool is a legitimate way to condition on nucleotide partners.
    Parser extension in §3.1; empirically fails to beat ProteinMPNN on the nucleotide split.
invented entities (1)
  • UMA-Inverse PairMixer + learned ligand-attention decoder
    purpose: Encode all residue–residue and residue–ligand pairs densely and read ligand context with position-specific attention for inverse folding.
    Architectural construct, not a physical entity; independent evidence is empirical recovery/KL/cofold numbers in this paper only.

how reviews work

0 comments
Cite this review

Pith. "Pith review of UMA-Inverse: Ligand-Conditioned Protein Inverse Folding with a Distogram-Supervised Dense Pair Encoder." pith.science (2026). https://pith.science/paper/2QRUDYSP

@misc{pith2026260707866,
  author       = {Pith},
  title        = {Pith review of: UMA-Inverse: Ligand-Conditioned Protein Inverse Folding with a Distogram-Supervised Dense Pair Encoder},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2QRUDYSP}},
  note         = {Machine review of arXiv:2607.07866}
}
abstract

Designing protein sequences that bind specific ligands benefits from an inverse-folding model conditioned on full ligand geometry. We present UMA-Inverse, which replaces the sparse graph backbone of LigandMPNN with a dense pair-representation encoder: a six-block PairMixer (triangle multiplication, no triangle self-attention or sequence track) refines all residue-residue and residue-ligand atom pairs, supervised by an auxiliary distogram objective, and an autoregressive decoder attends over ligand atoms through a learned, position-specific readout of the pair tensor. The model is compact ($\sim$3.3 M parameters). On the LigandMPNN test splits it reaches 56.1%/55.1%/35.3% interface recovery (small-molecule/metal/nucleotide). It trails LigandMPNN, but by less than the published numbers suggest: re-run under our identical protocol, LigandMPNN scores 59.8/64.4/53.3 (vs. published 63.3/77.5/50.5). In a pocket-fixed setting the redesigns are confidently folded and ligand-binding-competent under Boltz-2 cofolding, again modestly behind LigandMPNN. Its distinctive property is representational: the dense encoder propagates ligand identity to residues far beyond the interface, where LigandMPNN's signal decays. We offer UMA-Inverse as a compact baseline for ligand-conditioned inverse folding that trails LigandMPNN in accuracy, together with a characterization of how a dense all-pairs encoder distributes ligand information.

Figures

Figures reproduced from arXiv: 2607.07866 by the authors.

Figure 1
Figure 1. UMA-Inverse architecture. Backbone and ligand/nucleic-acid atoms are featurized into node and pairwise (RBF distance) features; a six-block PairMixer encoder refines the pair tensor Z over all residue–residue and residue–ligand atom pairs. Each block uses triangle multiplication and a pair transition and omits triangle self-attention [6]. An auxiliary distogram head supervises Z during training, and an autoregressiv… view at source ↗
Figure 2
Figure 2. Three-stage training curves. Validation accuracy and loss across the curriculum; the auxiliary distogram top-1 accuracy is overlaid for Stage 3. The best checkpoint (red star) is at epoch 11. to interface residues (sidechain heavy atom within 5 Å of any non-protein heavy atom), per-PDB median over the 10 samples, headline = mean of per-PDB medians. We report all three ligand classes (small molecule, metal, nucleotid… view at source ↗
Figure 3
Figure 3. Interface sequence recovery on the LigandMPNN test splits. [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Constrained (pocket-fixed) design. (a) With the native pocket held fixed, both methods redesign the distal positions; LigandMPNN recovers them somewhat better than UMA-Inverse (N=104). (b) Boltz-2 cofold ligand-pose RMSD for the native sequence, UMA-Inverse, and Ligand…
Figure 5
Figure 5. Figure 5: Ligand-conditioning signal vs. distance, by ligand class. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

22 extracted references · 22 canonical work pages

  1. [1]

    Ragotte, Lukas F

    Justas Dauparas, Ivan Anishchenko, Nathaniel Bennett, Hua Bai, Robert J. Ragotte, Lukas F. Milles, Basile I. M. Wicky, Alexis Courbet, Rob J. de Haas, Neville Bethel, Philip J. Y. Leung, Timothy F. Huddy, Samuel Pellock, Doug Tischer, Frederick Chan, Brian Gross, Vikram Bhatt, Asim Kang, et al. Robust deep learning–based protein sequence design using Prot...

  2. [2]

    Atomic context-conditioned protein sequence design using LigandMPNN.Nature Methods, 22(4):717–723, 2025

    Justas Dauparas, Gyu Rie Lee, Robert Pecoraro, Linna An, Ivan Anishchenko, Cameron Glasscock, and David Baker. Atomic context-conditioned protein sequence design using LigandMPNN.Nature Methods, 22(4):717–723, 2025. doi: 10.1038/s41592-025-02626-1. Preprint: bioRxiv 2023.12.22.573103

  3. [3]

    Kai Yi, Kiarash Jamali, and Sjors H. W. Scheres. All-atom inverse protein folding through discrete flow matching. InInternational Conference on Machine Learning (ICML), 2025. arXiv:2507.14156

  4. [4]

    Highly accu- rate protein structure prediction with AlphaFold.Nature, 596(7873):583–589, 2021

    John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ron- neberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Žídek, et al. Highly accu- rate protein structure prediction with AlphaFold.Nature, 596(7873):583–589, 2021. doi: 10.1038/s41586-021-03819-2

  5. [5]

    Ballard, Joshua Bambrick, et al

    Josh Abramson, Jonas Adler, Jack Dunger, Richard Evans, Tim Green, Alexander Pritzel, Olaf Ronneberger, Lindsay Willmore, Andrew J. Ballard, Joshua Bambrick, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3.Nature, 630:493–500, 2024. doi: 10.1038/s41586-024-07487-w

  6. [6]

    Diaz, Gianluca Scarpellini, Richard Strong Bowen, Nate Gruver, Adam Klivans, Philipp Krähenbühl, Aleksandra Faust, and Maruan Al- Shedivat

    Jeffrey Ouyang-Zhang, Pranav Murugan, Daniel J. Diaz, Gianluca Scarpellini, Richard Strong Bowen, Nate Gruver, Adam Klivans, Philipp Krähenbühl, Aleksandra Faust, and Maruan Al- Shedivat. Triangle Multiplication Is All You Need For Biomolecular Structure Representations. arXiv:2510.18870 [q-bio.QM], 2025. Genesis Research; UT Austin

  7. [7]

    Bowen Jing, Stephan Eismann, Patricia Suriana, Raphael J. L. Townshend, and Ron Dror. Learning from protein structure with geometric vector perceptrons. InInternational Conference on Learning Representations (ICLR), 2021. URL https://openreview.net/forum?id=1YLJ DvSx6J4

  8. [8]

    Learning inverse folding from millions of predicted structures

    Chloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin, Brian Hie, Tom Sercu, Adam Lerer, and Alexander Rives. Learning inverse folding from millions of predicted structures. InInternational Conference on Machine Learning (ICML), pages 8946–8970. PMLR, 2022

Show all 22 references
  1. [9]

    Zhangyang Gao, Cheng Tan, and Stan Z. Li. PiFold: Toward effective and efficient protein inverse folding. InInternational Conference on Learning Representations (ICLR), 2023. URL https://openreview.net/forum?id=oMsN9TYwJ0j

  2. [10]

    Accurate and robust protein sequence design with CarbonDesign.Nature Machine Intelligence, 6:536–547, 2024

    Milong Ren, Chungong Yu, Dongbo Bu, and Haicang Zhang. Accurate and robust protein sequence design with CarbonDesign.Nature Machine Intelligence, 6:536–547, 2024. doi: 10.1038/s42256-024-00838-2

  3. [11]

    Boltz-2: Towards accurate and efficient binding affinity prediction.bioRxiv, 2025

    Saro Passaro, Gabriele Corso, Jeremy Wohlwend, Mateo Reveiz, Stephan Thaler, Vignesh Ram Somnath, Noah Getz, Tally Portnoi, Julien Roy, Hannes Stark, David Kwabi-Addo, Dominique Beaini, Tommi Jaakkola, and Regina Barzilay. Boltz-2: Towards accurate and efficient binding affini...

  4. [12]

    Watson, David Juergens, Nathaniel R

    Joseph L. Watson, David Juergens, Nathaniel R. Bennett, Brian L. Trippe, Jason Yim, Helen E. Eisenach, Woody Ahern, Andrew J. Borst, Robert J. Ragotte, Lukas F. Milles, et al. De novo design of protein structure and function with RFdiffusion.Nature, 620:1089–1100, 2023. doi: 1...

  5. [13]

    J. K. V. Butcher, R. Krishna, R. Mitra, R. I. Brent, Y. Li, N. Corley, P. Kim, J. Funk, S. V. Mathis, S. Salike, et al. De novo design of all-atom biomolecular interactions with RFdiffusion3. bioRxiv, 2025. doi: 10.1101/2025.09.18.676967

  6. [14]

    Shuai, Tianyu Lu, Subhang Bhatti, Petr Kouba, and Po-Ssu Huang

    Richard W. Shuai, Tianyu Lu, Subhang Bhatti, Petr Kouba, and Po-Ssu Huang. Ensemble- conditioned protein sequence design with Caliby.bioRxiv, 2025. doi: 10.1101/2025.09.30.679633

  7. [15]

    Krapp, Fernando A

    Lucien F. Krapp, Fernando A. Meireles, Luciano A. Abriata, and Matteo Dal Peraro. Context- aware geometric deep learning for protein sequence design.Nature Communications, 15:6273,

  8. [16]

    doi: 10.1038/s41467-024-50571-y

  9. [17]

    RNA sequence design and protein–DNA specificity prediction with NA-MPNN.bioRxiv, 2025

    Andrew Kubaney, Andrew Favor, Lilian McHugh, Raktim Mitra, Robert Pecoraro, Justas Dauparas, Cameron Glasscock, and David Baker. RNA sequence design and protein–DNA specificity prediction with NA-MPNN.bioRxiv, 2025. doi: 10.1101/2025.10.03.679414

  10. [18]

    Benchmarking and consensus ranking of inverse folding models for protein–ligand interface design

    Yao Wei, Uliano Guerrini, and Ivano Eberini. Benchmarking and consensus ranking of inverse folding models for protein–ligand interface design. InCompanion Proceedings of the 16th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics (BCB ...

  11. [19]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInternational Conference on Learning Representations (ICLR), 2019. URL https://openreview.net/for um?id=Bkg6RiCqY7

  12. [20]

    Smith, Stuart Firestein, and John F

    Paul C. Smith, Stuart Firestein, and John F. Hunt. The crystal structure of the olfactory marker protein at 2.3 å resolution.Journal of Molecular Biology, 319(3):807–821, 2002. doi: 10.1016/S0022-2836(02)00242-5

  13. [21]

    Vincenzo Laveglia, Andrea Giachetti, Davide Sala, Claudia Andreini, and Antonio Rosato. Learning to identify physiological and adventitious metal-binding sites in the three-dimensional structures of proteins by following the hints of a deep neural network.Journal of Chemical I...

  14. [22]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv:2210.02747 [cs.LG], 2022. 16

Pith tools

Reviewed July 10, 2026 · model on record in the stance chip above.