{"id":"b02c5bc0-8b83-4634-a508-292023690844","arxiv_id":"2412.01388","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A preference-fine-tuned protein language model ranks single and double mutants of CAR binding domains better than the pretrained model, finding several mutants that outperform their parent.","lead":"The authors fine-tuned a protein language model on preference pairs built from high-throughput cell assays, and found that the model's loss correlates with how strongly engineered CAR receptors activate T cells. If it holds up, this offers a way to use AI to guide small numbers of mutations in cell-therapy proteins.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No random-mutant control and evaluation restricted to model-selected mutants leave the central hit-maturation claim unsupported; a matched random-mutant arm would settle it.","rationale":"The reader's verdict of CONDITIONAL is well aligned with the evidence level of this preprint. The strongest claim has two parts: that model loss correlates with biological assay readouts, and that the fine-tuned model reliably discovers improved mutants via few-shot exploration. The first part is supported by Figures 5 and 6, but only on mutants that were selected by the very model being evaluated. The second part is supported by Figure 7, but with no random-mutant baseline. These are distinct gaps from the reader's weakest_assumption about assay noise and missing replicates; the assay-noise concern is real, but the missing random-mutant control is more directly load-bearing for the 'hit maturation' claim because it tests whether the model-guided selection itself adds value. The proposed concrete test would provide that control: a matched random-mutant arm evaluated on the same plates would measure both the background improvement rate and the incremental value of the model ranking. Since the paper already frames itself as a proof-of-concept and includes explicit caveats about statistical significance, the appropriate verdict remains CONDITIONAL rather than ACCEPT or REJECT, and the reader's conditional judgement need not be changed.","tokens_in":10402,"tokens_out":7494,"duration_ms":77231,"concrete_test":"Run the same 96-well activation assay on a matched control plate of random single and double mutants for candidates 1, 2, and S8-S10, drawn from the same parent CDR3s, excluding any training-set sequence, and using the same number of mutants per parent as the model-selected sets (15/15, 45, and 8 respectively). Compare (i) the fraction of random mutants exceeding parent activity against the model-selected fraction, and (ii) the Pearson correlation between model loss and activity computed jointly on selected and random mutants. If random mutants show a comparable improvement fraction, or if the joint correlation drops below significance, the model-guided selection step is not necessary for the claimed hit-maturation benefit; if model-selected mutants clearly dominate, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 evaluates only mutants chosen by the model itself: greedy selection of the top 15 single and double mutants, exhaustive selection of the top 45, and few-shot selection of the top 8. Figures 5 and 6 then regress the same model loss against activation within this non-random, range-restricted set. This is not an unbiased estimate of how well model loss ranks arbitrary mutants, and the reported p-values are conditional on a selected sample rather than a random draw from the design space. Figure 7's 'improved mutants', including cases with more than double parent activity, has no random-mutant comparator; without knowing the improvement rate of random single or double substitutions at matched positions, the hit-maturation result cannot be distinguished from chance. The authors' own caveat that S8-S10 validation differences are not statistically significant (Section 4.1) makes 'reliably find improved mutants' too strong. The decisive missing control is a random-mutant arm on the same assay plates.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes using preference-optimized protein language models for hit maturation of CAR VHH domains. It constructs preference pairs from high-throughput co-culture activation data, fine-tunes ProGen2-medium with KTO loss, and evaluates model loss against GFP activation for selected mutants around parental CARs. It reports Pearson correlations around -0.57 to -0.70 for the fine-tuned model's loss and negligible correlations for the pretrained model, and it identifies several mutants with more than double the parent activity. The authors conclude that fine-tuned model loss is a useful ranking signal for guided exploration of the CAR design space.","tokens_in":10614,"tokens_out":5534,"duration_ms":48792,"significance":"If the central claim holds, this is an interesting proof-of-concept: preference optimization on noisy cellular assay data can produce a sequence-scoring function that correlates with biological activity better than pretrained PLM likelihood, with potential generalization to other therapeutic proteins. Strengths include the direct comparison to the pretrained model, evaluation on disjoint mutant sequences, a thoughtful analysis of loss-function pathologies (KTO vs. hinge vs. sigmoid), and explicit caveats that validation-set differences are not statistically significant. However, the current evidence is insufficient to establish that the method 'reliably finds improved mutants': evaluation is restricted to model-selected mutants, no random-mutant control is reported, and no replicate plates or error bars are provided. These are fixable with additional experiments, but they are load-bearing for the paper's main claims.","major_comments":[{"comment":"The correlation analysis is performed only on mutants selected by the model itself (top 15 greedy, top 45 exhaustive, top 8 few-shot). This range-restricted, non-random sample cannot provide an unbiased estimate of how well model loss ranks arbitrary single/double mutants, and the reported p-values are conditional on the selection rule. A matched random-mutant arm evaluated on the same plates is required to establish that the correlation is not an artifact of selection; please add such an arm and report correlations on the combined random plus selected set.","section":"Section 4, Figures 5-6"},{"comment":"The hit-maturation claim that the method 'reliably find[s] improved mutants' is not supported without a baseline improvement rate for random single/double substitutions. All evaluated mutants were chosen for high model likelihood, so observing several mutants above the parent could reflect the chance distribution of substitutions rather than the model's guidance. The authors' own caveat that validation differences for S8-S10 are not statistically significant (Section 4.1) further weakens the reliability claim. A random-mutant control on the same plates, with replicate wells, is needed to distinguish model-guided improvement from chance.","section":"Section 4, Figure 7"},{"comment":"The scalar activation score (ΔGFP) used to build preference labels and to validate mutants comes from a high-throughput assay described in Section 3.2 as potentially lacking 'sufficient precision to resolve the finest difference between candidate CARs.' No replicate plates, error bars, or assay-variability estimates are reported. If this score is noisy or plate-dependent, both training labels and validation readouts are affected, and the reported p-values would not be meaningful. Please report replicate measurements for the validation plates and quantify assay noise.","section":"Section 3.1-3.2"},{"comment":"Preference pairs are constructed from the same high-throughput assay family used for evaluation, so the model may learn to reproduce the assay's systematic ranking rather than the biological activity of the CAR. Removing training-set mutants and using validation parents addresses overlap in sequences, but it does not remove a plate-level or protocol-level bias shared by all experiments. Please either provide an independent readout (e.g., binding or orthogonal activation assay) or explicitly bound the effect of assay bias on the conclusions.","section":"Section 3.3 vs. Section 4"}],"minor_comments":[{"comment":"Report exact Pearson r values, p-values, and sample sizes for each panel; 'p < 0.05' is insufficient, especially with small selected samples.","section":"Figures 5-6"},{"comment":"The notation 'arg maxtopk(logits(concat(L, R))p)' is unclear; specify that the top-k amino-acid substitutions are taken per position p and define what 'valid(s)' means.","section":"Algorithm 1"},{"comment":"The normalization of ΔGFP on the second plate using reference CAR performance is described only verbally; provide the formula and state whether each plate had replicate reference wells.","section":"Section 4.1"},{"comment":"Minor typo: 'start of with' should be 'start off with'.","section":"Section 3.1"},{"comment":"Provide a data/code availability statement or clearly state which parts of the pipeline are proprietary, since the absence of public data limits reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a proof-of-concept with proprietary data and moderate novelty. The main technical concerns—missing random-mutant control, absent replicate assays, and the shared assay family for training and validation—are fixable but require new experiments rather than simple edits. I would encourage the editor to request a major revision that adds a random-mutant arm, replicate measurements, and exact statistics, rather than to reject, because the core correlation result is interesting and the authors are appropriately cautious in some of their claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid proof-of-concept that preference-fine-tuned ProGen2 loss correlates with a biological activation readout, but the paper's stronger claim—reliable hit maturation—is not yet supported because there is no random-mutant control arm.\n\nWhat's actually new: applying DPO/KTO to CAR VHH CDR3 hit maturation using a proprietary high-throughput platform. The core comparison is well designed. Fine-tuned model loss correlates with ΔGFP activation (r ≈ -0.574 overall, -0.699 and -0.418 per candidate, p < 0.05), while the pretrained model's loss does not (0.218 and 0.072). They also remove training-set mutants from evaluation and use disjoint context/candidate sequences for the validation preference pairs. That is real evidence that fine-tuning adds signal beyond the base model.\n\nThe soft spots are mostly experimental. The biggest issue is that all evaluated mutants are selected by the model itself, so the loss–activity regression is range-restricted and conditional on a non-random sample. Without a matched set of random single and double mutants on the same plates, the \"improved mutants\" in Figure 7, including the >2× cases, cannot be distinguished from chance. The authors' own caveat that S8–S10 validation differences are not statistically significant (Section 4.1) makes \"reliably find improved mutants\" too strong. Also: single-plate assays with no replicates or error bars, proprietary data and no code, and a small candidate count. The preference pairs are built from the same high-throughput assay family used for validation, so there is a mild circularity, though the disjoint sequence evaluation mitigates it.\n\nNone of this is fatal to the central correlation result, which stands as a proof-of-concept. The paper is honest about its preliminary nature and the KTO-vs-DPO discussion is thoughtful. It would be a good reading-group case study on selection bias in ML-guided screening.\n\nWho it's for: ML-for-protein engineers and immunotherapy ML people. It deserves a serious referee, but the review should make a random-mutant arm and replicate plates the condition for acceptance. I'd send it to review rather than desk-reject.","headline":"A credible proof-of-concept that preference-fine-tuned PLM loss tracks a CAR activation assay, but the hit-maturation claim needs a random-mutant control before it can be trusted.","tokens_in":11198,"tokens_out":1721,"would_cite":false,"duration_ms":17540,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that after preference fine-tuning of an auto-regressive protein language model, the model's next-token loss is highly correlated with T-cell activation measurements for CAR mutants, allowing few-shot hit maturation that…","keywords":["protein language model","preference optimisation","direct preference optimization","KTO","CAR-T","VHH","hit maturation","few-shot exploration"],"falsifier":"Re-assay the same CAR mutants from Figures 5-7 on at least three replicate plates with plate-normalised controls and compute the correlation between average model loss and $\\Delta$GFP; if the Pearson r no longer reaches significance (p $\\geq$ 0.05, or |r| below roughly 0.3) or the 'more than double parent' mutants fall within the parent's assay noise, the central claim is falsified. A second check: re-derive the preference labels from an independent repeat of the high-throughput co-culture run; if the chosen/rejected CDR3 pairs flip on replication, the training signal itself is not reproducible.","tokens_in":10209,"feed_emoji":"🧬","tokens_out":9244,"duration_ms":69075,"temperature":0.7,"pith_summary":"The paper reports that fine-tuning an auto-regressive protein language model with preference optimisation on high-throughput CAR activation data makes the model's next-token loss a reliable ranking score for CAR variant activity. Across single and double mutants of two parent CARs, the average model loss correlates with measured T-cell activation at Pearson r = -0.574 (p < 0.05), with the correlation improving after fine-tuning compared to the pretrained model. Using this score to select mutants yielded variants that outperform their parent in most of ten cases, and more than double the parent's activity in three. The authors present this as a proof-of-concept for ML-guided hit maturation in cell therapy, with the preference signal coming from low-precision but high-throughput FACS/NGS measurements.","feed_headline":"Protein LM loss predicts CAR activity in cell assays","feed_subtitle":"A preference-fine-tuned protein language model ranks CAR mutants by T-cell activation, enabling few-shot hit maturation.","key_machinery":"The load-bearing mechanism is the combination of (i) a preference-fine-tuning scheme that conditions an auto-regressive PLM on a context of five good-performing CDR3 sequences and optimises the difference in log-likelihood between chosen (good) and rejected (poor) CDR3 completions, and (ii) the use of the fine-tuned model's next-token cross-entropy as a ranking score for unseen mutants. Among the three losses tested (sigmoid DPO, hinge, and KTO), the authors select KTO because it raises the likelihood of chosen completions without over-penalising rejected ones, avoiding the collapsed or trivial outputs seen with hinge loss. The model loss is averaged over all 120 context permutations to produce a per-mutant score; this score is what correlates with the biological activation readout $\\Delta$GFP.","core_discovery":"The central claim is that an auto-regressive protein language model (ProGen2, 764M parameters) fine-tuned with Kahneman-Tversky Optimization (KTO) on preference pairs built from high-throughput CAR-T activation data learns a sequence-level reward that is well approximated by its own cross-entropy loss. Concretely, for mutants generated greedily or by exhaustive search around a parent VHH CDR3, the model loss averaged over five-CDR3 context prompts correlates strongly with the change in GFP reporter activation (all mutants: r = -0.574; candidate 1: r = -0.699; candidate 2: r = -0.418), and the correlation is significant at p < 0.05. The fine-tuned model also finds mutants that exceed parent activation, including cases with more than double the parent's activity, in a search space of $10^{4}$-$10^{5}$ variants. The authors interpret this as evidence that preference-fine-tuned PLMs can guide few-shot hit maturation despite the noisiness of high-throughput biological data.","pith_inferences":["An untested extension is that the five-CDR3 context softly encodes the target antigen, so prompting with good performers for a new target might enable zero-shot maturation without retraining; the paper says this is not its focus but the architecture would permit it.","The preference-label construction, which groups by CDR3 and retains the maximum-performing variant, may itself encode a strong prior that the model learns to exploit; comparing against a model trained on regression labels from the same scalar would isolate what preference optimisation contributes.","The reported correlations and 'double parent activity' results come from single plates without replicate wells or error bars, so the practical claim depends on the assay scalar being reproducible across plates and days, which the paper does not demonstrate.","If the high-throughput scalar is too noisy to resolve fine differences, the preference pairs may be mislabelled; a useful diagnostic would be to measure label agreement between two independent runs of the same CAR library."],"forward_implications":["If the correlation holds in independent replicates, model loss can act as an in-silico ranking filter, letting labs evaluate hundreds of mutants computationally before synthesising a handful for 96-well assays.","Exhaustive scoring of all single and double mutants around a parent CDR3, despite its compute cost, finds high-performing mutants more reliably than greedy left-to-right generation, and is the recommended mode for hit maturation.","Fine-tuning with preferences derived from high-throughput cell assays can transfer to other therapeutic protein modalities such as bispecific antibodies or cytokines, since the approach only needs preference pairs and a pretrained auto-regressive PLM.","The next-token loss of the fine-tuned model is a usable proxy for activation, meaning the model can serve as a zero-shot or few-shot fitness predictor without a separate reward model or structural information."],"supporting_citations":[{"why":"Supplies the pretrained ProGen2 autoregressive protein language model (151M-6.4B) that is fine-tuned throughout the paper; its 764M 'medium' variant carries the reported results.","marker":"[Nijkamp et al., 2023]"},{"why":"Formulates Direct Preference Optimisation, the sigmoid-based preference objective that defines the DPO baseline and the reference point for the KTO and hinge variants used here.","marker":"[Rafailov et al., 2024]"},{"why":"Proposes Kahneman-Tversky Optimization (KTO), the loss function selected for the model used in the final evaluation; the paper relies on KTO's asymmetric handling of chosen and rejected completions.","marker":"[Ethayarajh et al., 2024]"},{"why":"Justifies the preference-optimisation framing for biological sequences; the authors cite it as related work for fine-tuning models on preference-style tasks and as a future direction for sampling.","marker":"[Hayes et al., 2024]"},{"why":"Provides the hinge loss variant tested in the loss-function comparison; the paper uses it to show that margin-based alternatives produce trivial completions, motivating the KTO choice.","marker":"[Liu et al., 2024]"}],"fun_headline_variants":["Protein LM loss predicts CAR-T activation in assays","Few-shot CAR hits via preference-tuned protein LM","Model loss mirrors CAR activity for hit maturation","KTO-tuned protein LM ranks CAR mutants by activity","Preference-optimized protein LM guides CAR hit discovery"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The single scalar derived from the co-culture FACS/NGS fractions is treated as a stable ground-truth measure of CAR activation; if that measurement is noisy or target-line dependent in ways not captured by the reported single-plate assays, the preference labels and the measured correlations may not reproduce.","fun_headline_variants_meta":{"raw":{"variants":["Protein LM loss predicts CAR-T activation in assays","Few-shot CAR hits via preference-tuned protein LM","Model loss mirrors CAR activity for hit maturation","KTO-tuned protein LM ranks CAR mutants by activity","Preference-optimized protein LM guides CAR hit discovery"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000255,"raw_usage":{"total_tokens":1547,"prompt_tokens":896,"completion_tokens":651,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":576}},"tokens_in":512,"tokens_out":651,"duration_ms":6074,"temperature":1.0,"reasoning_tokens":576,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:24:41.786198+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-assay the same CAR mutants from Figures 5-7 on at least three replicate plates with plate-normalised controls and compute the correlation between average model loss and $\\Delta$GFP; if the Pearson r no longer reaches significance (p $\\geq$ 0.05, or |r| below roughly 0.3) or the 'more than double parent' mutants fall within the parent's assay noise, the central claim is falsified. A second check: re-derive the preference labels from an independent repeat of the high-throughput co-culture run; if the chosen/rejected CDR3 pairs flip on replication, the training signal itself is not reproducible.","supporting_citations":[],"review_version":1}