{"id":"8a7eab8c-0ff8-42fc-854a-006764b70e83","arxiv_id":"2606.31126","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"TabPFN3 and TabICL paired with ESMC or ECFP/RDKit features are competitive or better than task-specific baselines on ProteinGym, PpEST, and several molecular benchmarks, especially in few-shot regimes.","lead":"Tabular in-context models pretrained on synthetic causal tables work surprisingly well for protein fitness and small-molecule property prediction when paired with strong fixed embeddings. This offers a practical few-shot path for wet-lab-limited protein engineering and ADMET-style screening without task-specific retraining.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The protein SOTA claim rests on local random-split aggregates and incomplete external leaderboard confirmation, so the central “achieve or exceed” result is not yet fully secured.","rationale":"The reader correctly flags incomplete external ProteinGym evaluation, PCA rescues, and partial MoleculeNet coverage as reasons for CONDITIONAL rather than ACCEPT. That diagnosis is right on the completeness axis. The single most load-bearing soft spot for the central claim, however, is narrower and more decisive than the general “representation-geometry” assumption: the protein SOTA language in the abstract and strongest claim is currently supported only by local random-split aggregates that the paper itself declines to treat as a finished public ranking. Representation quality is necessary for the method to work, but the claim that the tabular models “achieve or exceed SOTA” stands or falls on whether those local numbers survive official ProteinGym evaluation. The molecular half of the paper is already carefully hedged (“no single pairing dominates”), so the protein ranking is the piece that most directly underwrites the headline. Pending that external check, the verdict remains CONDITIONAL; a clean leaderboard confirmation would move it toward ACCEPT, while a material drop would require softening the abstract. This is a verification gap, not an internal inconsistency, and does not require inventing a deeper theoretical flaw.","tokens_in":15199,"tokens_out":708,"duration_ms":7185,"concrete_test":"Submit the exact TabPFN3 and TabICL packages (including the two PCA-rescue assays) to the official ProteinGym supervised-substitution leaderboard under random, modulo, and contiguous schemes; if either model falls below the current top public entries (Kermut/ProteinNPT) on the random scheme or loses the modest edge over the other on all three schemes, revise the abstract and Table 3 from “achieve or exceed SOTA” to “competitive under local evaluation.”","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim states that TabPFN3/TabICL + ESMC “achieve or exceed state-of-the-art results on ProteinGym.” Table 1 and the abstract report local random 5-fold means (Spearman 0.767 / 0.753) that sit above the public-context numbers in Table 3 (Kermut 0.745, ProteinNPT 0.741). The paper itself, however, labels Table 3 as “external context rather than as a new public ranking claim, because the TabICL and TabPFN3 submission packages must still be evaluated through ProteinGym’s external process” (§4.1). Two of the largest assays further use PCA-rescue rows rather than the same full-feature configuration as the rest of the suite. Official modulo and contiguous holdouts already drop mean Spearman to ~0.59 and ~0.52 (Table 2), so the headline ranking is sensitive both to split protocol and to unfinished external verification. If the external evaluation does not reproduce the local ordering, the protein half of the central claim weakens even while the broader “competitive learner–representation pair” framing remains intact.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"This paper empirically tests whether tabular in-context learners (TabPFN3, TabICL), pretrained on synthetic causal tables, transfer to biomolecular property prediction when used as frozen predictors over fixed domain representations. For proteins, sequences are encoded with ESMC and evaluated on ProteinGym (217 DMS assays; random/modulo/contiguous holdouts; full-train and few-shot) and the PpEST esterase dataset. For small molecules, TabICL/TabPFN3/XGBoost are paired with ECFP, RDKit, or ECFP+RDKit descriptors and compared to ChemProp and ChemProp+CheMeleon across TDC ADMET, MoleculeNet, FS-Mol, and DrugOOD (including ID/OOD gaps). The authors report strong protein results under fixed ESMC, competitive but non-dominant molecular results that depend heavily on representation, and conclude that tabular foundation models are viable biomolecular predictors only as learner–representation pairs.","tokens_in":15566,"tokens_out":1539,"duration_ms":23215,"significance":"The work addresses a practical bottleneck—data-efficient prediction once strong frozen encoders exist—and documents a counter-intuitive transfer of synthetic tabular ICL priors to protein and molecular feature tables. Strengths include multi-assay coverage, official ProteinGym split schemes, few-shot learning curves with explicit support-set protocols, paired TabPFN3–TabICL deltas, fixed-representation controls that isolate the predictor, DrugOOD OOD analysis, and an honest limitations section. If the protein results hold under external ProteinGym evaluation and careful claim language, the paper is a useful empirical audit for protein engineering and cheminformatics practice, even without a new architecture. The learner–representation-pair framing is methodologically clean and should influence how tabular foundation models are reported in scientific domains.","major_comments":[{"comment":"Abstract and opening claim that TabPFN3/TabICL + ESMC “achieve or exceed state-of-the-art results on ProteinGym,” but §4.1 and Table 3 explicitly treat public leaderboard numbers as “external context rather than as a new public ranking claim,” because submission packages still require ProteinGym’s external process. Local random 5-fold means (Table 1: Spearman 0.767/0.753) sit above the contextual public figures (Kermut 0.745, ProteinNPT 0.741), yet this ordering is not yet externally secured. The abstract’s SOTA language should be tempered to match the body’s caution until external evaluation is complete, or the external results should be included.","section":"Abstract; §4.1; Table 3"},{"comment":"Table 2 shows large drops under official modulo and contiguous schemes (TabPFN3 mean Spearman 0.744 random → 0.595 modulo → 0.522 contiguous). The paper correctly notes that random-split numbers are not the only generalization estimate, but the abstract and high-level protein claims still lead with the strongest random-split aggregates. Load-bearing claims should foreground split dependence (or report all three schemes as co-primary) so readers do not take random-split SOTA-style numbers as the main ProteinGym result.","section":"Abstract; Table 2; Figure 3; §7"},{"comment":"Two of the largest ProteinGym assays (HIS7_YEAST_Pokusaeva_2019, SPG1_STRSG_Olson_2014) use PCA128 rescue (and 8 estimators for TabPFN3) rather than the same full-feature configuration as the remaining suite (§4.1; Limitations). Because these assays can move aggregate means, the paper should quantify sensitivity: report aggregates with and without the rescued assays, and/or with a uniform PCA setting across all methods, so the headline ProteinGym ranking is not partly an artifact of incomplete full-feature coverage.","section":"§4.1; Limitations"},{"comment":"For molecules, the abstract’s “competitive with the existing task-specific state-of-the-art” is only weakly anchored: Table 7 reports best internal pairs by family, MoleculeNet TabPFN3 coverage excludes PCBA (667/806), and there is little direct numerical comparison to published benchmark SOTA numbers (as opposed to the paper’s own ChemProp/XGBoost baselines). Either add a compact external-SOTA reference table per family or soften the abstract to “competitive with strong descriptor and graph baselines under our protocol,” which the body already supports.","section":"Abstract; Table 7; §6; Limitations"}],"minor_comments":[{"comment":"The abstract says models “outperform task-specific supervised regressors on a diverse esterase catalytic activity dataset,” while Table 5 shows TabICL/TabPFN3 trading wins with each other and beating FT ESM/HGBR overall—fine, but “outperform” could specify rank vs MSE and that FT ESM is the strongest non-tabular baseline.","section":"Abstract; Table 5"},{"comment":"Equations (1)–(3) define monotone best-so-far envelopes for plots; this is reasonable for legibility, but figure captions should state more prominently that hollow markers are raw means and that source tables retain full variability, so readers do not over-read envelope smoothness.","section":"§4 Evaluation metrics; Figures 4, 6–8"},{"comment":"RBF sampler collapse on ESMC (Table 1 Spearman ≈ 0) is attributed to median-heuristic scaling; a one-sentence note on whether alternative kernel scales were tried would prevent the impression that kernel methods were dismissed after a single default.","section":"Table 1; §4.1"},{"comment":"PpEST is introduced via Ahmed et al. 2026 with overlapping authorship; a brief disclosure that this is a related experimental resource (not an independent third-party benchmark) would improve transparency without diminishing its value as a diverse-sequence test.","section":"§4 Datasets; References"},{"comment":"Minor consistency: the abstract uses “TabPFN” generically while the body standardizes on TabPFN3; align naming in the abstract and keywords.","section":"Abstract; §2"},{"comment":"Figure 5 (assay size vs Spearman) is under-interpreted; a short quantitative statement (e.g., correlation of performance with assay size) would make the diagnostic more useful.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The empirical audit is careful and the body is more restrained than the abstract. The main risk for the journal is overstated ProteinGym SOTA language before external leaderboard confirmation, not a flawed experimental design. If the authors temper claims and add the requested sensitivity/split framing, this is a solid contribution; I would not reject on novelty grounds alone—the transfer result and protocol quality are the value. PpEST author overlap is minor if disclosed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is simple: once you freeze a strong encoder (ESMC for proteins; ECFP/RDKit for molecules), TabPFN3 and TabICL act as strong few-shot tabular heads. That is counter-intuitive given their synthetic causal-table pretraining, and the paper shows it with a clean predictor–representation protocol rather than a new architecture.\n\nWhat is new is the measurement, not the models. ProteinGym (217 assays, random/modulo/contiguous, few-shot 8–64 with 30 draws), PpEST, and four molecular suites under fixed features give a careful cross-domain audit. Protein side is the stronger half: under fixed ESMC, TabPFN3/TabICL beat ridge, HistGB, fine-tuned ESM, and RBF on local random 5-fold and few-shot curves; PpEST is consistent. Molecular side is honest—no single pair wins TDC/MoleculeNet/FS-Mol/DrugOOD; representation and OOD gaps matter as much as the learner. Framing methods as pairs, not pure architectures, is the right methodological move.\n\nSoft spots are real but bounded. The abstract’s “achieve or exceed SOTA on ProteinGym” sits on local random aggregates (Spearman ~0.77/0.75) above the public-context numbers they print for Kermut/ProteinNPT. The paper itself labels Table 3 as external context only and notes that submission packages still need ProteinGym’s process; two large assays use PCA rescue. Official modulo/contiguous splits drop Spearman to ~0.59/0.52. So the ranking claim is not fully secured yet; the broader “competitive learner–representation pair” claim is. MoleculeNet full-train TabPFN3 also skips PCBA for resource reasons. Citations and baselines look appropriate; no circularity load.\n\nThis is for people who care about data-efficient heads on frozen bio embeddings—protein engineering, ADMET screening, few-shot molecular work. Not a theory paper. It deserves a serious referee: multi-benchmark, multi-split, cautious molecular conclusions, and the incompleteness is already disclosed. I would engage, cite the protein few-shot numbers once external ProteinGym confirms, and treat the SOTA wording as provisional until then.","headline":"Solid empirical audit: TabPFN3/TabICL + ESMC are competitive few-shot protein heads; the ProteinGym SOTA wording is ahead of the external verification the paper itself flags.","tokens_in":16163,"tokens_out":560,"would_cite":true,"duration_ms":5684,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Tabular in-context models trained on synthetic tables become strong biomolecular predictors when paired with expressive fixed embeddings.","keywords":["tabular foundation models","in-context learning","protein fitness prediction","small-molecule property prediction","few-shot learning","ProteinGym","ESMC embeddings","ECFP RDKit descriptors"],"falsifier":"On the same ProteinGym assays and support sizes, replace ESMC with a deliberately weak or random fixed embedding of equal dimension and check whether TabPFN3/TabICL still beat ridge and HistGradientBoosting; if they collapse to baseline or worse, the claim that the tabular prior transfers once geometry is good fails.","tokens_in":16121,"feed_emoji":"🧬","tokens_out":609,"duration_ms":5939,"temperature":0.7,"pith_summary":"Predicting protein fitness or small-molecule properties from few labels is a core bottleneck once good pretrained embeddings already exist. This paper shows that tabular foundation models—pretrained only on synthetic tables from random causal graphs—can fill that role surprisingly well. When sequences are encoded with fixed ESMC embeddings, the models match or exceed strong baselines on ProteinGym and a diverse esterase family, including in few-shot regimes. With ECFP/RDKit descriptors they stay competitive across ADMET, MoleculeNet, FS-Mol, and DrugOOD without dominating every family. The practical claim is that the bottleneck has shifted: choose a representation that already exposes local task structure, then let an amortized tabular in-context learner read support labels at inference time.","feed_headline":"Synthetic-table models predict proteins and molecules well","feed_subtitle":"With fixed ESMC or fingerprint features they match task-specific baselines, even from few labels","key_machinery":"The predictor–representation pair: a frozen domain encoder (ESMC for proteins; ECFP, RDKit, or both for molecules) produces fixed-length rows of a table; TabICL or TabPFN3 then performs in-context prediction by conditioning on labeled support rows without task-specific gradient updates.","core_discovery":"Tabular foundation models (TabPFN3 and TabICL), despite a prior with no obvious biological correspondence, act as strong data-efficient predictors for biomolecular property tasks when treated as predictor–representation pairs: ESMC embeddings plus either model achieve competitive or better ProteinGym and PpEST protein-fitness results, and ECFP/RDKit pairings stay competitive with task-specific methods on small-molecule classification, especially under few-shot budgets.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Tabular ICL models with ESMC match ProteinGym fitness SOTA","Synthetic-table priors transfer to few-shot protein and molecule tasks","TabPFN and TabICL competitive on biomolecular properties via fixed embeddings","ESMC plus tabular models beat supervised regressors on esterase activity","Tabular foundation models act as strong few-shot biomolecular predictors"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The fixed pretrained embeddings already place similar-labeled molecules near each other so a tabular prior trained only on synthetic causal tables can interpolate without learning new structure or using graphs.","fun_headline_variants_meta":{"raw":{"variants":["Tabular ICL models with ESMC match ProteinGym fitness SOTA","Synthetic-table priors transfer to few-shot protein and molecule tasks","TabPFN and TabICL competitive on biomolecular properties via fixed embeddings","ESMC plus tabular models beat supervised regressors on esterase activity","Tabular foundation models act as strong few-shot biomolecular predictors"]},"model":"grok-4.5","effort":"low","cost_usd":0.007778,"raw_usage":{"total_tokens":1921,"prompt_tokens":841,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":77780000,"prompt_tokens_details":{"text_tokens":841,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1003,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":841,"tokens_out":77,"duration_ms":8661,"temperature":1.0,"reasoning_tokens":1003,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T10:21:48.538938+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same ProteinGym assays and support sizes, replace ESMC with a deliberately weak or random fixed embedding of equal dimension and check whether TabPFN3/TabICL still beat ridge and HistGradientBoosting; if they collapse to baseline or worse, the claim that the tabular prior transfers once geometry is good fails.","supporting_citations":[],"review_version":2}