{"id":"7ad47a18-325e-47f7-872f-e5a0fbe2cd76","arxiv_id":"2507.09982","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A joint omics-plus-text conditioning signal can generate chemically valid, novel hit-like molecules, demonstrated on the new TextOmics benchmark.","lead":"This preprint introduces TextOmics, a dataset that links gene expression profiles and molecular text descriptions to the molecules that induce them, and ToDi, a diffusion model that generates hit-like molecules from both signals. The authors report that ToDi beats existing omics-only and text-only baselines across validity, diversity, structural, and zero-shot disease benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The text condition in TextOmics is generated from the target molecule's own SELFIES, so the reported multimodal advantage may be a structure-reconstruction artifact rather than evidence of independent semantic guidance.","rationale":"The reader identified the same weakest assumption: text is generated from the target molecule, so it may not be an independent signal. My analysis confirms this is the most load-bearing issue and that the claimed multimodal advantage is confounded. A conditional verdict is appropriate because the concern is testable: an independent-text control would either preserve or collapse the reported gains. Additional concerns (self-referential hit-ratio threshold in Algorithm E.1, SELFIES-guaranteed validity, lambda selected on the test set) reinforce but do not replace this central issue. Since the reader already conditioned acceptance on resolving this, the verdict remains CONDITIONAL/UNCHANGED.","tokens_in":20513,"tokens_out":5709,"duration_ms":68396,"concrete_test":"Re-run the Table 2 / D.6 comparison while replacing the BioT5-generated text descriptions with independently authored descriptions that were not derived from the target SELFIES (e.g., ChEBI-20 annotations for the same compounds, or BioT5 text generated from held-out SMILES with structure-bearing tokens masked). If the full-ToDi margin over ToDiw/o T in FCD, Morgan, MACCS, or Novelty disappears or reverses, the multimodal advantage is an artifact of text re-encoding the target structure; if the margin persists, the semantic-guidance claim survives.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 and Appendix A.3 state that every molecular textual description is produced by BioT5 from the SELFIES representation of the target molecule, then manually verified. This makes each text prompt a lossy paraphrase of the very structure the model is asked to generate. The central claim that joint omics+text conditioning beats either modality alone (Table 2, Table D.6) therefore rests on a confound: ToDi's cross-attention can read structural tokens such as functional-group names directly out of the text, so the improvement over ToDiw/o T is expected even if the model has learned no biological or semantic association. The zero-shot section (4.6) does not resolve this, because it replaces molecular descriptions with out-of-domain patient narratives and evaluates only two Alzheimer's examples. The manuscript's own limitation section concedes there is no experimental validation and only a case study. Consequently, the headline \"outperforms SOTA\" conflates representation/architecture gains with genuine multimodal grounding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TextOmics, a benchmark that links omics expression profiles, molecular textual descriptions, and SELFIES molecular representations in a one-to-one manner, and proposes ToDi, a diffusion-based generative framework that conditions on both omics embeddings (OmicsEn) and text embeddings (TextEn) to generate hit-like molecules. Experiments on ChemInduced, TargetPerturb, and DiseaseSign report validity, uniqueness, novelty, Levenshtein similarity, FCD, Morgan/MACCS Tanimoto, and a hit-ratio metric, with the paper claiming state-of-the-art performance and strong zero-shot potential.","tokens_in":20666,"tokens_out":4241,"duration_ms":52248,"significance":"If the claims are substantiated, TextOmics and ToDi would be a useful contribution to multimodal molecular generation: the dataset is a new resource, the architecture is modular and clearly described, the code is promised to be open, and ablations (Table D.6) isolate the contribution of each modality. However, the central ``outperforms SOTA'' claim is weakened by the benchmark construction: the textual descriptions are generated from the target molecule's own SELFIES via BioT5 (Section 3.1, Appendix A.3), so the text condition is a re-encoding of the structure rather than an independent semantic source. In addition, the validity metric is predetermined by the SELFIES representation, and the hit-ratio metric in Algorithm E.1 uses a self-referential threshold. These issues do not necessarily invalidate the framework, but they must be corrected before the headline claims are credible.","major_comments":[{"comment":"Every molecular textual description in TextOmics is produced by BioT5 from the SELFIES representation of the target molecule and then manually verified. Consequently, the text condition is a lossy paraphrase of the very structure the model is asked to generate. The improvement of full ToDi over ToDiw/o T in Tables 2 and D.6 can therefore be explained by the cross-attention mechanism reading functional-group tokens directly from the text, without the model learning any independent biological or semantic association. Please address this confound, for example by evaluating on independently authored descriptions (e.g., ChEBI-20 natural-language captions) or by a control experiment in which the text condition is perturbed while the target structure remains fixed, and show that the multimodal gain is not purely a structure-reconstruction artifact.","section":"Section 3.1 and Appendix A.3"},{"comment":"The hit-ratio metric defines the threshold δ as the average noise estimation error across the model's own generated samples, and counts a molecule as a hit if its noise error is below this self-derived threshold OR if a functional-group match succeeds. This makes the hit ratio self-referential: a model whose errors are tightly concentrated can inflate its hit ratio without any semantic alignment. Please define δ independently of the model's output distribution, or better, compute the hit ratio purely from functional-group matching, and report the functional-group match and noise-error criteria separately.","section":"Section 4.2, Eq. (E.3), and Algorithm E.1"},{"comment":"The reported 100% validity scores are a property of the SELFIES representation, not a learned achievement, because every SELFIES string decodes to a valid molecular graph (as the paper itself states in Appendix A.1). Presenting this value as a distinguishing model result in Table 2 and in the surrounding discussion overstates the contribution. The validity column should either be removed from the comparison or replaced by a more informative check, such as validity after conversion to SMILES and RDKit sanitization, with a note that SELFIES guarantees syntactic validity by construction.","section":"Table 2 and Section 4.3"},{"comment":"After transferring ToDi from ChemInduced to TargetPerturb, the model achieves a MACCS Tanimoto score of 1.00 on six of ten target proteins. Since the model was not trained on TargetPerturb molecules, perfect MACCS agreement with the reference ligands is surprising and needs explanation. Please clarify how the reference ligands for each target were selected, whether there is any overlap with the ChemInduced training set, and report how many of the generated molecules are exact or near-duplicates of the reference molecules. Without this, the cross-target generalization claim in Table 3 is not fully supported.","section":"Section 4.5, Table 3"},{"comment":"The zero-shot therapeutic-generation claim rests on similarity scores for two Alzheimer's disease examples (Memantine and Donepezil), and the limitation section explicitly states that the generated molecules have not undergone experimental validation and that the evaluation is confined to a case study. The abstract's phrase \"remarkable potential in zero-shot therapeutic molecular generation\" is disproportionate to two case-study similarity numbers. Please either expand the zero-shot evaluation to a larger set of disease queries with statistical testing, or temper the claim in the abstract and conclusion.","section":"Section 4.6 and Section 5"}],"minor_comments":[{"comment":"In the paragraph on MolT5, \"hider\" should be \"hinder\".","section":"Section 2"},{"comment":"The DiffGen paragraph contains \"molecular stings\" and should read \"molecular strings\".","section":"Section 3.2"},{"comment":"The phrase \"atomic-level chemcial environments\" contains a typo: \"chemcial\" should be \"chemical\".","section":"Section 4.2"},{"comment":"The sentence fragment \"donoted as w ⊂ C\" should be \"denoted as w ⊂ C\".","section":"Section 3.2"},{"comment":"The abstract and Section 4.3 state that ToDi \"outperforms existing SOTA approaches\" without noting that the comparison is modality-conditional; Table D.6 shows that the text-only variant ToDiw/o O has lower novelty than GxV AEs on the ChemInduced benchmark, so the headline claim should be qualified to the multimodal setting.","section":"Abstract and Section 4.3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript presents a useful dataset and a clean modular architecture, but the benchmark construction and the hit-ratio metric need substantial rework before the state-of-the-art claim is credible. I see a viable path to revision: use independent text-molecule pairs, define the hit-ratio threshold a priori, and clarify the TargetPerturb evaluation protocol. I would not reject the paper outright, but the current version does not support the abstract's claims as stated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about arXiv:2507.09982 is that it's a useful new dataset wrapped in an overstated SOTA claim. The TextOmics benchmark is genuinely new: one-to-one pairing of omics expression, BioT5-generated molecular text, and SELFIES for 13,755 MCF7 compounds, plus ten target perturbation profiles. That's a real resource for the field. The ToDi framework is a sensible combination of a VAE omics encoder, frozen SciBERT, and diffusion over SELFIES with cross-attention. It's not architecturally groundbreaking, but it's clean and the ablations suggest both modalities add something.\n\nThe soft spots are real and they matter. The text descriptions are produced by BioT5 from the very SELFIES string the model is asked to generate. So the 'semantic' condition is a lossy paraphrase of the target structure. When ToDi outperforms its omics-only variant, the gain can come from reading functional-group names directly off the text, not from learning biological semantics. The hit-ratio metric in Algorithm E.1 sets its threshold to the average noise error of the model's own samples, which makes the 'hit' criterion self-referential. Validity is 100% because SELFIES guarantees validity by construction; that's a property of the encoding, not the model. There are no error bars anywhere, lambda is selected on the test set (Table C.5), and the zero-shot section is two Alzheimer's molecules with patient narratives. The manuscript itself concedes no experimental validation.\n\nThat said, the central direction is plausible. The dataset is the kind of thing the subfield needs, and the framework is a reasonable first attempt at joint conditioning. The problems are in the evaluation and the claims, not in the basic idea. If the authors released the dataset and code, re-ran the metrics with independent thresholds and held-out lambda, and reframed the text condition as a paraphrase rather than independent knowledge, this could be a solid contribution.\n\nFor a colleague: this is worth a serious referee because the resource is real and the question matters. But I wouldn't cite the SOTA numbers, and I'd treat the benchmark with caution until the artifacts are out and the metrics are cleaned up. If you're working on molecular generation, it's worth a look for the dataset design; the method itself you can probably improve on.","headline":"A genuinely useful heterogeneous dataset and a plausible joint-conditioning framework, but the SOTA claim rests on text prompts that paraphrase the target structure and a self-referential hit-ratio threshold.","tokens_in":21215,"tokens_out":2137,"would_cite":false,"duration_ms":22424,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that jointly conditioning diffusion-based molecular generation on gene expression profiles and textual descriptions produces hit-like molecules that beat single-modality baselines in validity, novelty, diversity, and…","keywords":["hit-like molecular generation","conditional diffusion","omics-guided generation","text-guided generation","heterogeneous data integration","SELFIES","drug discovery","zero-shot therapeutic generation"],"falsifier":"Regenerate the benchmark's textual descriptions using descriptions authored independently of the target molecule, or sourced from curated chemistry texts, then rerun the ToDi versus text-only and omics-only ablations on the same test split; if the joint gains in novelty and structural similarity shrink or reverse, the claimed multimodal advantage is an artifact of text re-encoding the target structure rather than a genuine integration of complementary signals.","tokens_in":20278,"feed_emoji":"🧬","tokens_out":5401,"duration_ms":57067,"temperature":0.7,"pith_summary":"This paper introduces TextOmics, a dataset that pairs omics expression profiles with molecular textual descriptions and SELFIES structures, and ToDi, a diffusion framework that conditions generation on both omics and text. The authors aim to show that jointly integrating biological and semantic signals yields chemically valid, novel, diverse, hit-like molecules—plausible starting points for drug discovery—better than conditioning on either modality alone. If true, this would let researchers generate candidate molecules directly from a cell line's or patient's transcriptional state plus a short chemical description, speeding up early-stage target-specific drug discovery.","feed_headline":"Omics + text beats single-modality molecule generation","feed_subtitle":"ToDi's diffusion model jointly conditions on gene expression and text, hitting 100% validity and higher novelty than prior baselines.","key_machinery":"The load-bearing mechanism is a joint conditioning pipeline built on three modules: OmicsEn, a variational autoencoder that maps gene expression profiles into a latent embedding $Z_O$; TextEn, a frozen SciBERT encoder that extracts a semantic embedding $Z_D$ from molecular text; and DiffGen, a diffusion Transformer that denoises vocabulary-aware remapped SELFIES sequences under the fused condition $Z = T(X_t + PE + TE_t, Z_D) \\oplus Z_O$. SELFIES, a molecular string format whose grammar guarantees chemically valid outputs, serves as the bridge representation linking omics and text to molecular structure and is what allows the reported perfect validity.","core_discovery":"The paper's central claim is that a conditional diffusion generator, jointly guided by an omics expression embedding and a molecular textual description embedding, produces hit-like molecules that are more valid, more unique, more novel, and structurally closer to known ligands than generation conditioned on omics alone or text alone. The authors support this with benchmark results on their own TextOmics dataset, with transfer experiments across ten cancer-related targets, and with a zero-shot Alzheimer's disease case study in which patient symptom narratives replace molecular descriptions.","pith_inferences":["Editorial inference: because each textual description is generated from the target molecule itself via BioT5, the text condition may largely re-encode the molecular structure rather than supply independent biological or semantic knowledge; the reported multimodal advantage could be an artifact of this text–structure coupling.","Editorial inference: a fairer test of genuine multimodal synergy would pair omics profiles with human-written or externally sourced descriptions for held-out molecules, or shuffle the text–molecule pairing to see whether the joint gains persist when the text is not derived from the target structure.","Editorial inference: the zero-shot therapeutic claim is limited to a single case study; extending it to rare diseases would require new disease-specific omics datasets and experimental validation of the generated candidates before any therapeutic conclusion can be drawn.","Editorial inference: the omics encoder's reconstruction fidelity, shown by PCA overlap, establishes that the latent represents the input profile, but it does not by itself establish that the omics condition carries pharmacological relevance beyond what the text already encodes."],"forward_implications":["On the ChemInduced test set, the full ToDi model reaches 100% validity, 98.45% uniqueness, and 97.30% novelty, improving novelty by 5.1 points over the omics-only variant and by 8.05 points over the previous omics-guided baseline.","Omitting omics still leaves a text-only variant that outperforms the text-diffusion baseline TGM on the ChEBI-20 benchmark across most reported metrics, indicating the text encoder alone carries substantial structural information.","Transferring ToDi to ten cancer-relevant targets yields 100% validity on every target and perfect MACCS Tanimoto scores on six of ten, suggesting the joint conditioning generalizes to target-specific ligand generation.","In the zero-shot Alzheimer's study, molecules generated under symptom-narrative guidance achieve higher Morgan and MACCS similarity to approved drugs than the prior baseline, suggesting the framework can operate without target-specific training text.","The ablation results imply that omics and text contribute complementary information, since each single-modality variant outperforms prior single-modality baselines while the joint model improves further."],"supporting_citations":[{"why":"Provides the primary omics-guided baseline (GxVAEs) that ToDi compares against on both ChemInduced and target-transfer tasks.","marker":"[10]"},{"why":"Provides the text-diffusion baseline TGM used for semantic and ChEBI-20 comparisons.","marker":"[14]"},{"why":"Supplies the SELFIES representation whose grammar guarantees chemical validity, the basis for ToDi's reported 100% validity.","marker":"[30]"},{"why":"SciBERT is the frozen encoder used as TextEn to produce semantic embeddings from molecular textual descriptions.","marker":"[36]"},{"why":"BioT5 is used to generate every molecular textual description in the TextOmics dataset, making it the source of the text condition.","marker":"[37]"},{"why":"LINCS L1000 is the source of the ChemInduced and TargetPerturb omics expression profiles.","marker":"[39]"},{"why":"CREEDS is the source of the DiseaseSign Alzheimer's disease omics profiles used in the zero-shot study.","marker":"[40]"}],"fun_headline_variants":["Joint omics-text conditioning boosts hit-like molecule design","Diffusion model fuses gene expression and text for drug discovery","TextOmics benchmark plus ToDi: multimodal molecules","Zero-shot therapeutic molecules from omics and text","Multimodal diffusion generates valid, novel hit-like molecules"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the BioT5-generated textual descriptions provide semantic information independent of the molecule's structure, yet each description is produced from the very molecule it is paired with, so the text may simply restate the structure and thereby predetermine the observed multimodal gains.","fun_headline_variants_meta":{"raw":{"variants":["Joint omics-text conditioning boosts hit-like molecule design","Diffusion model fuses gene expression and text for drug discovery","TextOmics benchmark plus ToDi: multimodal molecules","Zero-shot therapeutic molecules from omics and text","Multimodal diffusion generates valid, novel hit-like molecules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000165,"raw_usage":{"total_tokens":1181,"prompt_tokens":808,"completion_tokens":373,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":424,"completion_tokens_details":{"reasoning_tokens":295}},"tokens_in":424,"tokens_out":373,"duration_ms":4398,"temperature":1.0,"reasoning_tokens":295,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:42:03.805889+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Regenerate the benchmark's textual descriptions using descriptions authored independently of the target molecule, or sourced from curated chemistry texts, then rerun the ToDi versus text-only and omics-only ablations on the same test split; if the joint gains in novelty and structural similarity shrink or reverse, the claimed multimodal advantage is an artifact of text re-encoding the target structure rather than a genuine integration of complementary signals.","supporting_citations":[{"cited_title":"GxV AEs: Two joint V AEs generate hit molecules from gene expression profiles","cited_arxiv_id":null,"evidence_quote":"Provides the primary omics-guided baseline (GxVAEs) that ToDi compares against on both ChemInduced and target-transfer tasks."},{"cited_title":"Text-guided molecule generation with diffusion language model","cited_arxiv_id":null,"evidence_quote":"Provides the text-diffusion baseline TGM used for semantic and ChEBI-20 comparisons."},{"cited_title":"Self-referencing embedded strings (SELFIES): A 100% robust molecular string representation","cited_arxiv_id":null,"evidence_quote":"Supplies the SELFIES representation whose grammar guarantees chemical validity, the basis for ToDi's reported 100% validity."},{"cited_title":"SciBERT: A pretrained language model for scientific text","cited_arxiv_id":null,"evidence_quote":"SciBERT is the frozen encoder used as TextEn to produce semantic embeddings from molecular textual descriptions."},{"cited_title":"BioT5: Enriching cross-modal integration in biology with chemical knowledge and natural language associations","cited_arxiv_id":null,"evidence_quote":"BioT5 is used to generate every molecular textual description in the TextOmics dataset, making it the source of the text condition."},{"cited_title":"Lincs Canvas Browser: interactive web app to query, browse and interrogate LINCS L1000 gene expression signatures","cited_arxiv_id":null,"evidence_quote":"LINCS L1000 is the source of the ChemInduced and TargetPerturb omics expression profiles."},{"cited_title":"Extraction and analysis of signatures from the gene expression omnibus by the crowd","cited_arxiv_id":null,"evidence_quote":"CREEDS is the source of the DiseaseSign Alzheimer's disease omics profiles used in the zero-shot study."}],"review_version":1}