REVIEW 3 major objections 4 minor 1 cited by
A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A multi-modal copilot lets researchers drive single-cell annotation, generation, and drug prediction with natural-language instructions and matches foundation-model performance.
desk verdict Useful instruction-following architecture and dataset, but random cell-level splits leave the headline generalization claims unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a three-part multimodal cell-language architecture. A Q-Former-style cell encoder with eight learnable query vectors reads the raw gene-expression vector and produces a small set of cell embeddings; a pretrained text-to-text language model processes these embeddings together with the text tokens, so the same transformer stack reasons over both modalities; and a conditional variational autoencoder, using a zero-inflated negative binomial output distribution, reconstructs or generates exact integer expression counts when the model emits a special signal token. The special tokens that mark where cell data enters the token stream and the signal token that marks the handoff from text to cell generation are what let one set of weights answer with text, with a vector, or with an interleaved mix.
What would settle it
Train InstructCell on one set of tissues or sequencing batches and test it on tissues or batches that never appear in training; if cell type annotation and drug sensitivity accuracy fall to near chance, the claim of adapting to diverse experimental conditions is disproved.
Extended reading notes
Core claim
The central claim is that the discrete, numerical language of single-cell transcription and the flexible language of human instructions can be coupled in one end-to-end model, and that coupling is enough to make instruction following competitive with dedicated pretrained foundation models. InstructCell interleaves text and cell profiles in one token sequence: a cell profile is marked by special tokens, encoded by a learned query-based transformer, and fused with text in a pre-trained language-model backbone; for generation tasks the model emits a signal token whose hidden state conditions a count-based decoder that outputs a new gene-expression profile. The paper reports that on cell type annotation and drug sensitivity prediction InstructCell matches or exceeds the foundation-model baselines despite no large-scale single-cell pretraining, and on conditional pseudo-cell generation it produces distributions closer to real cells than the generative baselines it is compared with. It further claims this works across human and mouse data from multiple tissues and remains stable when instruction wording changes.
Load-bearing premise
The evaluation assumes that a random 8:1:1 split of each dataset, where test cells come from the same tissues and batches as training cells, tells us how the model will perform on genuinely new biological samples.
Editorial extensions
If this is right
- A single instruction-tuned model can replace task-specific pipelines for cell type annotation, conditional pseudo-cell generation, and drug sensitivity prediction.
- Because cell profiles stay in their native count modality, outputs preserve low-expression genes and exact integer values, which matters for downstream count-based analyses.
- The model generalizes to instruction templates and multiple-choice formats it has never seen, so users can phrase queries freely without retraining.
- Multi-task instruction tuning beats single-task tuning on every task, implying that the shared cell-language representation is a genuine asset rather than a compromise.
Reading between the lines
- The paper's random-split evaluation leaves untested how the model would fare on a wholly unseen tissue or sequencing batch, so a natural stress test is to hold out entire batches or organs before concluding that it adapts to diverse experimental conditions.
- The instruction-template synthesis recipe, which varies personality, motivation, and proficiency traits when prompting a large language model, could plausibly transfer to other molecular modalities such as chromatin accessibility or spatial transcriptomics with minimal architectural change.
- The marker-gene results hint that saliency over this architecture could double as a hypothesis generator for novel cell-type markers, but the paper only checks agreement with known markers, so prospective validation in the wet lab would be needed.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. InstructCell is a multi-modal instruction-tuned model for scRNA-seq that combines a Q-Former cell encoder, a T5-base language model, and a ZINB-based conditional VAE cell decoder. The authors synthesize a large GPT-4o-generated instruction template set, train instruct and chat variants on 11 public datasets, and evaluate the model on cell type annotation (CTA), conditional pseudo-cell generation (CPCG), and drug sensitivity prediction (DSP). The paper compares against scGPT, scBERT, Geneformer, scDiffusion, scGAN, and Cell2Sentence, validates marker-gene findings against CellMarker2.0, and includes ablations of the Q-Former, query count, multi-task tuning, pre-trained weights, and template diversity. The headline claims are that InstructCell meets or exceeds single-cell foundation models and adapts to diverse experimental conditions while preserving discrete gene-expression information.
Significance. If the evaluation supported them, the claims would be significant: InstructCell would provide a lightweight, instruction-following alternative to large single-cell foundation models, unifying understanding and generation tasks in one multi-modal framework. The paper has clear strengths: the models and code are publicly released; the architecture math in Eqs. (1)-(13) is standard and largely sound; comparisons are made against externally published baselines; marker-gene results are checked against an external database; and the template-diversity study is a useful analysis. The weakness is evaluational rather than conceptual: the random cell-level split protocol does not test the core generalization claim, and most headline comparisons lack uncertainty estimates. These gaps are fixable with additional experiments, which is why I recommend major revision rather than rejection.
major comments (3)
- [Methods, Experimental setup; Fig. 9; Figs. 4-5] The central claim that InstructCell 'adapts to diverse experimental conditions' is not supported by the current evaluation protocol. The Methods state that 'each dataset is divided into training, validation, and test sets in an 8:1:1 ratio' and that multi-task tuning 'merge[s] all training splits into a single mixed training set.' For DSP, the datasets GSE117872, GSE149383, and GSE110894 are composed of a small number of patients/samples, and all cells from one sample share the same drug-sensitivity label; a random cell-level split places cells from the same sample in both training and test sets, allowing the model to memorize sample-specific expression signatures. The near-perfect accuracies in Fig. 4(c) (>0.95) are consistent with such leakage. The same random-split design is used for CTA and CPCG, so the model is never tested on an unseen tissue, batch, donor, or study; the 'unseen template' analysis in Fig. 5(a) varies only instruction phrasing on the same held-out cells. I request a held-out-patient or leave-one-dataset-out evaluation (with per-fold metrics) for DSP, and a leave-one-dataset-out or cross-tissue protocol for CTA and CPCG, before the generalization claim can be accepted.
- [Results, Figs. 3(a), 4(a), 7(b-e)] The main model-vs-baseline comparisons are reported as single values without error bars, confidence intervals, or significance tests. Fig. 3(a) and Fig. 4(a) show colored bars for CTA and DSP, and the ablations in Fig. 7(b-e) compare single average values. Because the headline claim is that InstructCell 'consistently meets or exceeds' several baselines, differences that could be within run-to-run noise need to be quantified. The authors should report mean and standard deviation over at least 3-5 random seeds or split instantiations, and include a paired significance test (e.g., McNemar's test or a bootstrap) for the classification comparisons, with analogous confidence estimates for the CPCG metrics.
- [Results, 'InstructCell enables conditional pseudo-cell generation'; Methods, 'Metrics'] The second strong claim—that InstructCell 'preserves the discrete nature of gene expression profiles'—is not actually tested by the CPCG evaluation. All CPCG metrics (MMD, ΔsKNN, pKNN) are computed after normalizing counts to 10,000, applying a log1p transform, reducing to 50 principal components, and embedding with UMAP (Methods, Metrics, Eqs. 14-17). These transformations remove the count-level information that the discrete-modality design is intended to protect; a model producing smoothed continuous values could receive similar scores. I ask the authors to add count-level fidelity checks (e.g., raw-count histograms, zero-inflation statistics, NB/ZINB likelihood, or comparisons of the generated count distributions against real data) or to moderate the claim accordingly.
minor comments (4)
- [Fig. 4 caption] The caption says 'Evaluation of InstructCell's CTA performance across human oral, lung, and mouse bone datasets' and refers to 'predicted cell types'; both should refer to drug sensitivity prediction and drug-sensitivity labels.
- [Results, 'InstructCell enables conditional pseudo-cell generation'] The text says the CPCG experiments use '9 tissues—bladder, blood, liver, lung, spleen, thymus, and vasculature' but lists only seven tissues; the count and list should be reconciled with the dataset table in Fig. 9.
- [Results, 'InstructCell boosts the performance of cell type annotation'] The sentence that InstructCell works 'despite not relying on large-scale unlabeled pre-training' is imprecise because the backbone is a pre-trained T5-base model, and Fig. 7(e) shows that removing pre-trained weights degrades classification performance; the authors should qualify what kind of pre-training they mean.
- [References] The T5 paper is cited twice as references [25] and [70]; this duplicate should be removed, and reference [49] should give the full author list and venue for 'Attention is all you need'.
Circularity Check
No significant circularity: headline results are benchmarked against external baselines and an external marker-gene database, and no load-bearing derivation reduces to its own inputs.
full rationale
InstructCell's central claims are supported by comparisons against external baselines (scBERT, scGPT, Geneformer, scDiffusion, scGAN, Cell2Sentence) using fixed 8:1:1 data splits, and marker-gene findings are checked against the external CellMarker2.0 database. The generative objective is a standard conditional-VAE ELBO (Eq. 7) trained to reconstruct target expression profiles, not defined in terms of the reported MMD/sKNN/pKNN metrics. The 'LLM-as-a-judge' quality check uses Claude 3.5 Sonnet on outputs of a model trained with GPT-4o-generated templates; while this is an internal quality assessment rather than human evaluation, it is not load-bearing for the headline performance claims, and the judge is architecturally different from the template generator. No load-bearing self-citation, imported uniqueness theorem, or ansatz-smuggling citation is present; self-citations such as Mol-Instructions appear only in related-work context. The random cell-level split concern raised in the reader's take is a generalization/data-leakage validity concern, not a circularity of the derivation chain, so it does not affect this score.
Assumptions & free parameters
free parameters (5)
- Q-Former query count k =
8
- KL weight alpha in CVAE loss =
not reported
- Rare cell type exclusion threshold =
fewer than 20 cells
- Gene filtering thresholds =
min 200 genes/cell; min 8 cells/gene; top 3,600 HVGs per dataset
- Template ratio for final model =
100% of templates
assumptions (5)
- domain assumption scRNA-seq counts are well described by a Zero-Inflated Negative Binomial distribution
- domain assumption The 256-dimensional Gaussian latent z_s with an unconditional prior N(0,I) is sufficient to represent cell identity when conditioned through the decoder functions f1-f4
- domain assumption MMD, sKNN, and pKNN computed in a 50-PCA then 2D-UMAP embedding space capture biologically meaningful differences between real and generated cells
- domain assumption GPT-4o-generated instruction templates are a valid proxy for real researcher queries
- standard math Standard ELBO and reparameterization identities for VAEs
Cite this review
Pith. "Pith review of A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following." pith.science (2026). https://pith.science/paper/NA5ALG6F
@misc{pith2026250108187,
author = {Pith},
title = {Pith review of: A Multi-Modal AI Copilot for Single-Cell Analysis with Instruction Following},
year = {2026},
howpublished = {\url{https://pith.science/paper/NA5ALG6F}},
note = {Machine review of arXiv:2501.08187}
}
read the original abstract
Large language models excel at interpreting complex natural language instructions, enabling them to perform a wide range of tasks. In the life sciences, single-cell RNA sequencing (scRNA-seq) data serves as the "language of cellular biology", capturing intricate gene expression patterns at the single-cell level. However, interacting with this "language" through conventional tools is often inefficient and unintuitive, posing challenges for researchers. To address these limitations, we present InstructCell, a multi-modal AI copilot that leverages natural language as a medium for more direct and flexible single-cell analysis. We construct a comprehensive multi-modal instruction dataset that pairs text-based instructions with scRNA-seq profiles from diverse tissues and species. Building on this, we develop a multi-modal cell language architecture capable of simultaneously interpreting and processing both modalities. InstructCell empowers researchers to accomplish critical tasks-such as cell type annotation, conditional pseudo-cell generation, and drug sensitivity prediction-using straightforward natural language commands. Extensive evaluations demonstrate that InstructCell consistently meets or exceeds the performance of existing single-cell foundation models, while adapting to diverse experimental conditions. More importantly, InstructCell provides an accessible and intuitive tool for exploring complex single-cell data, lowering technical barriers and enabling deeper biological insights.
Figures
Figures from the paper (10 more)
Forward citations
Cited by 1 Pith paper
-
Cell-o1: Training LLMs to Solve Single-Cell Reasoning Puzzles with Reinforcement Learning
A 7B model trained with reasoning distillation and reinforcement learning reaches 32.9% batch-level accuracy on a new single-cell annotation benchmark, versus 19.0% for OpenAI's o1.
Reference graph
Works this paper leans on
-
[1]
https://openai.com/gpt-4, Accessed 15 Apr 2024
GPT-4. https://openai.com/gpt-4, Accessed 15 Apr 2024
2024
-
[2]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradb...
2023
-
[3]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ ee Lacroix, Baptiste Rozi` ere, Naman Goyal, Eric Hambro, Faisal Azhar, Aur´ elien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language models. CoRR, abs/2302.13971, 2023
arXiv 2023
-
[4]
https://www.anthropic.com/claude, Accessed 27 Jun 2024
Claude anthropic. https://www.anthropic.com/claude, Accessed 27 Jun 2024
2024
-
[5]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human fee...
2022
-
[6]
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian- Jian Jiang, Han Wang, Matteo Manica, ...
2022
-
[7]
Gene expression data analysis
Alvis Brazma and Jaak Vilo. Gene expression data analysis. FEBS letters, 480(1):17–24, 2000
2000
-
[8]
Cell type atlas and lineage tree of a whole complex animal by single-cell transcriptomics
Mireya Plass, Jordi Solana, F Alexander Wolf, Salah Ayoub, Aristotelis Misios, Petar Glaˇ zar, Benedikt Obermayer, Fabian J Theis, Christine Kocks, and Nikolaus Rajewsky. Cell type atlas and lineage tree of a whole complex animal by single-cell transcriptomics. Science, 360(6391):eaaq1723, 2018
2018
Show all 116 references
-
[9]
The single-cell transcriptional landscape of mammalian organogenesis
Junyue Cao, Malte Spielmann, Xiaojie Qiu, Xingfan Huang, Daniel M Ibrahim, Andrew J Hill, Fan Zhang, Stefan Mundlos, Lena Christiansen, Frank J Steemers, et al. The single-cell transcriptional landscape of mammalian organogenesis. Nature, 566(7745):496–502, 2019
2019
-
[10]
Ncbi geo: archive for functional genomics data sets—update
Tanya Barrett, Stephen E Wilhite, Pierre Ledoux, Carlos Evangelista, Irene F Kim, Maxim Tomashevsky, Kimberly A Marshall, Katherine H Phillippy, Patti M Sherman, Michelle Holko, et al. Ncbi geo: archive for functional genomics data sets—update. Nucleic acids research, 41(D1):D...
2012
-
[11]
The human cell atlas
Aviv Regev, Sarah A Teichmann, Eric S Lander, Ido Amit, Christophe Benoist, Ewan Birney, Bernd Bodenmiller, Peter Campbell, Piero Carninci, Menna Clatworthy, et al. The human cell atlas. elife, 6:e27041, 2017
2017
-
[12]
scbert as a large-scale pretrained deep language model for cell type annotation of single-cell rna-seq data
Fan Yang, Wenchuan Wang, Fang Wang, Yuan Fang, Duyu Tang, Junzhou Huang, Hui Lu, and Jianhua Yao. scbert as a large-scale pretrained deep language model for cell type annotation of single-cell rna-seq data. Nature Machine Intelligence, 4(10):852–866, 2022
2022
-
[13]
Transfer learning enables predictions in network biology
Christina V Theodoris, Ling Xiao, Anant Chopra, Mark D Chaffin, Zeina R Al Sayed, Matthew C Hill, Helene Mantineo, Elizabeth M Brydon, Zexian Zeng, X Shirley Liu, et al. Transfer learning enables predictions in network biology. Nature, 618(7965):616–624, 2023. 31
2023
-
[14]
scgpt: toward building a foundation model for single-cell multi-omics using generative ai
Haotian Cui, Chloe Wang, Hassaan Maan, Kuan Pang, Fengning Luo, Nan Duan, and Bo Wang. scgpt: toward building a foundation model for single-cell multi-omics using generative ai. Nature Methods, pages 1–11, 2024
2024
-
[15]
Large-scale foundation model on single-cell transcriptomics
Minsheng Hao, Jing Gong, Xin Zeng, Chiming Liu, Yucheng Guo, Xingyi Cheng, Taifeng Wang, Jianzhu Ma, Xuegong Zhang, and Le Song. Large-scale foundation model on single-cell transcriptomics. Nature Methods, pages 1–11, 2024
2024
-
[16]
Cell2sentence: Teaching large language models the language of biology
Daniel LeVine, Syed Asad Rizvi, Sacha L´ evy, Nazreen Pallikkavaliyaveetil, David Zhang, Xingyu Chen, Sina Ghadermarzi, Ruiming Wu, Zihe Zheng, Ivan Vrkic, Anna Zhong, Daphne Raskin, Insu Han, Antonio Hen- rique de Oliveira Fonseca, Josue Ortega Caro, Amin Karbasi, Rahul Madha...
2024
-
[17]
Assessing gpt-4 for cell type annotation in single-cell rna-seq analysis
Wenpin Hou and Zhicheng Ji. Assessing gpt-4 for cell type annotation in single-cell rna-seq analysis. Nature Methods, pages 1–4, 2024
2024
-
[18]
Simple and effective embedding model for single-cell biology built from chatgpt
Yiqun Chen and James Zou. Simple and effective embedding model for single-cell biology built from chatgpt. Nature Biomedical Engineering, pages 1–11, 2024
2024
-
[19]
scelmo: Embeddings from language models are good learners for single-cell data analysis
Tianyu Liu, Tianqi Chen, Wangjie Zheng, Xiao Luo, and Hongyu Zhao. scelmo: Embeddings from language models are good learners for single-cell data analysis. bioRxiv, pages 2023–12, 2023
2023
-
[20]
https://openai.com/index/hello-gpt-4o , Accessed 29 May 2024
GPT-4o. https://openai.com/index/hello-gpt-4o , Accessed 29 May 2024
2024
-
[21]
Enhancing chat language models by scaling high-quality instructional conversations
Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. Enhancing chat language models by scaling high-quality instructional conversations. In EMNLP, pages 3029–3051. Association for Computational Linguistics, 2023
2023
-
[22]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven C. H. Hoi. BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In ICML, volume 202 of Proceedings of Machine Learning Research, pages 19730–19742. PMLR, 2023
2023
-
[23]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019
2019
-
[24]
BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. In ACL, pages 7871–7880. Associ...
2020
-
[25]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21:140:1–140:67, 2020
2020
-
[26]
Towards open-ended visual quality comparison
Haoning Wu, Hanwei Zhu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Chunyi Li, Annan Wang, Wenxiu Sun, Qiong Yan, Xiaohong Liu, Guangtao Zhai, Shiqi Wang, and Weisi Lin. Towards open-ended visual quality comparison. In ECCV (3), volume 15061 of Lecture Notes in Compu...
2024
-
[27]
Mantis: Interleaved multi-image instruction tuning
Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning. arXiv preprint arXiv:2405.01483, 2024
2024 arXiv
-
[28]
Next-gpt: Any-to-any multimodal llm
Shengqiong Wu, Hao Fei, Leigang Qu, Wei Ji, and Tat-Seng Chua. Next-gpt: Any-to-any multimodal llm. arXiv preprint arXiv:2309.05519, 2023
2023 arXiv
-
[29]
Generating images with multimodal language models
Jing Yu Koh, Daniel Fried, and Russ Salakhutdinov. Generating images with multimodal language models. In NeurIPS, 2023
2023
-
[30]
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learning structured output representation using deep conditional generative models. In NIPS, pages 3483–3491, 2015. 32
2015
-
[31]
The generalization of ‘student’s’problem when several different population varlances are involved
Bernard L Welch. The generalization of ‘student’s’problem when several different population varlances are involved. Biometrika, 34(1-2):28–35, 1947
1947
-
[32]
scdiffusion: conditional generation of high-quality single-cell data using diffusion model
Erpai Luo, Minsheng Hao, Lei Wei, and Xuegong Zhang. scdiffusion: conditional generation of high-quality single-cell data using diffusion model. Bioinformatics, 40(9):btae518, 2024
2024
-
[33]
Realistic in silico generation and augmentation of single-cell rna-seq data using generative adversarial networks
Mohamed Marouf, Pierre Machart, Vikas Bansal, Christoph Kilian, Daniel S Magruder, Christian F Krebs, and Stefan Bonn. Realistic in silico generation and augmentation of single-cell rna-seq data using generative adversarial networks. Nature communications, 11(1):166, 2020
2020
-
[34]
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. Deep inside convolutional networks: Visualising image classification models and saliency maps. In ICLR (Workshop Poster), 2014
2014
-
[35]
Cellmarker 2.0: an updated database of manually curated cell markers in human/mouse and web tools based on scrna-seq data
Congxue Hu, Tengyue Li, Yingqi Xu, Xinxin Zhang, Feng Li, Jing Bai, Jing Chen, Wenqi Jiang, Kaiyue Yang, Qi Ou, Xia Li, Peng Wang, and Yunpeng Zhang. Cellmarker 2.0: an updated database of manually curated cell markers in human/mouse and web tools based on scrna-seq data. Nucl...
2023
-
[36]
A single-cell transcriptome atlas of the human pancreas
Mauro J Muraro, Gitanjali Dharmadhikari, Dominic Gr¨ un, Nathalie Groen, Tim Dielen, Erik Jansen, Leon Van Gurp, Marten A Engelse, Francoise Carlotti, Eelco Jp De Koning, et al. A single-cell transcriptome atlas of the human pancreas. Cell systems, 3(4):385–394, 2016
2016
-
[37]
Generation of human islet cell type-specific identity genesets
L´ eon van Gurp, Leon Fodoulian, Daniel Oropeza, Kenichiro Furuyama, Eva Bru-Tari, Anh Nguyet Vu, John S Kaddis, Iv´ an Rodr ´ ıguez, Fabrizio Thorel, and Pedro L Herrera. Generation of human islet cell type-specific identity genesets. Nature communications, 13(1):2020, 2022
2020
-
[38]
Xing, Hao Zhang, Joseph E
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS, 2023
2023
-
[39]
Llms as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. Llms as narcissistic evaluators: When ego inflates evaluation scores. CoRR, abs/2311.09766, 2023
2023 arXiv
-
[40]
Rabiul Awal, Rui Cao, Roy Ka-Wei Lee, and Sandra Mitrovic
Md. Rabiul Awal, Rui Cao, Roy Ka-Wei Lee, and Sandra Mitrovic. Angrybert: Joint learning target and emotion for hate speech detection. In PAKDD (1), volume 12712 of Lecture Notes in Computer Science, pages 701–713. Springer, 2021
2021
-
[41]
A survey on multi-task learning
Yu Zhang and Qiang Yang. A survey on multi-task learning. IEEE Trans. Knowl. Data Eng., 34(12):5586–5609, 2022
2022
-
[42]
Predicting transcriptional outcomes of novel multigene perturba- tions with gears
Yusuf Roohani, Kexin Huang, and Jure Leskovec. Predicting transcriptional outcomes of novel multigene perturba- tions with gears. Nature Biotechnology, 42(6):927–935, 2024
2024
-
[43]
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le. Finetuned language models are zero-shot learners. arXiv preprint arXiv:2109.01652, 2021
2021 arXiv
-
[44]
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al. Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53, 2024
2024
-
[45]
Zero-shot generalization during instruction tuning: Insights from similarity and granularity
Bingxiang He, Ning Ding, Cheng Qian, Jia Deng, Ganqu Cui, Lifan Yuan, Huan-ang Gao, Huimin Chen, Zhiyuan Liu, and Maosong Sun. Zero-shot generalization during instruction tuning: Insights from similarity and granularity. arXiv preprint arXiv:2406.11721, 2024
2024 arXiv
-
[46]
Peakvi: A deep generative model for single-cell chromatin accessibility analysis
Tal Ashuach, Daniel A Reidenbach, Adam Gayoso, and Nir Yosef. Peakvi: A deep generative model for single-cell chromatin accessibility analysis. Cell reports methods, 2(3), 2022
2022
-
[47]
Muse-gnn: learning unified gene representation from multimodal biological graph data
Tianyu Liu, Yuge Wang, Rex Ying, and Hongyu Zhao. Muse-gnn: learning unified gene representation from multimodal biological graph data. Advances in neural information processing systems, 36, 2024. 33
2024
-
[48]
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pages 74–81, 2004
2004
-
[49]
Attention is all you need
Ashish Vaswani. Attention is all you need. arXiv preprint arXiv:1706.03762, 2017
2017 arXiv
-
[50]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016
2016
-
[51]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023
2023
-
[52]
Auto-encoding variational bayes
DP Kingma. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013
2013 arXiv
-
[53]
Moderated estimation of fold change and dispersion for rna-seq data with deseq2
Michael I Love, Wolfgang Huber, and Simon Anders. Moderated estimation of fold change and dispersion for rna-seq data with deseq2. Genome biology, 15:1–21, 2014
2014
-
[54]
Comparison and evaluation of statistical error models for scrna-seq
Saket Choudhary and Rahul Satija. Comparison and evaluation of statistical error models for scrna-seq. Genome biology, 23(1):27, 2022
2022
-
[55]
Systematic comparison of single-cell and single-nucleus rna-sequencing methods
Jiarui Ding, Xian Adiconis, Sean K Simmons, Monika S Kowalczyk, Cynthia C Hession, Nemanja D Marjanovic, Travis K Hughes, Marc H Wadsworth, Tyler Burks, Lan T Nguyen, et al. Systematic comparison of single-cell and single-nucleus rna-sequencing methods. Nature biotechnology, 3...
2020
-
[56]
Consequences and opportunities arising due to sparser single-cell rna-seq datasets
Gerard A Bouland, Ahmed Mahfouz, and Marcel JT Reinders. Consequences and opportunities arising due to sparser single-cell rna-seq datasets. Genome biology, 24(1):86, 2023
2023
-
[57]
powsimr: power analysis for bulk and single cell rna-seq experiments
Beate Vieth, Christoph Ziegenhain, Swati Parekh, Wolfgang Enard, and Ines Hellmann. powsimr: power analysis for bulk and single cell rna-seq experiments. Bioinformatics, 33(21):3486–3488, 2017
2017
-
[58]
Single-cell rna-seq denoising using a deep count autoencoder
G¨ okcen Eraslan, Lukas M Simon, Maria Mircea, Nikola S Mueller, and Fabian J Theis. Single-cell rna-seq denoising using a deep count autoencoder. Nature communications, 10(1):390, 2019
2019
-
[59]
Deep generative modeling for single-cell transcriptomics
Romain Lopez, Jeffrey Regier, Michael B Cole, Michael I Jordan, and Nir Yosef. Deep generative modeling for single-cell transcriptomics. Nature methods, 15(12):1053–1058, 2018
2018
-
[60]
Controlvae: Controllable variational autoencoder
Huajie Shao, Shuochao Yao, Dachun Sun, Aston Zhang, Shengzhong Liu, Dongxin Liu, Jun Wang, and Tarek Abdelzaher. Controlvae: Controllable variational autoencoder. In International conference on machine learning, pages 8655–8664. PMLR, 2020
2020
-
[61]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598, 2022
2022 arXiv
-
[63]
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 2019
1907 arXiv
-
[64]
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. ALBERT: A lite BERT for self-supervised learning of language representations. In ICLR. OpenReview.net, 2020
2020
-
[65]
Deberta: decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. Deberta: decoding-enhanced bert with disentangled attention. In ICLR. OpenReview.net, 2021
2021
-
[66]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language understanding by generative pre-training. 2018
2018
-
[67]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. 34
1901
-
[68]
Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021
Ben Wang and Aran Komatsuzaki. Gpt-j-6b: A 6 billion parameter autoregressive language model, 2021
2021
-
[69]
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon. Unified language model pre-training for natural language understanding and generation. In NeurIPS, pages 13042–13054, 2019
2019
-
[70]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020
2020
-
[71]
Pegasus: Pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International conference on machine learning, pages 11328–11339. PMLR, 2020
2020
-
[72]
A bayesian/information theoretic model of learning to learn via multiple task sampling
Jonathan Baxter. A bayesian/information theoretic model of learning to learn via multiple task sampling. Machine learning, 28:7–39, 1997
1997
-
[73]
An overview of multi-task learning in deep neural networks
Sebastian Ruder. An overview of multi-task learning in deep neural networks. arXiv preprint arXiv:1706.05098, 2017
2017 arXiv
-
[74]
xfinder: Robust and pinpoint answer extraction for large language models
Qingchen Yu, Zifan Zheng, Shichao Song, Zhiyu Li, Feiyu Xiong, Bo Tang, and Ding Chen. xfinder: Robust and pinpoint answer extraction for large language models. CoRR, abs/2405.11874, 2024
2024 arXiv
-
[75]
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426, 2018
2018 arXiv
-
[76]
A kernel two-sample test
Arthur Gretton, Karsten M Borgwardt, Malte J Rasch, Bernhard Sch¨ olkopf, and Alexander Smola. A kernel two-sample test. The Journal of Machine Learning Research, 13(1):723–773, 2012
2012
-
[77]
Removal of batch effects using distribution-matching residual networks
Uri Shaham, Kelly P Stanton, Jun Zhao, Huamin Li, Khadir Raddassi, Ruth Montgomery, and Yuval Kluger. Removal of batch effects using distribution-matching residual networks. Bioinformatics, 33(16):2539–2546, 2017
2017
-
[78]
Goodfellow, Moritz Hardt, and Been Kim
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian J. Goodfellow, Moritz Hardt, and Been Kim. Sanity checks for saliency maps. In NeurIPS, pages 9525–9536, 2018
2018
-
[79]
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of machine learning research, 9(11), 2008
2008
-
[80]
Dimensionality reduction for visualizing single-cell data using umap
Etienne Becht, Leland McInnes, John Healy, Charles-Antoine Dutertre, Immanuel WH Kwok, Lai Guan Ng, Florent Ginhoux, and Evan W Newell. Dimensionality reduction for visualizing single-cell data using umap. Nature biotechnology, 37(1):38–44, 2019
2019
-
[81]
Scanpy: large-scale single-cell gene expression data analysis
F Alexander Wolf, Philipp Angerer, and Fabian J Theis. Scanpy: large-scale single-cell gene expression data analysis. Genome biology, 19:1–5, 2018
2018
-
[82]
Scater: pre-processing, quality control, normalization and visualization of single-cell rna-seq data in r
Davis J McCarthy, Kieran R Campbell, Aaron TL Lun, and Quin F Wills. Scater: pre-processing, quality control, normalization and visualization of single-cell rna-seq data in r. Bioinformatics, 33(8):1179–1186, 2017
2017
-
[83]
Rna sequencing of single human islet cells reveals type 2 diabetes genes
Yurong Xin, Jinrang Kim, Haruka Okamoto, Min Ni, Yi Wei, Christina Adler, Andrew J Murphy, George D Yancopoulos, Calvin Lin, and Jesper Gromada. Rna sequencing of single human islet cells reveals type 2 diabetes genes. Cell metabolism, 24(4):608–615, 2016
2016
-
[84]
A transcriptional cross species map of pancreatic islet cells
Sophie Tritschler, Moritz Thomas, Anika B¨ ottcher, Barbara Ludwig, Janine Schmid, Undine Schubert, Elisabeth Kemter, Eckhard Wolf, Heiko Lickert, and Fabian J Theis. A transcriptional cross species map of pancreatic islet cells. Molecular Metabolism, 66:101595, 2022
2022
-
[85]
https://github.com/openvax/pyensembl?tab=readme-ov-file
pyEnsembl. https://github.com/openvax/pyensembl?tab=readme-ov-file
-
[86]
https://github.com/jrderuiter/pybiomart
pybiomart. https://github.com/jrderuiter/pybiomart
-
[87]
Spatial reconstruction of single-cell gene expression data
Rahul Satija, Jeffrey A Farrell, David Gennert, Alexander F Schier, and Aviv Regev. Spatial reconstruction of single-cell gene expression data. Nature biotechnology, 33(5):495–502, 2015. 35
2015
-
[88]
Comprehensive integration of single-cell data
Tim Stuart, Andrew Butler, Paul Hoffman, Christoph Hafemeister, Efthymia Papalexi, William M Mauck, Yuhan Hao, Marlon Stoeckius, Peter Smibert, and Rahul Satija. Comprehensive integration of single-cell data. cell, 177(7):1888–1902, 2019
1902
-
[89]
Single-cell transcriptome profiling of an adult human cell atlas of 15 major organs
Shuai He, Lin-He Wang, Yang Liu, Yi-Qi Li, Hai-Tian Chen, Jing-Hong Xu, Wan Peng, Guo-Wang Lin, Pan-Pan Wei, Bo Li, et al. Single-cell transcriptome profiling of an adult human cell atlas of 15 major organs. Genome biology, 21:1–34, 2020
2020
-
[90]
Single-cell transcriptome profiling of human pancreatic islets in health and type 2 diabetes
˚Asa Segerstolpe, Athanasia Palasantza, Pernilla Eliasson, Eva-Marie Andersson, Anne-Christine Andr´ easson, Xiaoyan Sun, Simone Picelli, Alan Sabirsh, Maryam Clausen, Magnus K Bjursell, et al. Single-cell transcriptome profiling of human pancreatic islets in health and type 2...
2016
-
[91]
Chromatin potential identified by shared single-cell profiling of rna and chromatin
Sai Ma, Bing Zhang, Lindsay M LaFave, Andrew S Earl, Zachary Chiang, Yan Hu, Jiarui Ding, Alison Brack, Vinay K Kartha, Tristan Tay, et al. Chromatin potential identified by shared single-cell profiling of rna and chromatin. Cell, 183(4):1103–1116, 2020
2020
-
[92]
Comprehensive single cell mrna profiling reveals a detailed roadmap for pancreatic endocrinogenesis
Aim´ ee Bastidas-Ponce, Sophie Tritschler, Leander Dony, Katharina Scheibner, Marta Tarquis-Medina, Ciro Salinno, Silvia Schirge, Ingo Burtscher, Anika B¨ ottcher, Fabian J Theis, et al. Comprehensive single cell mrna profiling reveals a detailed roadmap for pancreatic endocri...
2019
-
[93]
Longitudinal single-cell rna sequencing of patient-derived primary cells reveals drug-induced infidelity in stem cell hierarchy
Ankur Sharma, Elaine Yiqun Cao, Vibhor Kumar, Xiaoqian Zhang, Hui Sun Leong, Angeline Mei Lin Wong, Neeraja Ramakrishnan, Muhammad Hakimullah, Hui Min Vivian Teo, Fui Teen Chong, et al. Longitudinal single-cell rna sequencing of patient-derived primary cells reveals drug-induc...
2018
-
[94]
Single-cell transcriptional changes associated with drug tolerance and response to combination therapies in cancer
Alexandre F Aissa, Abul BMMK Islam, Majd M Ariss, Cammille C Go, Alexandra E Rader, Ryan D Conrardy, Alexa M Gajda, Carlota Rubio-Perez, Klara Valyi-Nagy, Mary Pasquinelli, et al. Single-cell transcriptional changes associated with drug tolerance and response to combination th...
2021
-
[95]
Targeting enhancer switching overcomes non-genetic drug resistance in acute myeloid leukaemia
Charles C Bell, Katie A Fennell, Yih-Chih Chan, Florian Rambow, Miriam M Yeung, Dane Vassiliadis, Luis Lara, Paul Yeh, Luciano G Martelotto, Aljosja Rogiers, et al. Targeting enhancer switching overcomes non-genetic drug resistance in acute myeloid leukaemia. Nature communicat...
2019
-
[96]
Massively parallel digital transcriptional profiling of single cells
Grace XY Zheng, Jessica M Terry, Phillip Belgrader, Paul Ryvkin, Zachary W Bent, Ryan Wilson, Solongo B Ziraldo, Tobias D Wheeler, Geoff P McDermott, Junjie Zhu, et al. Massively parallel digital transcriptional profiling of single cells. Nature communications, 8(1):14049, 2017
2017
-
[97]
The tabula sapiens: A multiple-organ, single-cell transcriptomic atlas of humans
The Tabula Sapiens Consortium*, Robert C Jones, Jim Karkanias, Mark A Krasnow, Angela Oliveira Pisco, Stephen R Quake, Julia Salzman, Nir Yosef, Bryan Bulthaup, Phillip Brown, et al. The tabula sapiens: A multiple-organ, single-cell transcriptomic atlas of humans. Science, 376...
2022
-
[98]
Nature, 583(7817):590–595, 2020
A single-cell transcriptomic atlas characterizes ageing tissues in the mouse. Nature, 583(7817):590–595, 2020
2020
-
[99]
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning, pages 4596–4604. PMLR, 2018
2018
-
[100]
Robust enumeration of cell subsets from tissue expression profiles
Aaron M Newman, Chih Long Liu, Michael R Green, Andrew J Gentles, Weiguo Feng, Yue Xu, Chuong D Hoang, Maximilian Diehn, and Ash A Alizadeh. Robust enumeration of cell subsets from tissue expression profiles. Nature methods, 12(5):453–457, 2015
2015
-
[101]
Visualization and analysis of gene expression in tissue sections by spatial transcriptomics
Patrik L St ˚ ahl, Fredrik Salm´ en, Sanja Vickovic, Anna Lundmark, Jos´ e Fern´ andez Navarro, Jens Magnusson, Stefania Giacomello, Michaela Asp, Jakub O Westholm, Mikael Huss, et al. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics....
2016
-
[102]
Metabolic reprogramming: a hallmark of viral oncogenesis
P Levy and B Bartosch. Metabolic reprogramming: a hallmark of viral oncogenesis. Oncogene, 35(32):4155–4164, 2016. 36
2016
-
[103]
Single-cell rna-seq reveals new types of human blood dendritic cells, monocytes, and progenitors
Alexandra-Chlo´ e Villani, Rahul Satija, Gary Reynolds, Siranush Sarkizova, Karthik Shekhar, James Fletcher, Morgane Griesbeck, Andrew Butler, Shiwei Zheng, Suzan Lazo, et al. Single-cell rna-seq reveals new types of human blood dendritic cells, monocytes, and progenitors. Sci...
2017
-
[104]
Tools for the analysis of high-dimensional single-cell rna sequencing data
Yan Wu and Kun Zhang. Tools for the analysis of high-dimensional single-cell rna sequencing data. Nature Reviews Nephrology, 16(7):408–421, 2020
2020
-
[105]
Causal machine learning for single-cell genomics
Alejandro Tejada-Lapuerta, Paul Bertin, Stefan Bauer, Hananeh Aliee, Yoshua Bengio, and Fabian J Theis. Causal machine learning for single-cell genomics. arXiv preprint arXiv:2310.14935, 2023
2023 arXiv
-
[106]
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. In NAACL-HLT (1), pages 4171–4186. Association for Computational Linguistics, 2019
2019
-
[107]
Machine intelligence in single-cell data analysis: advances and new challenges
Jiajia Liu, Zhiwei Fan, Weiling Zhao, and Xiaobo Zhou. Machine intelligence in single-cell data analysis: advances and new challenges. Frontiers in Genetics, 12:655536, 2021
2021
-
[108]
Algorithmic advances in machine learning for single-cell expression analysis
Sergio Oller-Moreno, Karin Kloiber, Pierre Machart, and Stefan Bonn. Algorithmic advances in machine learning for single-cell expression analysis. Current Opinion in Systems Biology, 25:27–33, 2021
2021
-
[109]
Machine learning for perturbational single-cell omics
Yuge Ji, Mohammad Lotfollahi, F Alexander Wolf, and Fabian J Theis. Machine learning for perturbational single-cell omics. Cell Systems, 12(6):522–537, 2021
2021
-
[110]
Single cells make big data: new challenges and opportunities in transcriptomics
Philipp Angerer, Lukas Simon, Sophie Tritschler, F Alexander Wolf, David Fischer, and Fabian J Theis. Single cells make big data: new challenges and opportunities in transcriptomics. Current opinion in systems biology, 4:85–91, 2017
2017
-
[111]
Langcell: Language-cell pre-training for cell identity understanding
Suyuan Zhao, Jiahuan Zhang, Yushuai Wu, Yizhen Luo, and Zaiqing Nie. Langcell: Language-cell pre-training for cell identity understanding. In ICML. OpenReview.net, 2024
2024
-
[112]
Translation between molecules and natural language
Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. Translation between molecules and natural language. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 375–413, Abu Dhabi, United Arab Emirates, December...
2022
-
[113]
Mol-instructions: A large-scale biomolecular instruction dataset for large language models
Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. Mol-instructions: A large-scale biomolecular instruction dataset for large language models. ICLR, 2024
2024
-
[114]
Chatmol: interactive molecular discovery with natural language
Zheni Zeng, Bangchen Yin, Shipeng Wang, Jiarui Liu, Cheng Yang, Haishen Yao, Xingzhi Sun, Maosong Sun, Guotong Xie, and Zhiyuan Liu. Chatmol: interactive molecular discovery with natural language. Bioinformatics, 40(9):btae534, 2024
2024
-
[115]
Mollm: a unified language model for integrating biomedical text with 2d and 3d molecular representations
Xiangru Tang, Andrew Tran, Jeffrey Tan, and Mark B Gerstein. Mollm: a unified language model for integrating biomedical text with 2d and 3d molecular representations. Bioinformatics, 40(Supplement 1):i357–i368, 2024
2024
-
[116]
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge. Nature, 620(7972):172–180, 2023
2023
-
[117]
Boiko Daniil, MacKnight Robert, Kline Ben, and Gomes Gabe
A. Boiko Daniil, MacKnight Robert, Kline Ben, and Gomes Gabe. Autonomous chemical research with large language models. Nature, pages 570–578, 2023. 37
2023
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.