A new benchmarking framework shows virtual cell models overestimate performance on standard tests, drop sharply on unseen contexts and perturbations, and produce inconsistent rankings across metrics.
Simple controls exceed best deep learning algorithms and reveal foundation model effectiveness for predicting genetic perturbations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-bio.CB 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Benchmarking virtual cell models for in-the-wild perturbation response
A new benchmarking framework shows virtual cell models overestimate performance on standard tests, drop sharply on unseen contexts and perturbations, and produce inconsistent rankings across metrics.