{"paper":{"title":"Benchmarking and Evaluation of AI Models in Biology: Outcomes and Recommendations from the CZI Virtual Cells Workshop","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Aditi Krishnapriyan, Alma Andersson, Aly Khan, Angela Oliveira Pisco, Ankit Gupta, Anthony Gitter, Benjamin Chang, Bo Wang, Daniel Burkhardt, Elana Simon, Elizabeth Fahsbender, Emma Lundberg, Genevieve Haliburton, Georg K. Gerber, Gustavo Stolovitzky, Ivana Jelic, James Zou, Jeremy Ash, Jon M. Laurent, Julio Saez-Rodriguez, Katherine S. Pollard, Katrina Kalantar, Marc Valer, Patrick Godau, Polina Binder, Rob Moccia, Shalin B. Mehta, Siyu He, Srinivasan Sivanandan, Suresh Ramani, Tianyu Liu, Trey Ideker, Xikun Zhang, Yang-Joon Kim, Yasin Senbabaoglu","submitted_at":"2025-07-14T17:25:28Z","abstract_excerpt":"Artificial intelligence holds immense promise for transforming biology, yet a lack of standardized, cross domain, benchmarks undermines our ability to build robust, trustworthy models. Here, we present insights from a recent workshop that convened machine learning and computational biology experts across imaging, transcriptomics, proteomics, and genomics to tackle this gap. We identify major technical and systemic bottlenecks such as data heterogeneity and noise, reproducibility challenges, biases, and the fragmented ecosystem of publicly available resources and propose a set of recommendation"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2507.10502","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2507.10502/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}