REVIEW 3 major objections 6 minor 52 references
JUMP-lite: Compact, reproducible benchmarking of cell representations
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read JPEG XL compression at high and medium quality preserves biological signal in Cell Painting images, so a 116 GB subset of the 115 TB JUMP dataset can stand in for the full resource when benchmarking cell representations.
desk verdict The compression-fidelity result is solid and the dataset is a real contribution, but the headline model ranking is confounded by a site-count mismatch between CellProfiler and every deep model. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a compression-fidelity ladder built on JPEG XL codec settings (near-lossless HQ, medium MQ, and aggressive d20), combined with a standardized evaluation harness. Ground truth comes from the intersection of RefChem records and MOTIVE drug-target graphs, and every representation is scored on phenotypic activity, phenotypic consistency, and cross-modality recall, normalized to NAP. The harness uses an orchestration pipeline for segmentation and featurization, and an inter-process bridge that runs each deep-learning model in an isolated, reproducible environment so all methods are processed identically.
What would settle it
Recompute the classical CellProfiler profiles from the same four imaging sites per well used for deep-learning embeddings and rerun the phenotypic activity and consistency tasks; if CellProfiler no longer ranks with MorphEM at the top, the three-tier ordering is an artifact of site count rather than of representation quality.
Extended reading notes
Core claim
On its own terms, the paper's central discovery is that moderate JPEG XL lossy compression preserves enough biological signal in Cell Painting images to support fair benchmarking of representation methods. Across the full JUMP-lite dataset, high-quality compression changes mean per-task performance by -1.2% and stays within ±6% on every task, medium-quality compression costs -6.5% on average, and the relative ordering of models is preserved at both settings. Only the aggressive d20 setting destroys signal, with phenotypic activity dropping 20-34% and phenotypic consistency up to 31%. Using this evaluation, the paper reports three performance tiers among representations: CellProfiler features and MorphEM lead, DINOv2, SubCell, and OpenPhenom sit in the middle, and cell-count and random-weight ViT baselines trail.
Load-bearing premise
The paper's comparison of representation methods assumes that the classical feature pipeline, which uses 6-9 imaging sites per well, and the deep learning models, which use 4 sites per well, are directly comparable.
Editorial extensions
If this is right
- A researcher with a modest server can run the full JUMP-lite benchmark at MQ compression and expect model rankings to match those from the uncompressed 115 TB dataset.
- The ~97–99% storage reduction at HQ and MQ comes with average per-task performance changes of -1.2% and -6.5%, so conclusions about which representations are strong or weak do not hinge on using raw images.
- Because aggressive d20 compression degrades PA by 20–34% and PC by up to 31%, the paper establishes a usable compression range and identifies settings that should be avoided.
- CellProfiler features and MorphEM lead the benchmark, meaning classical engineered features remain a competitive baseline for image-based profiling, despite being roughly 200x slower than deep-learning pipelines.
Reading between the lines
- If the compression-fidelity trade-off transfers to other fluorescence and bright-field assays, the same HQ/MQ recipe could shrink other large microscopy repositories, making them benchmarkable on commodity hardware.
- The observation that some models score higher at intermediate compression than at raw hints that lossy compression may act as a mild denoiser, suppressing inter-site resolution noise; testing this directly would clarify when compression improves rather than degrades profiles.
- Releasing JUMP-lite at four compression levels lets future work study how compression interacts with specific cell types, assays, and annotation schemes, and could guide compression choices for specialized benchmarks.
- The same harness should make it straightforward to add new representation methods and re-run the entire benchmark without re-tuning the evaluation, which would accelerate comparisons as new foundation models appear.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces JUMP-lite, a 116 GB curated and JPEG XL-compressed subset of the 115 TB JUMP Cell Painting dataset that retains roughly 24,401 perturbations, and Nahual, a Nix-based framework for deploying models with incompatible dependencies. The authors characterize the effects of lossy compression on segmentation, feature correlation, and downstream retrieval tasks (phenotypic activity, phenotypic consistency, and MOTIVE cross-modality recall), reporting that HQ compression yields a 1.2% mean performance change and MQ a 6.5% change while reducing storage by 97–98.8%. They benchmark five representation methods (CellProfiler, MorphEM, OpenPhenom, SubCell, DINOv2) plus simple baselines across eleven tasks and find that CellProfiler and MorphEM form a top tier, DINOv2/SubCell/OpenPhenom a middle tier, and that model ranking is stable across HQ/MQ compression.
Significance. If the central claims hold, JUMP-lite would be a valuable community resource: a roughly 1000-fold storage reduction with a multi-lab, multi-task evaluation suite, combined with Nahual's reproducibility infrastructure, could make systematic benchmarking of cell representations accessible to groups without petabyte-scale storage. The compression-fidelity analysis fills a real gap in the Cell Painting literature, where previous compact benchmarks such as RxRx3-core did not assess the impact of compression on downstream signal. The finding that engineered CellProfiler features remain competitive with modern self-supervised models is an important and falsifiable result, and the authors should be credited for their public code release, detailed curation documentation, and transparent reporting of the normalization sweep. However, the head-to-head model ranking is currently confounded by a site-count mismatch, and the selection of post-processing configurations on the evaluation metric leaves the absolute performance claims open to optimistic bias; these issues must be resolved before the benchmark's headline conclusions are reliable.
major comments (3)
- [3.5, Figure 5, S1.2.7] The headline ranking of representation methods is confounded by the input sampling protocol. Section S1.2.7 states that the precomputed CellProfiler features were derived from six to nine imaging sites per well, whereas all deep-learning embeddings in JUMP-lite are computed from exactly four sites per well. Because well-level aggregation over more imaging sites reduces noise, CellProfiler's top-tier NAP values (e.g., 0.815 vs. 0.777 on CRISPR PA in Figure 5) may reflect data quantity rather than representation quality. Moreover, the released JUMP-lite images contain four sites per well, so a user cannot reproduce the CellProfiler row from the released data. The acknowledgment in S1.2.7 does not justify the unadjusted comparison in Figure 5; the authors should either recompute CellProfiler features on the same four sites per well (or a matched subsample) or explicitly restrict the ranking claim to deep-learning models.
- [3.1, S1.2.6, Figure 5 caption] The post-processing protocol is internally inconsistent and, under one reading, selects configurations on the evaluation metric. Section 3.1 states 'We report the configuration with the highest balanced PA/PC (rescaled variant)' and S1.2.6 says 'The post-processed profiles from the best-performing configuration were then used for downstream analyses,' but the Figure 5 caption reports 'the average across normalization configurations.' These are different quantities: the former is a maximum over 280–420 configurations, which optimistically biases all reported scores; the latter is a mean and gives some protection against overfitting. The authors must specify which quantity is used for each reported result, and for the best-configuration variant they should provide unbiased estimates (e.g., evaluation on a held-out split or at a fixed configuration). This is load-bearing for the quantitative claims in Table 1 and the tier assignment in Figure 5.
- [Figure 5, Table S5] The three-tier ranking is presented without uncertainty quantification. Figure 5 shows only the mean normalized score per model, and Table S5 reports standard deviations across normalization configurations in percentage-change units; for example, the middle-tier scores (DINOv2 0.721, SubCell 0.715, OpenPhenom 0.688) are close, and the per-configuration spread in Table S5 is comparable to these differences. Without error bars or paired statistical tests on the normalized scores, the claim that these three models form a distinct middle tier is not supported. The authors should report the distribution across configurations for the normalized scores in Figure 5, or provide a statistical test of the ranking.
minor comments (6)
- [Figure 3c caption] The caption states that the pooled distribution includes '48 normalization configurations each,' but S1.2.6 describes 280 configurations for engineered features and up to 420 for deep-learning embeddings; please reconcile these numbers.
- [3.1, Table S2, Figure 2a] The main text says 'We evaluated four compression levels' (raw, HQ, MQ, d20), but Table S2 and Figure 2a present additional levels (zstd, jxl-effort-3, jxl-d2-e8, jxl-lq, jxl-d10, jxl-d15, jxl-d25); please clarify which levels are used in which analyses and why the additional levels are omitted from the main narrative.
- [S1.4.1, Data availability] S1.4.1 states that links to software, code, and repositories are unavailable due to anonymization, while the Data availability section provides GitHub and Zenodo URLs; in the final version these statements should be reconciled.
- [Figure 5, Methods] The 'Cell Count' baseline used in Figure 5 is not defined anywhere in the Methods or Supplementary Information; please specify how cell count is computed, aggregated to the well level, and normalized before being used as a profile.
- [3.4, 3.5, S1.3.2] The main text does not state whether the reported MOTIVE results use the 'full' or 'strict' annotation set described in S1.3.2; because the composition of the edge sets differ substantially (e.g., the full CC graph includes chemical-similarity RESEMBLES edges), this choice should be stated explicitly wherever MOTIVE recall values are reported.
- [Discussion] The sentence 'appreciable degradation appearing only upon aggressive compression (d20)' is inconsistent with the MQ results in Table 1, where some tasks lose up to 15.4% (PC on Diverse); please qualify the claim to reflect the task-dependent MQ degradation.
Circularity Check
No circularity: JUMP-lite is an empirical benchmark whose compression, rank-stability, and model-tier claims are measured on external retrieval tasks, not derived from the dataset construction or from self-citations.
full rationale
JUMP-lite is an empirical benchmark paper, not a derivational one, and I found no step where a claimed result reduces to its own input by construction. The central claims—that HQ/MQ JPEG XL compression preserves biological signal, that model ranks are stable across compression, and that representations fall into three tiers—are all measurements on retrieval tasks, not consequences of the dataset curation. The curation step (Section 3.1) does select compounds present in both RefChem and MOTIVE, and the same annotation sources later define PC and recall@K tasks (Section S1.3); this is a design overlap, but the tasks are not satisfied by construction: PA compares replicates against plate negative controls, PC compares same-target versus different-target perturbations, and MOTIVE recall@K requires ranking image-derived profiles across modalities. No equation-level equivalence, no fitted parameter renamed as a prediction, and no uniqueness theorem imported from the authors appears in the paper. Self-citations to copairs, MOTIVE, ALIBY, and cp measure are references to published tools and metrics whose validity is external to this work; they do not by themselves force any reported performance delta. The acknowledged confound that precomputed CellProfiler features come from 6-9 imaging sites per well while deep-learning embeddings use four sites (S1.2.7) is a comparability limitation and a correctness risk for the Figure 5 ranking, but it is not circularity: the CellProfiler scores are still measured against external replicates and annotations. The paper is self-contained as an empirical benchmark, so the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (3)
- JPEG XL compression levels =
jxl-hq, jxl-mq, jxl-d20-e2
- RefChem support threshold =
>1 independent assay record
- Normalization configuration =
best balanced PA/PC (rescaled) per model-codec
assumptions (4)
- domain assumption Retrieval of known drug-target annotations from morphology is a valid proxy for representation quality.
- domain assumption Negative controls (DMSO) provide a valid per-plate baseline for normalization.
- domain assumption The Target-2 pilot subset (four plates) is representative of compression effects across the full JUMP-lite dataset.
- domain assumption RefChem and MOTIVE annotations are ground truth for target relationships.
Cite this review
Pith. "Pith review of JUMP-lite: Compact, reproducible benchmarking of cell representations." pith.science (2026). https://pith.science/paper/WOIL77FY
@misc{pith2026260807632,
author = {Pith},
title = {Pith review of: JUMP-lite: Compact, reproducible benchmarking of cell representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/WOIL77FY}},
note = {Machine review of arXiv:2608.07632}
}
read the original abstract
Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting now provide millions of images for systematic study. However, the scale of these resources, 115 TB for JUMP alone, and fragmented evaluation practices make systematic comparison of representation methods intractable for many researchers. Here we present Nahual, an open-source framework for reproducible model deployment, and JUMP-lite, a curated 116 GB subset of JUMP that is 1000 times smaller while preserving phenotypic diversity through careful selection of perturbations with high-confidence annotations and a storage reduction via lossy JPEG XL compression. With these, we benchmark five representation methods, including classical features (CellProfiler) and deep learning models (MorphEM, OpenPhenom, SubCell, DINOv2), and demonstrate that compression preserves downstream signal while standardized phenotypic activity and consistency metrics reveal meaningful performance differences across methods. Together, JUMP-lite and Nahual provide a foundation for accessible, reproducible benchmarking of image-based cell representations.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Vidit Agrawal, John Peters, Tyler N. Thompson, Moham- mad Vali Sanian, Chau Pham, Nikita Moshkov, Arshad Kazi, Aditya Pillai, Jack Freeman, Byunguk Kang, Samouil L. Farhi, Ernest Fraenkel, Ron Stewart, Lassi Paavolainen, Bryan A. Plummer, and Juan C. Caicedo. CHAMMI-75: Pre-Training 7 Figure 5.Model performance per task, normalized by per-task maximum (co...
work page 2025
-
[2]
John Arevalo, Ellen Su, Anne E. Carpenter, and Shantanu Singh. MOTIVE: A Drug-Target Interaction Graph For Induc- tive Link Prediction.Advances in Neural Information Process- ing Systems, 37:140320–140333, 2024
work page 2024
-
[3]
Nicolas Bourriez, Ihab Bendidi, Ethan Cohen, Gabriel Watkin- son, Maxime Sanchez, Guillaume Bollot, and Auguste Gen- ovesio. ChAda-ViT : Channel Adaptive Attention for Joint Rep- resentation Learning of Heterogeneous Microscopy Images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11556–11565, 2024
work page 2024
-
[4]
Davis, Blake Borgeson, Cathy L
Mark-Anthony Bray, Shantanu Singh, Han Han, Chadwick T. Davis, Blake Borgeson, Cathy L. Hartland, Maria Kost- Alimova, Sigrun Gustafsdottir, Christopher C. Gibson, and Anne E. Carpenter. Cell Painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes.Nature Protocols, 11(9):1757–1774, 2016
work page 2016
-
[5]
Data-Analysis Strategies for Image-Based Cell Profiling.Nature Methods, 14(9):849–863, 2017
Juan C Caicedo, Sam Cooper, Florian Heigwer, Scott War- chal, Peng Qiu, Csaba Molnar, Aliaksei S Vasilevich, Joseph D Barry, Harmanjit Singh Bansal, Oren Kraus, Mathias Wawer, Lassi Paavolainen, Markus D Herrmann, Mohammad Rohban, Jane Hung, Holger Hennig, John Concannon, Ian Smith, Paul A Clemons, Shantanu Singh, Paul Rees, Peter Horvath, Roger G Liningt...
work page 2017
-
[6]
Caicedo, Allen Goodman, Kyle W
Juan C. Caicedo, Allen Goodman, Kyle W. Karhohs, Beth A. Cimini, Jeanelle Ackerman, Marzieh Haghighi, Cher-Keng Heng, Tim Becker, Minh Doan, Claire McQuin, Csaba Mora ˇn, Anne E. Carpenter, and Shantanu Singh. Nucleus segmenta- tion across imaging experiments: The 2018 Data Science Bowl. Nature Methods, 16(12):1247–1253, 2019
work page 2018
-
[7]
Peter D Caie, Rebecca E Walls, Alexandra Ingleston-Orme, Sandeep Daya, Tom Houslay, Rob Eagle, Mark E Roberts, and Neil O Carragher. High-content phenotypic profiling of drug re- sponse signatures across distinct cancer cells.Molecular cancer therapeutics, 9(6):1913–1926, 2010
work page 1913
-
[8]
Franck Cappello, Allison Baker, Ebru Bozda, Martin Burtscher, Kyle Chard, Sheng Di, Paul Christopher O Grady, Peng Jiang, Shaomeng Li, Erik Lindahl, et al. Lossy compression of sci- entific data: Applications constrains and requirements.arXiv preprint arXiv:2503.20031, 2025
arXiv 2025
Show all 52 references
-
[9]
Building, benchmarking, and exploring perturbative maps of transcriptional and morphological data.PLoS compu- tational biology, 20(10):e1012463, 2024
Safiye Celik, Jan-Christian H ¨utter, Sandra Melo Carlos, Nathan H Lazar, Rahul Mohan, Conor Tillinghast, Tommaso Biancalani, Marta M Fay, Berton A Earnshaw, and Imran S Haque. Building, benchmarking, and exploring perturbative maps of transcriptional and morphological data.PL...
2024
-
[10]
Lazar, Rahul Mohan, Conor Tillinghast, Tommaso 8 Biancalani, Marta M
Safiye Celik, Jan-Christian H ¨utter, Sandra Melo Carlos, Nathan H. Lazar, Rahul Mohan, Conor Tillinghast, Tommaso 8 Biancalani, Marta M. Fay, Berton A. Earnshaw, and Imran S. Haque. Building, benchmarking, and exploring perturbative maps of transcriptional and morphological d...
2024
-
[11]
Boyd, and Anne E
Srinivas Niranj Chandrasekaran, Hugo Ceulemans, Justin D. Boyd, and Anne E. Carpenter. Image-based profiling for drug discovery: Due for a machine-learning upgrade?Nature Re- views Drug Discovery, 20:145–159, 2020
2020
-
[12]
Michael Ando, John Arevalo, Melissa Bennion, Nicolas Boisseau, Adriana Borowa, Justin D
Srinivas Niranj Chandrasekaran, Jeanelle Ackerman, Eric Alix, D. Michael Ando, John Arevalo, Melissa Bennion, Nicolas Boisseau, Adriana Borowa, Justin D. Boyd, Laurent Brino, Patrick J. Byrne, Hugo Ceulemans, Carolyn Ch’ng, Beth A. Cimini, Djork-Arne Clevert, Nicole Deflaux, J...
2023
-
[13]
Three million images and morphological profiles of cells treated with matched chemical and genetic perturbations.Na- ture Methods, 21(6):1114–1121, 2024
Srinivas Niranj Chandrasekaran, Beth A Cimini, Amy Goodale, Lisa Miller, Maria Kost-Alimova, Nasim Jamali, John G Doench, Briana Fritchman, Adam Skepner, Michelle Melanson, et al. Three million images and morphological profiles of cells treated with matched chemical and geneti...
2024
-
[14]
Byrne, William G
Srinivas Niranj Chandrasekaran, Eric Alix, John Arevalo, Adri- ana Borowa, Patrick J. Byrne, William G. Charles, Zitong S. Chen, Beth A. Cimini, Boxiong Deng, John G. Doench, Jes- sica D. Ewald, Briana Fritchman, Colin J. Fuller, Jedidiah Gaetz, Amy Goodale, Marzieh Haghighi, ...
2025
-
[15]
Plummer, and Juan C
Zitong Chen, Chau Pham, Siqi Wang, Michael Doron, Nikita Moshkov, Bryan A. Plummer, and Juan C. Caicedo. CHAMMI: A benchmark for channel-adaptive models in microscopy imag- ing. InAdvances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track, 2023
2023
-
[16]
Cimini, Srinivas Niranj Chandrasekaran, Maria Kost- Alimova, Lisa Miller, Amy Goodale, Briana Fritchman, Patrick J
Beth A. Cimini, Srinivas Niranj Chandrasekaran, Maria Kost- Alimova, Lisa Miller, Amy Goodale, Briana Fritchman, Patrick J. Byrne, Sakshi Garg, Nasim Jamali, David J. Lo- gan, John Concannon, Charles-Hugues Lardeau, Elizabeth Mouchet, Shantanu Singh, Hamdah Shafqat Abbasi, Pet...
1981
-
[17]
The drug re- purposing hub: a next-generation drug library and information resource.Nature medicine, 23(4):405–408, 2017
Steven M Corsello, Joshua A Bittker, Zihan Liu, Joshua Gould, Patrick McCarren, Jodi E Hirschman, Stephen E Johnston, Anita Vrcic, Bang Wong, Mariya Khan, et al. The drug re- purposing hub: a next-generation drug library and information resource.Nature medicine, 23(4):405–408, 2017
2017
-
[18]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In2009 IEEE conference on computer vision and pattern recog- nition, pages 248–255. Ieee, 2009
2009
-
[19]
Nix Based Fully Automated Workflows and Ecosystem to Guaran- tee Scientific Result Reproducibility across Software Environ- ments and Systems
Adrien Devresse, Fabien Delalondre, and Felix Sch¨urmann. Nix Based Fully Automated Workflows and Ecosystem to Guaran- tee Scientific Result Reproducibility across Software Environ- ments and Systems. InProceedings of the 3rd International Workshop on Software Engineering for ...
2015
-
[20]
Nix: A Safe and Policy-Free System for Software Deployment
Eelco Dolstra, Merijn De Jonge, Eelco Visser, et al. Nix: A Safe and Policy-Free System for Software Deployment. InLISA, pages 79–92, 2004
2004
-
[21]
Fay, Oren Kraus, Mason Victors, Lakshmanan Aru- mugam, Kamal Vuggumudi, John Urbanik, Kyle Hansen, Safiye Celik, Nico Cernek, Ganesh Jagannathan, Jordan Christensen, Berton A
Marta M. Fay, Oren Kraus, Mason Victors, Lakshmanan Aru- mugam, Kamal Vuggumudi, John Urbanik, Kyle Hansen, Safiye Celik, Nico Cernek, Ganesh Jagannathan, Jordan Christensen, Berton A. Earnshaw, Imran S. Haque, and Ben Mabey. RxRx3: Phenomics Map of Biology, 2023
2023
-
[22]
Integration of the Drug–Gene Interaction Database (DGIdb 4.0) with open crowd- source efforts.Nucleic Acids Research, 49(D1):D1144–D1151, 2021
Sharon L Freshour, Susanna Kiwala, Kelsy C Cotto, Adam C Coffman, Joshua F McMichael, Jonathan J Song, Malachi Grif- fith, Obi L Griffith, and Alex H Wagner. Integration of the Drug–Gene Interaction Database (DGIdb 4.0) with open crowd- source efforts.Nucleic Acids Research, 4...
2021
-
[23]
Subcell: Proteome- aware vision foundation models for microscopy capture single- cell biology.bioRxiv, pages 2024–12, 2025
Ankit Gupta, Zoe Wefers, Konstantin Kahnert, Jan N Hansen, Mohini K Misra, Will Leineweber, Anthony Cesnik, Dan Lu, Ulrika Axelsson, Frederic Ballllosera, et al. Subcell: Proteome- aware vision foundation models for microscopy capture single- cell biology.bioRxiv, pages 2024–12, 2025
2024
-
[24]
Sokolnicki, Joshua Wilson, Deepika Walpita, Melissa M
Sigrun Gustafsdottir, Vebjorn Ljosa, Katherine L. Sokolnicki, Joshua Wilson, Deepika Walpita, Melissa M. Kemp, Kath- leen Petri Seiler, Hyman A. Carrel, Todd R. Golub, Stuart L. Schreiber, Paul A. Clemons, Anne E. Carpenter, and Alykhan F. Shamji. Multiplex Cytological Profili...
2013
-
[25]
Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network, 2025
Pin-Jui Huang, Yu-Hsuan Liao, SooHeon Kim, NoSeong Park, JongBae Park, and DongMyung Shin. Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network, 2025
2025
-
[26]
Workflow for defining reference chemicals for assessing performance of in vitro assays.Altex, 36(2):261, 2018
Richard S Judson, Russell S Thomas, Nancy Baker, Anita Simha, Xia Meng Howey, Carmen Marable, Nicole C Klein- streuer, and Keith A Houck. Workflow for defining reference chemicals for assessing performance of in vitro assays.Altex, 36(2):261, 2018
2018
-
[27]
Kalinin, John Arevalo, Erik Serrano, Loan Vul- liard, Hillary Tsang, Michael Bornholdt, Al ´an F
Alexandr A. Kalinin, John Arevalo, Erik Serrano, Loan Vul- liard, Hillary Tsang, Michael Bornholdt, Al ´an F. Mu˜noz, Sug- anya Sivagurunathan, Bartek Rajwa, Anne E. Carpenter, Gre- gory P. Way, and Shantanu Singh. A Versatile Information Re- trieval Framework for Evaluating P...
2025
-
[28]
Cheveralls, Manuel D
Hirofumi Kobayashi, Keith C. Cheveralls, Manuel D. Leonetti, and Loic A. Royer. Self-Supervised Deep Learning Encodes High-Resolution Features of Protein Subcellular Localization. Nature Methods, 19(8):995–1003, 2022
2022
-
[29]
The heterogeneous pharmacological medical biochemical net- work PharMeBINet.Scientific Data, 9(1):393, 2022
Cassandra K ¨onigs, Marcel Friedrichs, and Theresa Dietrich. The heterogeneous pharmacological medical biochemical net- work PharMeBINet.Scientific Data, 9(1):393, 2022
2022
-
[30]
Oren Kraus, Federico Comitani, John Urbanik, Kian Kenyon- Dean, Lakshmanan Arumugam, Saber Saberian, Cas Wognum, Safiye Celik, and Imran S. Haque. RxRx3-core: Benchmarking drug-target interactions in High-Content Microscopy, 2025
2025
-
[31]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. InEuro- pean conference on computer vision, pages 740–755. Springer, 2014
2014
-
[32]
Mu ˜noz.Phenotyping Single Cells of Saccharomyces Cerevisiae Using an End-to-End Analysis of High-Content Time-Lapse Microscopy
Al ´an F. Mu ˜noz.Phenotyping Single Cells of Saccharomyces Cerevisiae Using an End-to-End Analysis of High-Content Time-Lapse Microscopy. PhD thesis, University of Edinburgh, 2023
2023
-
[33]
Mu ˜noz, Tim Treis, Alexandr A
Al ´an F. Mu ˜noz, Tim Treis, Alexandr A. Kalinin, Shatavisha Dasgupta, Fabian Theis, Anne E. Carpenter, and Shantanu Singh. Cp measure: API-first Feature Extraction for Image- Based Profiling Workflows, 2025
2025
-
[34]
DI- NOv2: Learning Robust Visual Features without Supervision, 2024
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haz- iza, Francisco Massa, Alaaeldin El-Nouby, Mahmoud Ass- ran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Mich...
2024
-
[35]
Analysis of the human pro- tein atlas image classification competition.Nature methods, 16 (12):1254–1261, 2019
Wei Ouyang, Casper F Winsnes, Martin Hjelmare, Anthony J Cesnik, Lovisa ˚Akesson, Hao Xu, Devin P Sullivan, Shubin Dai, Jun Lan, Park Jinmo, et al. Analysis of the human pro- tein atlas image classification competition.Nature methods, 16 (12):1254–1261, 2019
2019
-
[36]
BioImage Model Zoo: A Community-Driven Resource for Ac- cessible Deep Learning in BioImage Analysis, 2022
Wei Ouyang, Fynn Beuttenmueller, Estibaliz G ´omez-de- Mariscal, Constantin Pape, Tom Burke, Carlos Garcia-L ´opez- de-Haro, Craig Russell, Luc´ıa Moya-Sans, Cristina de-la-Torre- Guti´errez, Deborah Schmidt, Dominik Kutra, Maksim Novikov, Martin Weigert, Uwe Schmidt, Peter Ba...
2022
-
[37]
Cellpose-sam: superhuman generalization for cellular segmen- tation.BioRxiv, pages 2025–04, 2025
Marius Pachitariu, Michael Rariden, and Carsen Stringer. Cellpose-sam: superhuman generalization for cellular segmen- tation.BioRxiv, pages 2025–04, 2025
2025
-
[38]
Self-Supervised Vision Transformers Accurately De- code Cellular State Heterogeneity, 2023
Ramon Pfaendler, Jacob Hanimann, Sohyon Lee, and Berend Snijder. Self-Supervised Vision Transformers Accurately De- code Cellular State Heterogeneity, 2023
2023
-
[39]
Mapping information-rich genotype-phenotype landscapes with genome-scale perturb-seq.Cell, 185(14):2559– 2575, 2022
Joseph M Replogle, Reuben A Saunders, Angela N Pogson, Jeffrey A Hussmann, Alexander Lenail, Alina Guna, Lauren Mascibroda, Eric J Wagner, Karen Adelman, Gila Lithwick- Yanai, et al. Mapping information-rich genotype-phenotype landscapes with genome-scale perturb-seq.Cell, 185...
2022
-
[40]
Carpenter
Srijit Seal, Maria-Anna Trapotsi, Ola Spjuth, Shantanu Singh, Jordi Carreras-Puigvert, Nigel Greene, Andreas Bender, and Anne E. Carpenter. Cell Painting: A Decade of Discovery and Innovation in Cellular Imaging.Nature Methods, 22(2):254– 268, 2025
2025
-
[41]
Cell- Profiler 4: Improvements in speed, utility and usability.BMC bioinformatics, 22(1):433, 2021
David R Stirling, Madison J Swain-Bowden, Alice M Lucas, Anne E Carpenter, Beth A Cimini, and Allen Goodman. Cell- Profiler 4: Improvements in speed, utility and usability.BMC bioinformatics, 22(1):433, 2021
2021
-
[42]
Cellpose: A Generalist Algorithm for Cellular Seg- mentation.Nature Methods, 18(1):100–106, 2021
Carsen Stringer, Tim Wang, Michalis Michaelos, and Marius Pachitariu. Cellpose: A Generalist Algorithm for Cellular Seg- mentation.Nature Methods, 18(1):100–106, 2021
2021
-
[43]
RxRx1: A Dataset for Evaluating Ex- perimental Batch Correction Methods
Maciej Sypetkowski, Morteza Rezanejad, Saber Saberian, Oren Kraus, John Urbanik, James Taylor, Ben Mabey, Mason Victors, Jason Yosinski, Alborz Rezazadeh Sereshkeh, Imran Haque, and Berton Earnshaw. RxRx1: A Dataset for Evaluating Ex- perimental Batch Correction Methods. InPro...
2023
-
[44]
Rxrx1: A dataset for evaluating experimental batch correction methods
Maciej Sypetkowski, Morteza Rezanejad, Saber Saberian, Oren Kraus, John Urbanik, James Taylor, Ben Mabey, Mason Victors, Jason Yosinski, Alborz Rezazadeh Sereshkeh, et al. Rxrx1: A dataset for evaluating experimental batch correction methods. InProceedings of the IEEE/CVF conf...
2023
-
[45]
Understanding the effects of modern compressors on the community earth science model
Robert Underwood, Julie Bessac, Sheng Di, and Franck Cap- pello. Understanding the effects of modern compressors on the community earth science model. In2022 IEEE/ACM 8th In- ternational Workshop on Data Analysis and Reduction for Big Scientific Data (DRBSD), pages 1–10. IEEE, 2022
-
[46]
Way, Heba Sailem, Steven Shave, Richard Kasprow- icz, and Neil O
Gregory P. Way, Heba Sailem, Steven Shave, Richard Kasprow- icz, and Neil O. Carragher. Evolution and impact of high con- tent imaging.SLAS Discovery, 28(7):292–305, 2023
2023
-
[47]
Carpenter, Beth A
Erin Weisbart, Ankur Kumar, John Arevalo, Anne E. Carpenter, Beth A. Cimini, and Shantanu Singh. Cell Painting Gallery: An Open Resource for Image-Based Profiling.Nature Methods, 21 (10):1775–1777, 2024
2024
-
[48]
Image data resource: a bioimage data integration and publication platform
Eleanor Williams, Josh Moore, Simon W Li, Gabriella Rustici, Aleksandra Tarkowska, Anatole Chessel, Simone Leo, B ´alint Antal, Richard K Ferguson, Ugis Sarkans, et al. Image data resource: a bioimage data integration and publication platform. Nature methods, 14(8):775–781, 2017
2017
-
[49]
Morphological profiling dataset of eu- openscreen bioactive compounds over multiple imaging sites and cell lines.bioRxiv, pages 2024–08, 2024
Christopher Wolff, Martin Neuenschwander, Carsten J ¨orn Beese, Divya Sitani, Maria C Ramos, Alzbeta Srovnalova, Mar´ıa Jos ´e Varela, Pavel Polishchuk, Katholiki E Skopelitou, Ctibor ˇSkuta, et al. Morphological profiling dataset of eu- openscreen bioactive compounds over mul...
2024
-
[50]
Ubas, Richard de Borja, Valentine Svens- son, Nicole Thomas, Neha Thakar, Ian Lai, Aidan Winters, 10 Umair Khan, Matthew G
Jesse Zhang, Airol A. Ubas, Richard de Borja, Valentine Svens- son, Nicole Thomas, Neha Thakar, Ian Lai, Aidan Winters, 10 Umair Khan, Matthew G. Jones, John D. Thompson, Vuong Tran, Joseph Pangallo, Efthymia Papalexi, Ajay Sapre, Hoai Nguyen, Oliver Sanderson, Maria Nigos, Ol...
2025
-
[51]
Sdrbench: Scien- tific data reduction benchmark for lossy compressors
Kai Zhao, Sheng Di, Xin Lian, Sihuan Li, Dingwen Tao, Julie Bessac, Zizhong Chen, and Franck Cappello. Sdrbench: Scien- tific data reduction benchmark for lossy compressors. In2020 IEEE international conference on big data (Big Data), pages 2716–2724. IEEE, 2020
2020
-
[52]
Deep-Learning- Based Image Compression for Microscopy Images: An Empir- ical Study.Biological Imaging, 4:e16, 2024
Yu Zhou, Jan Sollmann, and Jianxu Chen. Deep-Learning- Based Image Compression for Microscopy Images: An Empir- ical Study.Biological Imaging, 4:e16, 2024. S1 Supplementary Material S1.1 Additional analyses Using the Target-2 subset of JUMP-lite we explored a wider array of si...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.