REVIEW 3 major objections 5 minor 15 references
Reply to "Issues arising from benchmarking single-cell RNA sequencing imputation methods"
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read SAVER's benchmark ranking survives a revised downsampling test.
desk verdict A coherent reply that fixes the univariate mismatch in the downsampling and flags scImpute's unfair advantage, but the central claim of agreement is asserted without the numbers needed to verify it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the amended downsampling protocol: random subsetting of genes and cells from a real dataset, calculation of the subset's mean expression and library size, and then generation of noisy counts $Y_{gc} \sim \mathrm{Poisson}(\tau_c \lambda_{gc})$ with $\tau_c$ drawn from a Gamma distribution matched to the selected library size. This protocol is what lets the authors claim that the synthetic test data share the first-order statistics of real data while still serving as a known ground truth. The evaluation then uses Pearson correlation with the reference, distance between gene-gene and cell-cell correlation matrices, and Seurat-based clustering compared by Jaccard index. The reply's rebuttal also rests on the choice of clustering method and on whether the number of clusters is provided to a method in advance.
What would settle it
Construct amended downsampled versions of the four datasets while preserving the empirical gene-gene covariance, or use technical replicate pairs as ground truth, then rerun SAVER, MAGIC, and scImpute: if the SAVER advantage in correlation with the reference or in clustering agreement disappears or reverses, the reply's validation claim is falsified.
Extended reading notes
Core claim
The paper claims that the earlier conclusion, that SAVER recovers reference expression levels more accurately than MAGIC and scImpute as measured by gene-level and cell-level correlation, correlation-matrix distance, and clustering agreement, remains true when the downsampling simulation is corrected to better match real data. The authors construct amended downsampled datasets by picking random gene and cell subsets from each original dataset, computing the mean expression and library size of that subset, and sampling noisy observations from a Poisson-Gamma noise model tuned to those quantities. The resulting datasets have nearly identical mean, standard deviation, and zero fraction per gene as the original data, unlike the first version. Running the same three methods and metrics on these new datasets gives results that the paper states are in agreement with the findings in its Brief Communication.
Load-bearing premise
The argument collapses if matching a simulated dataset's mean expression and library size to real data does not also reproduce the gene-gene correlations, dropout structure, and batch effects that real scRNA-seq noise has.
Editorial extensions
If this is right
- If the amended experiment is a fair test, the original ranking of SAVER over MAGIC and scImpute on the four datasets is stable under a more realistic noise model.
- The critique's objections about the old downsampled data, such as zero-fraction mismatch and low per-gene agreement with the original data, do not invalidate the benchmark because those properties are expected consequences of adding noise.
- Evaluations that compare imputed output to labels derived from the same noisy data, or that use an unstable clustering method, are not reliable ways to judge imputation.
- Giving scImpute the true number of clusters while other methods are blind to it makes the comparison unfair to the other methods.
Reading between the lines
- The amended protocol matches mean expression and library size but does not explicitly preserve gene-gene correlations or dropout patterns, so a stricter test would generate noise with the empirical covariance or use technical replicate pairs.
- The exchange suggests that benchmarking conclusions in single-cell imputation can hinge on protocol choices such as the clustering algorithm and whether simulated noise preserves higher-order structure, pointing to a need for standardized noise-generation benchmarks.
- One testable extension is to apply the amended downsampling protocol to additional imputation methods or to datasets with known spike-in controls and check whether the relative ranking of SAVER, MAGIC, and scImpute remains unchanged.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This short reply addresses Li & Li's comment on the downsampling experiment in the authors' SAVER Brief Communication. The authors first correct a misunderstanding: their downsampling procedure uses a Poisson noise model with a Gamma-distributed cell-specific sequencing efficiency, not the Poisson-Gamma prior assumed by SAVER. They then acknowledge that the original downsampled data had marginal distributions that deviate from real data, and they propose an amended procedure that randomly selects a subset of genes and cells from the original dataset and downsamples the reference data to match the selected mean expression and library size. They report that the new downsampled data match the original data in mean expression, standard deviation, and zero fraction (Figure 1), and they state that re-running SAVER, MAGIC, and scImpute on these new datasets reproduces the findings of the original Brief Communication (Figure 2). Finally, they argue that Li & Li's real-data evaluation using hierarchical clustering is flawed because imputation should recover true expression rather than reproduce the original clustering, because hierarchical clustering on the first ten PCs performs poorly on the raw data, and because scImpute was given the number of cell types as input while the other methods were not.
Significance. If the amended benchmark indeed reproduces the original ranking, the reply would strengthen confidence in the original SAVER evaluation and clarify an important misconception about the role of the Gamma distribution in the simulation. The amended downsampling procedure is a step toward more realistic synthetic benchmarks because it matches the marginal statistics of real datasets. However, the central claim of agreement is stated without any numerical results, error bars, or statistical comparison, and the procedure is only shown to preserve univariate properties, not the joint correlation structure that imputation methods are designed to recover. The reply is therefore more a clarification and proposal than a complete validation. Its contribution would be substantially increased by reporting the actual metric values and by checking joint distributional fidelity.
major comments (3)
- [Section 2 (page 2-3), 'Using this amended downsampling technique...'] The central claim that 'the results obtained on the new downsampled datasets are in agreement with the findings in our Brief Communication' is not supported by any numerical evidence in the manuscript. Figure 2 is described qualitatively, but no correlation coefficients, Jaccard indices, or correlation-matrix distances are reported for any dataset or method. Without these values, the reader cannot tell how close the agreement is, whether it is consistent across all four datasets, or whether it is within sampling variability. Please provide the actual numbers, ideally with bootstrap confidence intervals or standard errors, for Figure 2a-c for all datasets and methods.
- [Section 2 (page 2), 'Using this amended downsampling technique...'] The amended downsampling procedure matches the mean expression, standard deviation, and zero fraction of the original data at the gene level. However, the benchmarks in Figure 2 (gene/cell-level correlation with the reference, correlation-matrix distance, and clustering) depend on joint gene-gene and cell-cell correlations and on the gene-specific dropout pattern, none of which are verified. It is possible for a dataset to match all three univariate statistics while having very different correlation structure. The reply should check, for example, the distribution of pairwise gene-gene correlation coefficients and the per-gene dropout rates of the simulated data against the original data, or otherwise justify that the first-order matching suffices for the chosen evaluation metrics.
- [Section 3 (page 3-4), 'Furthermore, the choice of clustering method...'] The reply's valid point that scImpute received the number of major cell types as input while SAVER and MAGIC were blind to it should be made quantitative. As written, the reply does not state what value of K was provided, how SAVER and MAGIC were configured, or what clustering parameters (e.g., Seurat resolution) were used for the comparison. The claim that hierarchical clustering is unreliable relies on a single unquantified observation that the adjusted Rand index is 'only slightly above 0' for the raw data. Please report the actual ARI or Jaccard values for the raw data and for each imputed dataset so that the reader can evaluate the degree of bias introduced by the asymmetric cluster-number input.
minor comments (5)
- [Page 2, first paragraph] The sentence 'We agree that even though we tried to match the mean expression and total zero fraction of the downsampled datasets with the original datasets, the distributions of these characteristics differ from the original datasets' is wordy and could mislead. A clearer formulation would be: 'We agree that the distributions of these characteristics in the downsampled data differ from the original data, even though we matched the mean expression and total zero fraction.'
- [Page 2, 'To do this, we randomly selected...'] The reply does not describe how the 'reference data' were constructed (e.g., the criteria for 'high quality cells' and 'highly expressed genes'). A brief summary of that procedure, or a pointer to the Methods of the original Brief Communication, should be included for clarity.
- [Figure 2 caption] The caption for Figure 2 is incomplete: it does not define the color coding, the clustering algorithm parameters (such as Seurat resolution), or what is meant by 'reference classification.' This information is needed to interpret the figure and to reproduce the analysis.
- [Page 5, final sentence] The claim that third-party evaluations are the 'ultimate source of benchmarking' is asserted without specific results from references 13-15. Please state which findings from those evaluations support this view, or soften the statement.
- [Page 2-3, general] The term 'reference' is used to mean both the original data subset and the 'true' expression matrix. Define this term at first use to avoid ambiguity, e.g., 'reference (treated as true expression).'
Circularity Check
No significant circularity: the reply's amended downsampling benchmark is an external evaluation, not a derivation that reduces to its inputs.
full rationale
The paper's central claims are (1) that SAVER's original evaluation included both an RNA FISH validation and a downsampling experiment, and (2) that an amended downsampling procedure yields results in agreement with those earlier findings. Neither claim is a derivation in which an output is defined in terms of the claimed conclusion. The amended procedure randomly subsets genes and cells from the original data, computes mean expression and library size, and downsamples from a reference matrix to match those quantities; the three imputation methods are then scored by correlation with the reference, correlation-matrix distance, and clustering. The reference is a fixed stand-in for true expression and is not informed by the methods' outputs, so the comparison is not equivalent to the simulation input by construction. The reply's distinction between a Gamma prior on expression (SAVER) and a Gamma distribution on sequencing efficiency is a statistical modeling argument rather than a self-definitional reduction; even if the marginal noise model is related to SAVER's assumptions, that is a potential bias in benchmark design, not circularity. The agreement with the earlier Brief Communication is an empirical claim that could have failed, and no fitted parameter is relabeled as a prediction. Self-citations to the original SAVER paper identify the method being replied to but are not load-bearing evidence for the new benchmark, which is supported by the rerun and by independent third-party evaluations (references 13-15). Limitations on preserving joint gene-gene structure concern external validity and are appropriately assigned to correctness risk, not circularity.
Assumptions & free parameters
free parameters (1)
- Mean sequencing efficiency (tau_c) =
0.05 or 0.10 depending on the dataset
assumptions (4)
- domain assumption Observed UMI counts follow a Poisson noise model with cell-specific efficiency.
- domain assumption Sequencing efficiency tau_c follows a Gamma distribution.
- ad hoc to paper A random subset of genes and cells from real data, with mean expression and library size matched, is a realistic stand-in for true expression and a valid testbed for imputation.
- ad hoc to paper Imputation should be judged by recovery of the reference rather than by reproducing original cell-type labels.
Cite this review
Pith. "Pith review of Reply to "Issues arising from benchmarking single-cell RNA sequencing imputation methods"." pith.science (2026). https://pith.science/paper/OQI7RZBV
@misc{pith2026190902492,
author = {Pith},
title = {Pith review of: Reply to "Issues arising from benchmarking single-cell RNA sequencing imputation methods"},
year = {2026},
howpublished = {\url{https://pith.science/paper/OQI7RZBV}},
note = {Machine review of arXiv:1909.02492}
}
read the original abstract
In our Brief Communication (DOI: 10.1038/s41592-018-0033-z), we presented the method SAVER for recovering true gene expression levels in noisy single cell RNA sequencing data. We evaluated the performance of SAVER, along with comparable methods MAGIC and scImpute, in an RNA FISH validation experiment and a data downsampling experiment. In a Comment [arXiv:1908.07084v1], Li & Li were concerned with the use of the downsampled datasets, specifically focusing on clustering results obtained from the Zeisel et al. data. Here, we will address these comments and, furthermore, amend the data downsampling experiment to demonstrate that the findings from the data downsampling experiment in our Brief Communication are valid.
Reference graph
Works this paper leans on
-
[2]
Huang, M. et al. SAVER: gene expression recovery for single-cell RNA sequencing. Nat. Methods 15, 539–542 (2018)
work page 2018
-
[3]
van Dijk, D. et al. Recovering Gene Interactions from Single-Cell Data Using Data Diffusion. Cell 1–14 (2018)
work page 2018
-
[4]
Li, W. V. & Li, J. J. An accurate and robust imputation method scImpute for single-cell RNA-seq data. Nat. Commun. 9, 997 (2018)
work page 2018
-
[5]
Li, W. V. & Li, J. J. Issues arising from benchmarking single-cell RNA sequencing imputation methods. ArXiv (2019). arXiv:1908.07084v1 [stat.AP]
work page Pith review arXiv 2019
-
[6]
Zeisel, A. et al. Cell types in the mouse cortex and hippocampus revealed by single-cell RNA-seq. Science (80-. ). 347, 1138–1142 (2015)
work page 2015
-
[7]
Wang, J. et al. Gene expression distribution deconvolution in single-cell RNA sequencing. Proc. Natl. Acad. Sci. 115, E6437–E6446 (2018)
work page 2018
-
[8]
Kim, J. K., Kolodziejczyk, A. A., Illicic, T., Teichmann, S. A. & Marioni, J. C. Characterizing noise structure in single-cell RNA-seq distinguishes genuine from technical stochastic allelic expression. Nat. Commun. 6, 8687 (2015)
work page 2015
-
[9]
Svensson, V. et al. Power analysis of single-cell RNA-sequencing experiments. Nat. Methods 14, 381–387 (2017)
work page 2017
Show all 15 references
-
[10]
Baron, M. et al. A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell Population Structure. Cell Syst. 3, 346–360 (2016)
2016
-
[11]
& Zhang, Y
Chen, R., Wu, X., Jiang, L. & Zhang, Y. Single-Cell RNA-Seq Reveals Hypothalamic Cell Diversity. Cell Rep. 18, 3227–3241 (2017)
2017
-
[12]
La Manno, G. et al. Molecular Diversity of Midbrain Development in Mouse, Human, and Stem Cells. Cell 167, 566-580.e19 (2016)
2016
-
[13]
A., Gennert, D., Schier, A
Satija, R., Farrell, J. A., Gennert, D., Schier, A. F. & Regev, A. Spatial reconstruction of single-cell gene expression data. Nat. Biotechnol. 33, 495–502 (2015)
2015
-
[14]
Andrews, T. S. & Hemberg, M. False signals induced by single-cell imputation [version 2; peer review: 3 approved, 1 approved with reservations]. F1000Research 7, (2019)
2019
-
[15]
Tian, L. et al. scRNA-seq mixology: towards better benchmarking of single cell RNA-seq analysis methods. bioRxiv 433102 (2019). 6
2019
-
[16]
& Hellmann, I
Vieth, B., Parekh, S., Ziegenhain, C., Enard, W. & Hellmann, I. A Systematic Evaluation of Single Cell RNA-Seq Analysis Pipelines: Library preparation and normalisation methods have the biggest impact on the performance of scRNA-seq studies. bioRxiv 583013 (2019)
2019
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.