{"id":"768d452c-54fa-4976-b743-843205141bc5","arxiv_id":"1909.02492","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"The authors redo their downsampling benchmark with a revised procedure that matches real data characteristics and report that SAVER's original conclusions still hold.","lead":"This reply defends the authors' earlier benchmark of single-cell RNA sequencing imputation methods against a published critique, and adds an amended downsampling experiment that reproduces the original conclusions. It is a short methodological exchange rather than a new method.","discovery_kind":"replication","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The amended downsampling protocol matches only mean expression and library size; it does not verify that gene-gene correlations or gene-specific dropout are preserved, so the claimed agreement with the original benchmark cannot establish validity.","rationale":"The reader's weakest assumption concerned first-order moment matching. I agree and sharpen it: the paper matches univariate gene-level statistics but not the joint statistics (gene-gene correlation, gene-specific dropout) that imputation methods target. The concrete test would provide direct evidence. Because the paper also omits numeric agreement results, the CONDITIONAL verdict is appropriate: the central claim is plausible but unverified. No new objection beyond the reader's; I recommend UNCHANGED.","tokens_in":3239,"tokens_out":5151,"duration_ms":55814,"concrete_test":"Compare the gene-gene correlation matrix of each new downsampled dataset with the corresponding original dataset using the same correlation matrix distance metric reported in Figure 2b, and also compare the gene-specific zero fractions (probability of zero for each gene). If the simulated datasets differ from the original by more than a technical-replicate pair would, or if the gene-specific zero fractions deviate systematically (e.g., higher dropout for low-expression genes than observed), then the amended simulation does not preserve the joint structure on which imputation benchmarking depends. This check would settle whether the claimed agreement is meaningful.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim, that the new downsampled datasets reproduce the findings of the Brief Communication, rests on the assumption that the amended simulation is a faithful testbed. Section 2 (page 2) states that the authors randomly selected a subset of genes and cells from the original dataset, calculated mean expression and library size, and downsampled from the reference data to match these quantities. Figure 1 shows that univariate gene-level statistics (mean, standard deviation, zero fraction) match almost exactly. However, imputation methods are evaluated on their ability to recover joint structure: gene-gene correlations, cell-cell correlations, and gene-specific dropout patterns. The reply does not demonstrate that the new downsampled datasets preserve these properties. In fact, because the observed data are generated from a reference matrix that is itself a filtered subset (high-quality cells, highly expressed genes), and noise is added independently across genes, the joint distribution of the simulated data may differ substantially from the original. Thus, even if the rerun rankings agree with the original paper, this agreement could be an artifact of the simulation rather than evidence that SAVER outperforms MAGIC and scImpute on real scRNA-seq data. Additionally, the statement 'results obtained ... are in agreement' is not accompanied by numerical summaries or effect sizes, making it impossible to assess the strength or consistency of the agreement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This short reply addresses Li & Li's comment on the downsampling experiment in the authors' SAVER Brief Communication. The authors first correct a misunderstanding: their downsampling procedure uses a Poisson noise model with a Gamma-distributed cell-specific sequencing efficiency, not the Poisson-Gamma prior assumed by SAVER. They then acknowledge that the original downsampled data had marginal distributions that deviate from real data, and they propose an amended procedure that randomly selects a subset of genes and cells from the original dataset and downsamples the reference data to match the selected mean expression and library size. They report that the new downsampled data match the original data in mean expression, standard deviation, and zero fraction (Figure 1), and they state that re-running SAVER, MAGIC, and scImpute on these new datasets reproduces the findings of the original Brief Communication (Figure 2). Finally, they argue that Li & Li's real-data evaluation using hierarchical clustering is flawed because imputation should recover true expression rather than reproduce the original clustering, because hierarchical clustering on the first ten PCs performs poorly on the raw data, and because scImpute was given the number of cell types as input while the other methods were not.","tokens_in":3657,"tokens_out":4936,"duration_ms":56902,"significance":"If the amended benchmark indeed reproduces the original ranking, the reply would strengthen confidence in the original SAVER evaluation and clarify an important misconception about the role of the Gamma distribution in the simulation. The amended downsampling procedure is a step toward more realistic synthetic benchmarks because it matches the marginal statistics of real datasets. However, the central claim of agreement is stated without any numerical results, error bars, or statistical comparison, and the procedure is only shown to preserve univariate properties, not the joint correlation structure that imputation methods are designed to recover. The reply is therefore more a clarification and proposal than a complete validation. Its contribution would be substantially increased by reporting the actual metric values and by checking joint distributional fidelity.","major_comments":[{"comment":"The central claim that 'the results obtained on the new downsampled datasets are in agreement with the findings in our Brief Communication' is not supported by any numerical evidence in the manuscript. Figure 2 is described qualitatively, but no correlation coefficients, Jaccard indices, or correlation-matrix distances are reported for any dataset or method. Without these values, the reader cannot tell how close the agreement is, whether it is consistent across all four datasets, or whether it is within sampling variability. Please provide the actual numbers, ideally with bootstrap confidence intervals or standard errors, for Figure 2a-c for all datasets and methods.","section":"Section 2 (page 2-3), 'Using this amended downsampling technique...'"},{"comment":"The amended downsampling procedure matches the mean expression, standard deviation, and zero fraction of the original data at the gene level. However, the benchmarks in Figure 2 (gene/cell-level correlation with the reference, correlation-matrix distance, and clustering) depend on joint gene-gene and cell-cell correlations and on the gene-specific dropout pattern, none of which are verified. It is possible for a dataset to match all three univariate statistics while having very different correlation structure. The reply should check, for example, the distribution of pairwise gene-gene correlation coefficients and the per-gene dropout rates of the simulated data against the original data, or otherwise justify that the first-order matching suffices for the chosen evaluation metrics.","section":"Section 2 (page 2), 'Using this amended downsampling technique...'"},{"comment":"The reply's valid point that scImpute received the number of major cell types as input while SAVER and MAGIC were blind to it should be made quantitative. As written, the reply does not state what value of K was provided, how SAVER and MAGIC were configured, or what clustering parameters (e.g., Seurat resolution) were used for the comparison. The claim that hierarchical clustering is unreliable relies on a single unquantified observation that the adjusted Rand index is 'only slightly above 0' for the raw data. Please report the actual ARI or Jaccard values for the raw data and for each imputed dataset so that the reader can evaluate the degree of bias introduced by the asymmetric cluster-number input.","section":"Section 3 (page 3-4), 'Furthermore, the choice of clustering method...'"}],"minor_comments":[{"comment":"The sentence 'We agree that even though we tried to match the mean expression and total zero fraction of the downsampled datasets with the original datasets, the distributions of these characteristics differ from the original datasets' is wordy and could mislead. A clearer formulation would be: 'We agree that the distributions of these characteristics in the downsampled data differ from the original data, even though we matched the mean expression and total zero fraction.'","section":"Page 2, first paragraph"},{"comment":"The reply does not describe how the 'reference data' were constructed (e.g., the criteria for 'high quality cells' and 'highly expressed genes'). A brief summary of that procedure, or a pointer to the Methods of the original Brief Communication, should be included for clarity.","section":"Page 2, 'To do this, we randomly selected...'"},{"comment":"The caption for Figure 2 is incomplete: it does not define the color coding, the clustering algorithm parameters (such as Seurat resolution), or what is meant by 'reference classification.' This information is needed to interpret the figure and to reproduce the analysis.","section":"Figure 2 caption"},{"comment":"The claim that third-party evaluations are the 'ultimate source of benchmarking' is asserted without specific results from references 13-15. Please state which findings from those evaluations support this view, or soften the statement.","section":"Page 5, final sentence"},{"comment":"The term 'reference' is used to mean both the original data subset and the 'true' expression matrix. Define this term at first use to avoid ambiguity, e.g., 'reference (treated as true expression).'","section":"Page 2-3, general"}],"recommendation":"major_revision","confidential_remarks":"This is a reply to a published Comment, so the scope is necessarily narrow. The main issue is not the logic of the reply but the absence of the numerical results that would substantiate its central claim. If the authors can supply the missing numbers and add the joint-distribution check, the paper could be suitable for publication. I would not recommend rejection because the criticisms of Li & Li's hierarchical clustering design have merit and the Poisson-Gamma clarification is useful. However, in its current form the reply does not provide enough evidence for the editor or readers to evaluate the strength of the claimed agreement."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this reply is exactly what a comment/reply exchange should look like: it engages with each specific objection, concedes what is fair, and amends the experiment rather than just defending the original. The two sharpest points are well taken. The first is that scImpute was handed the true number of clusters, which is a real advantage in a clustering-based evaluation and does not reflect typical practice. The second is that Li & Li's hierarchical clustering on the raw Zeisel data yields near-zero ARI against the published labels, so using that clustering as a ground-truth test is shaky. Both points land.\n\nThe genuinely new thing is the amended downsampling procedure: randomly subset genes and cells, match mean expression and library size to the original data, then downsample from the reference. The univariate gene-level statistics look nearly identical to the original data, which directly answers Li & Li's Figure 1 comparison. That is a legitimate improvement and it shows the authors are willing to revise their own protocol.\n\nWhere the reply is soft is in its central empirical claim. The phrase \"the results obtained on the new downsampled datasets are in agreement with the findings in our Brief Communication\" appears without any numerical summary, effect size, or confidence measure. Figure 2 is described but not shown in a way the reader can verify from the text, and there is no code or data link. The stress-test concern also has merit: matching mean and library size does not guarantee that gene-gene correlations, cell-cell structure, or gene-specific dropout patterns match the real data. To be fair, the authors explicitly say they prefer downsampling over fully synthetic data precisely because it preserves gene-gene interactions from the reference, and the reference is a real filtered expression matrix. But the reply does not demonstrate preservation of those joint properties, so the agreement could still be an artifact of the simulation.\n\nNet: the logic is coherent, the critique of scImpute's advantage is convincing, and the amended procedure is a step forward. But the lack of numbers, code, and joint-structure checks means the reply is more persuasive as argument than as evidence. A serious referee would reasonably ask for the actual performance values, error bars, and a check that correlation structure is preserved. I would send it to review rather than desk reject, but I would not cite it without seeing the supplementary material.","headline":"A coherent reply that fixes the univariate mismatch in the downsampling and flags scImpute's unfair advantage, but the central claim of agreement is asserted without the numbers needed to verify it.","tokens_in":3946,"tokens_out":1385,"would_cite":false,"duration_ms":16989,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SAVER's benchmark ranking survives a revised downsampling test.","keywords":["single-cell RNA sequencing","imputation benchmarking","SAVER","MAGIC","scImpute","downsampling experiment","Poisson-Gamma noise model","method rebuttal"],"falsifier":"Construct amended downsampled versions of the four datasets while preserving the empirical gene-gene covariance, or use technical replicate pairs as ground truth, then rerun SAVER, MAGIC, and scImpute: if the SAVER advantage in correlation with the reference or in clustering agreement disappears or reverses, the reply's validation claim is falsified.","tokens_in":3030,"feed_emoji":"🧬","tokens_out":8166,"duration_ms":70575,"temperature":0.7,"pith_summary":"This reply defends the benchmarking of the SAVER gene-expression recovery method against a critique that the original downsampled test data did not look like real single-cell data. Its central assertion is that a new downsampling protocol, which randomly subsets genes and cells from real datasets and then matches their mean expression and library size while adding Poisson-Gamma noise, produces test data whose mean, variance, and zero-fraction match real data almost exactly. On those new test sets, the reply reports that SAVER, MAGIC, and scImpute produce the same relative performance as in the original Brief Communication, so the original findings stand. The reply also argues that the critique's alternative evaluation, clustering the full dataset with hierarchical clustering while giving scImpute the true number of clusters, does not measure what imputation is for.","feed_headline":"SAVER benchmark survives a stricter downsampling test","feed_subtitle":"A reply reruns the comparison with noise matched to real data and reports the same ranking holds.","key_machinery":"The load-bearing mechanism is the amended downsampling protocol: random subsetting of genes and cells from a real dataset, calculation of the subset's mean expression and library size, and then generation of noisy counts $Y_{gc} \\sim \\mathrm{Poisson}(\\tau_c \\lambda_{gc})$ with $\\tau_c$ drawn from a Gamma distribution matched to the selected library size. This protocol is what lets the authors claim that the synthetic test data share the first-order statistics of real data while still serving as a known ground truth. The evaluation then uses Pearson correlation with the reference, distance between gene-gene and cell-cell correlation matrices, and Seurat-based clustering compared by Jaccard index. The reply's rebuttal also rests on the choice of clustering method and on whether the number of clusters is provided to a method in advance.","core_discovery":"The paper claims that the earlier conclusion, that SAVER recovers reference expression levels more accurately than MAGIC and scImpute as measured by gene-level and cell-level correlation, correlation-matrix distance, and clustering agreement, remains true when the downsampling simulation is corrected to better match real data. The authors construct amended downsampled datasets by picking random gene and cell subsets from each original dataset, computing the mean expression and library size of that subset, and sampling noisy observations from a Poisson-Gamma noise model tuned to those quantities. The resulting datasets have nearly identical mean, standard deviation, and zero fraction per gene as the original data, unlike the first version. Running the same three methods and metrics on these new datasets gives results that the paper states are in agreement with the findings in its Brief Communication.","pith_inferences":["The amended protocol matches mean expression and library size but does not explicitly preserve gene-gene correlations or dropout patterns, so a stricter test would generate noise with the empirical covariance or use technical replicate pairs.","The exchange suggests that benchmarking conclusions in single-cell imputation can hinge on protocol choices such as the clustering algorithm and whether simulated noise preserves higher-order structure, pointing to a need for standardized noise-generation benchmarks.","One testable extension is to apply the amended downsampling protocol to additional imputation methods or to datasets with known spike-in controls and check whether the relative ranking of SAVER, MAGIC, and scImpute remains unchanged."],"forward_implications":["If the amended experiment is a fair test, the original ranking of SAVER over MAGIC and scImpute on the four datasets is stable under a more realistic noise model.","The critique's objections about the old downsampled data, such as zero-fraction mismatch and low per-gene agreement with the original data, do not invalidate the benchmark because those properties are expected consequences of adding noise.","Evaluations that compare imputed output to labels derived from the same noisy data, or that use an unstable clustering method, are not reliable ways to judge imputation.","Giving scImpute the true number of clusters while other methods are blind to it makes the comparison unfair to the other methods."],"supporting_citations":[{"why":"Presents the SAVER method and defines the original benchmark whose findings the reply defends.","marker":"[1]"},{"why":"MAGIC is one of the three imputation methods compared in the benchmark.","marker":"[2]"},{"why":"scImpute is one of the three imputation methods compared in the benchmark.","marker":"[3]"},{"why":"The Comment whose objections about downsampling and clustering the reply addresses.","marker":"[4]"},{"why":"Provides one of the four public datasets used for the amended downsampling experiments.","marker":"[5]"},{"why":"Supports the Poisson model for UMI technical noise in scRNA-seq.","marker":"[6]"},{"why":"Also supports the Poisson noise description for UMI-based single-cell data.","marker":"[7]"},{"why":"Supports the Gamma distribution used for cell-specific sequencing efficiency.","marker":"[8]"},{"why":"Provides another of the four public datasets used for the amended downsampling experiments.","marker":"[9]"},{"why":"Supplies the Seurat clustering workflow used to evaluate cell-type recovery in the benchmark.","marker":"[12]"}],"fun_headline_variants":["Amended downsampling reaffirms SAVER's edge","SAVER stands after fairer downsampling test","Corrected benchmark still favors SAVER","Noise-matched test backs SAVER findings","Reply strengthens SAVER's benchmark case"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument collapses if matching a simulated dataset's mean expression and library size to real data does not also reproduce the gene-gene correlations, dropout structure, and batch effects that real scRNA-seq noise has.","fun_headline_variants_meta":{"raw":{"variants":["Amended downsampling reaffirms SAVER's edge","SAVER stands after fairer downsampling test","Corrected benchmark still favors SAVER","Noise-matched test backs SAVER findings","Reply strengthens SAVER's benchmark case"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000346,"raw_usage":{"total_tokens":1840,"prompt_tokens":829,"completion_tokens":1011,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":952}},"tokens_in":445,"tokens_out":1011,"duration_ms":8086,"temperature":1.0,"reasoning_tokens":952,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:48:04.329624+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Construct amended downsampled versions of the four datasets while preserving the empirical gene-gene covariance, or use technical replicate pairs as ground truth, then rerun SAVER, MAGIC, and scImpute: if the SAVER advantage in correlation with the reference or in clustering agreement disappears or reverses, the reply's validation claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MAGIC is one of the three imputation methods compared in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"scImpute is one of the three imputation methods compared in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Comment whose objections about downsampling and clustering the reply addresses."},{"cited_title":"Issues arising from benchmarking single-cell RNA sequencing imputation methods","cited_arxiv_id":"1908.07084","evidence_quote":"Provides one of the four public datasets used for the amended downsampling experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the Poisson model for UMI technical noise in scRNA-seq."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Also supports the Poisson noise description for UMI-based single-cell data."},{"cited_title":"K., Kolodziejczyk, A","cited_arxiv_id":null,"evidence_quote":"Supports the Gamma distribution used for cell-specific sequencing efficiency."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides another of the four public datasets used for the amended downsampling experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Seurat clustering workflow used to evaluate cell-type recovery in the benchmark."}],"review_version":1}