{"id":"b7c8c9e7-754b-4c0d-b7a6-ebe17b3192eb","arxiv_id":"2607.04987","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"Data-driven soft labels that estimate each DNA read's conditional cell-type distribution let read-level classifiers scale to 39-type whole-body deconvolution and cut MSE 2.56× versus UXM.","lead":"A modular framework called Syto uses data-driven soft labels so DNA-read classifiers can deconvolve 39 whole-body cell types from methylation patterns. It cuts error more than twofold versus the prior state of the art and holds up on an out-of-distribution tissue set.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Pseudobulk + TCS proxies leave the 2.56× claim unanchored to real multi-type mixtures; that is the load-bearing gap.","rationale":"The reader correctly isolates the proxy-evaluation assumption as the weakest link. Soft labeling is a coherent, biologically motivated response to many-to-many methylation (Eq. 4 is the MLE of P(c|s); entropy-weighted CE and pooling are well-specified). The factorial design (360 paths), bootstrap CIs, and OOD TCS gains are internally consistent and improve on UXM under the chosen metrics. No internal contradiction or circularity appears. The remaining risk is external validity: without real multi-type ground truth the numerical 2.56× cannot be taken as a clinical or biological accuracy claim. That keeps the verdict CONDITIONAL (accept-shaped contribution pending real-mixture validation, broader baselines, and code). No stronger objection is warranted; the paper already flags the gap in §5.","tokens_in":71902,"tokens_out":635,"duration_ms":5910,"concrete_test":"Obtain or generate ≥20 independent multi-type mixtures with orthogonal ground-truth proportions (e.g., known-spike WGBS or FACS-sorted multi-lineage pools not used in GSE186458). Run the exact best Syto configuration of Table 1b and UXM U25 on those mixtures; if the MSE ratio falls below ~1.5× or loses significance under the same BCa bootstrap, the 2.56× claim does not transfer beyond the proxy regime.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline 2.56× MSE reduction (Table 1: Dismir + DD soft labels w/ pool. + PSLS + Lin. simplex, 1.17e-4 vs UXM 3.00e-4) is measured exclusively on Algorithm-2 pseudobulks subsampled from purified WGBS libraries (split 62:19:19, N=475k reads, K=100k mixtures). OOD transfer is scored only by Tissue Concordance Score (sum of predicted mass on Tabula-Sapiens-mapped expected types, §A.4.5), not by known proportions. Soft labels themselves are empirical frequencies of those same purified labels (Eq. 4 + Alg. 1). Consequently the entire pipeline never observes a true multi-type bulk whose ground-truth proportions are independent of the reference purification process. Section 5 candidly notes the absence of public known-proportion samples, yet the strongest claim is still stated as a deconvolution accuracy gain. If residual purification noise, coverage bias, or DMR-selection artifacts are shared between training soft labels and test pseudobulks, the reported MSE ratio can be inflated relative to real heterogeneous tissue. That is the single condition on which the central claim is least secure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces Syto, a modular read-level classification framework for methylome-based cell-type deconvolution that scales to dozens of classes. The central technical contribution is data-driven soft labeling: for each methylation signature the soft label is the (reweighted, Jaccard-pooled) empirical conditional distribution over cell types, presented as the MLE of P(c|s). This is combined with extensions of CancerDetector, Dismir and MethylBERT, five deconvolvers (including PSLS/NNLS and learned models), and a hyperparameter-free linear calibrator for simplex outputs. On 39-class WGBS pseudobulks from the Loyfer atlas, the best configuration (Dismir + soft labels with pooling + PSLS + linear simplex calibration) reports a 2.56\times MSE reduction versus UXM (1.17e-4 vs 3.00e-4); gains transfer to an OOD RRBS tissue atlas under a Tissue Concordance Score (TCS) proxy. Soft labeling is argued to be generally useful for many-to-many signal-to-label maps.","tokens_in":72368,"tokens_out":1411,"duration_ms":11286,"significance":"If the soft-label construction and modular pipeline hold under stronger evaluation, the work would remove a genuine bottleneck: existing read-level classifiers (CancerDetector, Dismir, MethylBERT) have been limited to few classes because hard labels conflict with non-discriminative off-target reads. Scaling to whole-body panels and framing tumor purity/classification as special cases of deconvolution would matter for liquid-biopsy and multi-omics applications. Strengths that should be credited include the exhaustive 360-permutation design with BCa bootstrap CIs, the explicit demonstration that MLE deconvolvers fail at this scale (Tables S1–S3), the feature-selection ablations, the linear calibrator comparison against Dirichlet/vector/temperature scaling, and the candid Limitations section. The soft-label MLE argument and entropy-weighted CE are clean and potentially portable beyond methylation.","major_comments":[{"comment":"The headline 2.56× MSE claim (Abstract; Table 1 best row: Dismir + DD soft labels w/ pool. + PSLS + Lin. simplex, 1.17e-4 vs UXM 3.00e-4) is measured exclusively on Algorithm-2 pseudobulks subsampled from purified WGBS libraries (N=475k, K=100k; §4.1, A.4.2). Soft labels themselves are empirical frequencies of those same purified labels (Eq. 4 + Alg. 1). No public multi-type bulk with independent known proportions is used. Section 5 acknowledges this, yet the strongest claim is still framed as a deconvolution accuracy gain. Residual purification noise, coverage bias, or DMR-selection artifacts shared between training soft labels and test pseudobulks can inflate the reported ratio relative to real heterogeneous tissue. Either (i) add at least one external known-proportion or spike-in multi-type evaluation, or (ii) restate the central claim as a large-scale pseudobulk result and demote the","section":null},{"comment":"OOD transfer is scored only by Tissue Concordance Score (TCS = sum of predicted mass on Tabula-Sapiens-mapped expected types, §A.4.5; Fig. 4, Tables S17–S21). TCS is an upper-bound proxy that does not penalize misallocation among expected types and depends on a CL/UBERON mapping plus a ≥1% threshold. The claim that gains “transfer” to 16 tissues is therefore weaker than a proportion-level claim. Report TCS alongside at least one secondary metric (e.g., expected-type KL or rank correlation with Tabula Sapiens fractions where available) and state explicitly that TCS is not a substitute for known-proportion MSE.","section":null},{"comment":"UXM is given a structural advantage: its atlas is built from all splits (§4.1), while Syto classifiers are trained on the 62% train split only. The 2.56× comparison therefore understates Syto’s relative gain if anything, but the absolute SoTA framing is slightly unfair. Either retrain UXM on the same train-only atlas or report both “UXM-all” and “UXM-train” baselines so the reader can calibrate the gap.","section":null}],"minor_comments":[{"comment":"Figure 3 shows non-convergence of hard labels but does not report the corresponding soft-label training curves; adding them would make the labeling comparison self-contained.","section":null},{"comment":"Notation: Δ^{C−1} is used for both C-class soft labels and C+1 hard-with-background labels; clarify when the background class expands the simplex.","section":null},{"comment":"Algorithm 1: the asymmetric Jaccard pooling and inverse-frequency reweighting are important; a short sensitivity table for τ and δ_max would strengthen reproducibility.","section":null},{"comment":"Table 1 caption and §A.4.4: “top 156 features” and “† = diagonal + background” are easy to miss; consider a one-line legend in the main table.","section":null},{"comment":"Related work on label enhancement (Xu et al., Wang & Geng) is cited; a sentence contrasting entropy-upweighting in prior LE with the paper’s low-entropy upweighting would help readers outside the methylation community.","section":null},{"comment":"Typos / polish: “Sytolays”, “Sytoconsistently”, “Sytoreduces” (missing spaces after Syto in Abstract/Introduction); “easurable” → “measurable” (A.4.3); “assesed” → “assessed” (A.3.2).","section":null}],"recommendation":"major_revision","confidential_remarks":"The technical core (soft-label MLE + modular pipeline + MLE-deconvolver failure analysis) is solid and the 360-permutation design is unusually thorough for this subfield. The load-bearing weakness is evaluation: pseudobulk + TCS. I would accept after the authors either add a real multi-type check or reframe the claim; I would not reject. Fit for a methods-oriented ML/comp-bio venue is good; for a pure biology journal the proxy metrics would need more defense."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The real news is that they finally got pattern-sensitive read classifiers to converge on a 39-type whole-body atlas. Hard labels fail because most reads are off-target and the methylation-to-cell-type map is many-to-many; their data-driven soft labels (empirical P(c|s) with inverse-frequency reweighting and asymmetric Jaccard pooling) plus an entropy-weighted CE loss fix that. They wrap it in a clean modular stack (classifier / deconvolver / linear simplex calibrator), extend CancerDetector, Dismir and MethylBERT, and show a 2.56× MSE drop versus UXM on 100k pseudobulks (1.17e-4 vs 3.00e-4) with bootstrap CIs. Gains transfer, via Tissue Concordance Score, to an OOD RRBS tissue set.\n\nWhat is new is the soft-label construction tailored to this biology, not label distribution learning in general. The factorial design (360 combinations), pure-mixture MLE failure analysis, feature-selection ablations, and candid limitations section are careful. Linear calibration for simplex outputs is simple and works. The soft-label idea is portable to other many-to-many settings.\n\nThe soft spot is evaluation, and it is load-bearing. Everything is measured on Algorithm-2 pseudobulks subsampled from purified libraries; OOD uses only TCS (sum of mass on Tabula-Sapiens-expected types), not known proportions. Soft labels themselves are built from the same purified counts. Section 5 admits there are no public multi-type ground-truth mixtures. That does not invent the result, but it means the headline accuracy claim is still one step removed from real bulk tissue. Baselines are also narrow (mainly UXM), DMR panels are small (25 per type), and no Syto code is shipped. Those are real but proportionate caveats, not circularity or math failure.\n\nThis is for people who do methylome deconvolution, ctDNA, or multi-class omics classification. It deserves a serious referee. I would engage: cite the soft-label construction and the scaling result, and push for real-mixture validation and code.","headline":"Solid engineering that makes read-level methylome deconvolution work at 39 classes via data-driven soft labels; the 2.56× MSE gain is real on their setup, but rests on pseudobulk + TCS proxies.","tokens_in":72981,"tokens_out":562,"would_cite":true,"duration_ms":8632,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Data-driven soft labels let DNA-read classifiers deconvolve 39 whole-body cell types, cutting error 2.56× versus prior methods.","keywords":["cell-type deconvolution","DNA methylation","read-level classification","soft labels","label enhancement","whole-body atlas","calibration","many-to-many mapping"],"falsifier":"Run the same best Syto configuration on a set of real bulk samples whose cell-type proportions have been independently measured by single-cell or flow-cytometry ground truth; if the reported 2.56-fold MSE reduction disappears, the central claim fails.","tokens_in":72766,"feed_emoji":"🧬","tokens_out":710,"duration_ms":6231,"temperature":0.7,"pith_summary":"Cell-type deconvolution estimates how much of each cell type is present in a mixed biological sample. Most methylation-based methods throw away the pattern carried by each individual DNA read and work only on averages; the few methods that keep read-level patterns never scaled past a handful of classes. Hard one-hot labels clash with biology: the same methylation pattern can arise from several cell types, and off-target reads dominate once dozens of types are included, so classifiers never converge. The paper shows that replacing hard labels with data-driven soft labels—empirical conditional distributions over cell types for each observed methylation signature—removes the conflict. These soft labels are embedded in Syto, a modular pipeline of classifier, deconvolver and calibrator. On a 39-cell-type whole-body atlas the best Syto configuration reduces mean-squared error 2.56-fold relative to the previous state of the art; the same gains appear on an out-of-distribution tissue panel. The result opens the door to larger reference panels and treats tumor purity and classification as special cases of deconvolution.","feed_headline":"Soft labels cut DNA deconvolution error 2.56× across 39 cell types","feed_subtitle":"Read-level classifiers finally scale to whole-body panels and transfer to new tissues","key_machinery":"Data-driven soft labels: for each unique methylation signature the empirical frequency of cell types among reads that share that signature (with class reweighting and nearest-neighbor pooling for low coverage). These soft targets replace one-hot labels inside an entropy-weighted cross-entropy loss, enabling the modular Syto pipeline of classifier, deconvolver and linear calibrator.","core_discovery":"Data-driven soft labels that estimate the true conditional cell-type distribution for each DNA-read methylation signature allow read-level classifiers to converge and deconvolve mixtures of 39 whole-body cell types, cutting mean-squared error by a factor of 2.56 relative to the prior state of the art and transferring the improvement to an out-of-distribution 16-tissue panel.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Soft labels scale DNA-read deconvolution to 39 whole-body cell types","Data-driven soft labels cut deconvolution MSE 2.56× on 39 cell types","Soft labels let read classifiers handle 39-cell whole-body mixtures","Read-level deconvolution reaches 39 types via conditional soft labels","Soft labeling transfers 2.56× DNA deconvolution gains to new tissues"],"cache_read_input_tokens":65664,"weakest_assumption_plain":"Performance is measured on artificial mixtures made by mixing purified cell-type libraries and on a proxy score that only checks whether expected tissue cell types appear, not on samples whose true multi-type proportions are known.","fun_headline_variants_meta":{"raw":{"variants":["Soft labels scale DNA-read deconvolution to 39 whole-body cell types","Data-driven soft labels cut deconvolution MSE 2.56× on 39 cell types","Soft labels let read classifiers handle 39-cell whole-body mixtures","Read-level deconvolution reaches 39 types via conditional soft labels","Soft labeling transfers 2.56× DNA deconvolution gains to new tissues"]},"model":"grok-4.5","effort":"low","cost_usd":0.006776,"raw_usage":{"total_tokens":1692,"prompt_tokens":796,"num_sources_used":0,"completion_tokens":105,"cost_in_usd_ticks":67760000,"prompt_tokens_details":{"text_tokens":796,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":791,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":796,"tokens_out":105,"duration_ms":6054,"temperature":1.0,"reasoning_tokens":791,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T10:29:47.992424+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same best Syto configuration on a set of real bulk samples whose cell-type proportions have been independently measured by single-cell or flow-cytometry ground truth; if the reported 2.56-fold MSE reduction disappears, the central claim fails.","supporting_citations":[],"review_version":1}