{"id":"c1fd3fb7-4bc8-4eb3-b9a7-43dd9a6a6c41","arxiv_id":"2607.06129","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"A deep learning pipeline applied to multiplexed DLBCL tissue images identifies distinct tumor microenvironment compositions and cell interaction patterns between GCB and ABC subtypes.","lead":"This paper uses a deep learning pipeline to analyze multiplexed fluorescence images of 106 DLBCL tumor samples, extracting cell composition, morphology, and spatial organization features to compare GCB and ABC subtypes. A smart generalist might read it to understand how automated tissue imaging can reveal microenvironment differences between cancer subtypes.","discovery_kind":"new_application","skeptic_critique":{"model":"glm-5.2","headline":"The 470-feature statistical comparison (Fig. 4) lacks multiple-testing correction, and the unit of analysis may be the tissue core rather than the patient, raising pseudoreplication risk. Either issue alone weakens the 'significantly different' claim; together they are the most load-bearing concern.","rationale":"The reader's verdict of CONDITIONAL is appropriate. The reader correctly identified multiple issues including the lack of multiple-testing correction, but placed Otsu thresholding as the single most load-bearing concern. I would re-rank: the statistical framework (uncorrected 470-feature tests plus possible pseudoreplication) is more directly load-bearing for the 'significantly different' part of the central claim, because it determines whether the observed differences are reliable rather than chance findings. The Otsu concern is real and foundational—misclassification would affect all downstream measurements—but the results are biologically plausible and consistent with prior literature (refs 14-19), which provides some indirect support for the cell classifications even without ground-truth validation. The Otsu concern is best addressed by the reader's suggested validation against expert annotation, but the multiple-testing and pseudoreplication concerns are more immediately damaging to the specific claim of statistical significance. Both concerns are addressable: FDR correction and patient-level aggregation are straightforward recomputations. If the findings survive both corrections, the CONDITIONAL verdict could be upgraded; if they do not, the descriptive observations would still be interesting but the significance claims would need retraction. The paper's pipeline is a legitimate methodological contribution regardless of the statistical issues, and the biological observations are consistent with existing literature on GCB/ABC TME differences. No change to the verdict is needed, but the specific concerns should be sharpened as described.","tokens_in":13509,"tokens_out":5141,"duration_ms":364677,"concrete_test":"Recompute all p-values from Fig. 4 with Benjamini-Hochberg FDR correction across all 470 features. Separately, aggregate all features at the patient level (averaging across cores from the same patient) and rerun the Mann-Whitney U tests with FDR correction. If fewer than 10 features survive FDR < 0.05 at the patient level, the claim of 'significantly different' TME features would need to be substantially qualified, and the top-20 feature list in Fig. 4B would be unreliable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper performs Mann-Whitney U tests on 470 features (10 cell types × 5 morphological features × 7 distribution metrics + 10 cell proportions + 55 interaction features) comparing GCB and ABC tumors (Methods 'Feature Aggregation'; Fig. 4). No multiple-testing correction—Bonferroni, Benjamini-Hochberg, or permutation-based—is mentioned anywhere. At a p<0.05 threshold across 470 tests, approximately 23 features would be expected to be significant by chance alone. The paper highlights the 'top 20' features (Fig. 4B) and uses p<0.05 as the significance boundary in the volcano plot (Fig. 4A), but several of the headline findings—morphometric differences in M2-macrophages and CD8+ T-cells, and specific cell interaction patterns—are drawn from these uncorrected tests. A secondary but related concern: the paper describes 106 patients with 'up to 3 tissue cores per patient' and 559 total tissue samples, yet the feature aggregation and statistical analysis refer to 'samples' without clarifying whether the unit is the patient or the individual core. If cores from the same patient are treated as independent observations, pseudoreplication would further inflate all p-values. The descriptive observations in Fig. 2C and Fig. 3 may hold regardless, but the formal statistical framework that underpins the 'significantly different' language in the central claim is compromised by these two issues. The reader identified the multiple-testing gap in their rationale but placed Otsu thresholding as the primary concern; I would elevate the statistical framework as more directly load-bearing for the 'significantly different' portion of the claim, since even with perfect cell classification, uncorrected tests on 470 features with possible pseudoreplication cannot support robust significance claims.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The manuscript presents a deep learning-based pipeline for analyzing multiplexed immunofluorescence images of DLBCL tumor microarrays. The pipeline segments nuclei (Cellpose), classifies cells into 10 types using Otsu-thresholded marker expression, and extracts 470 features per tissue sample spanning morphology, cell-type proportions, and spatial interaction patterns. These features are compared between GCB and ABC subtypes (as defined by the Hans classifier) using Mann-Whitney U tests. The authors report that ABC tumors are immune-rich with preferential M2-macrophage interactions, while GCB tumors are immune-poor, and identify morphometric differences in M2-macrophages and CD8+ T-cells as the most discriminating features. The pipeline is modular and the dataset (106 patients, 559 tissue samples, 14 markers) is substantial.","tokens_in":14558,"tokens_out":1399,"duration_ms":177114,"significance":"The study provides a useful, modular computational framework for quantitative TME characterization from multiplexed imaging data, and the DLBCL cohort is reasonably sized. The descriptive observations on cell-type composition differences between GCB and ABC subtypes (Fig. 2C, Fig. 3) are broadly consistent with published literature on immune-cold GCB and immune-hot ABC phenotypes. The label-permutation approach for assessing interaction enrichment is a sensible design choice. However, the formal statistical framework underlying the headline 'significantly different' claims has two load-bearing gaps that must be addressed before the central quantitative claims can be considered reliable.","major_comments":[{"comment":"§Results, 'Statistical analysis of aggregated features'; Fig. 4: The Mann-Whitney U tests are performed across 470 features (10 cell types × 5 morphological features × 7 distribution metrics + 10 cell proportions + 55 interaction features) with a p<0.05 threshold and no mention of multiple-testing correction (Bonferroni, Benjamini-Hochberg, or permutation-based FDR). At this threshold, approximately 23 features would be expected significant by chance alone. Several headline findings—morphometric differences in M2-macrophages and CD8+ T-cells, and specific cell interaction patterns—are drawn from these uncorrected tests. The authors should apply an appropriate multiple-testing correction and report which features survive.","section":null},{"comment":"§Methods, 'DLBCL tissue cohort assembly'; §Methods, 'Feature Aggregation': The cohort comprises 106 patients with 'up to 3 tissue cores per patient' (559 total tissue samples), but the feature aggregation and statistical analysis refer to 'samples' without clarifying whether the unit of analysis is the patient or the individual core. If multiple cores from the same patient are treated as independent observations, pseudoreplication would inflate all p-values. The authors should state the unit of analysis explicitly and, if cores are used, either aggregate to patient level or use a mixed-effects model with patient as a random effect.","section":null},{"comment":"§Methods, 'Data Extraction'; Table 4: Cell classification relies on Otsu thresholding of single-cell protein expression values to binarize each of the 11 markers. Otsu's method assumes a bimodal intensity distribution, but fluorescence signal distributions in multiplexed imaging are frequently continuous. If thresholds misclassify cells (e.g., calling a cell CD20+ when it is not), all downstream proportions, interaction analyses, and morphological comparisons are affected. The paper does not validate the Otsu thresholds against expert annotation, pathologist review, or any ground truth. The authors should provide at least a limited validation against manual annotation or justify the bimodality assumption with representative intensity histograms.","section":null}],"minor_comments":[{"comment":"§Methods, 'Spatial Organization of the DLBCLs': The proximity threshold d* = 0 is chosen 'based on qualitative analysis of graphs at various cutoffs.' A sensitivity analysis over a range of d* values, or at least a brief description of the qualitative criteria, would strengthen this choice and improve reproducibility.","section":null},{"comment":"§Discussion, paragraph 3: The tumor cell definition (CD20+ with nuclear area ≥ 2× mean of CD20+ population) is introduced only in the Discussion. This is a load-bearing definition for several results and should be described in the Methods, with the nuclear area threshold listed as a parameter.","section":null},{"comment":"Fig. 2C: The y-axis label and tick marks are unclear. Error bars or interquartile ranges should be specified in the figure legend.","section":null},{"comment":"Fig. 3A: The color scale legend for the interaction enrichment heatmaps should be clarified—specifically, whether the color represents a z-score, log-fold change, or raw fraction of samples.","section":null},{"comment":"Fig. 4B: The y-axes are labeled 'Arbitrary units' without further specification. For a statistical comparison figure, the axes should indicate the actual metric being plotted (e.g., U statistic, effect size).","section":null},{"comment":"§Methods, 'Feature Aggregation': The text states '10 cell types, 5 morphological features and 7 distribution metrics, resulting in a vector of length 350,' but Table 6 lists 10 cell types including 'Other' and 'Tumor.' The total feature count of 470 (350 + 10 + 110) should be reconciled with the cell-type count used.","section":null},{"comment":"Table 3: CD138, PD-L1, and CD56 are excluded due to poor staining quality, but the criteria for exclusion (e.g., signal-to-noise ratio threshold) are qualitative. A brief quantitative criterion or representative images would be useful.","section":null},{"comment":"§Discussion, paragraph 5: The authors note that CD31 is expressed in monocytes and dendritic cells to a lesser extent, which could confound the endothelial cell classification. This limitation should also be acknowledged in the Methods or Results where endothelial interactions are reported.","section":null},{"comment":"The manuscript would benefit from a data and code availability statement. If the pipeline code or processed data are available in a repository, this should be stated; if not, the authors should indicate how the pipeline can be accessed (e.g., upon request, pathologist review).","section":null}],"recommendation":"major_revision","confidential_remarks":"The core pipeline is methodologically reasonable and the descriptive results are plausible and consistent with the DLBCL literature. The major concerns are addressable: multiple-testing correction can be applied post hoc, the unit of analysis can be clarified, and a limited Otsu validation is feasible within the manuscript's scope. I would place the Otsu validation as the most important of the three, since it affects every downstream result. If the authors can show that key findings survive BH correction and patient-level aggregation, the paper should be acceptable for publication."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive reading of our manuscript. The referee identifies three important methodological points: (1) the absence of multiple-testing correction across 470 features tested with Mann-Whitney U tests, (2) potential pseudoreplication arising from multiple tissue cores per patient, and (3) the lack of validation for Otsu-based cell classification thresholds. We agree with all three points and will address each in a revised manuscript. Specifically, we will apply Benjamini-Hochberg FDR correction and report surviving features, clarify the unit of analysis and re-run statistics at the patient level (aggregating cores), and provide validation of Otsu thresholds against expert pathologist annotation on a representative subset of images. We also provide below our honest assessment of what can and cannot be fully resolved within the scope of the current dataset.","responses":[{"response":"The referee is correct. We performed 470 Mann-Whitney U tests without multiple-testing correction, and at a nominal p<0.05 threshold the expected number of false positives is indeed approximately 23. This is a genuine gap in our statistical framework. We will apply Benjamini-Hochberg FDR correction (q<0.05) to all 470 tests and report which features survive correction. We will update Figure 4 and the associated text to reflect corrected p-values, and we will revise all headline claims to reference only features that survive FDR correction. If certain features that we currently highlight do not survive correction, we will state this transparently and adjust our conclusions accordingly. We note that several of our findings—particularly the compositional differences shown in Figure 2C (cell-type proportions between GCB and ABC)—are broadly consistent with the published literature on immune-cold GCB and immune-hot ABC phenotypes, which provides external corroboration independent of our statistical testing. However, we agree that the formal statistical claims must rest on corrected tests.","revision_made":"yes","referee_comment":"Mann-Whitney U tests across 470 features with p<0.05 threshold and no multiple-testing correction; approximately 23 false positives expected. Headline findings drawn from uncorrected tests."},{"response":"The referee raises a valid and important concern. In our current analysis, the unit of analysis is the individual tissue core (sample), not the patient. Multiple cores from the same patient were treated as independent observations. We agree that this constitutes pseudoreplication and could inflate our p-values. In the revised manuscript, we will aggregate features to the patient level by computing the median across cores from the same patient, and re-run all Mann-Whitney U tests (with FDR correction as addressed above) at the patient level (n=106). We will state the unit of analysis explicitly in the Methods. We will also report the number of patients with 1, 2, and 3 cores, respectively, so that the degree of within-patient sampling is transparent. We note that aggregating to patient level will reduce our effective sample size, which may reduce statistical power for some features. We will report which findings persist and which do not, and we will adjust our conclusions accordingly.","revision_made":"yes","referee_comment":"Unit of analysis unclear: 106 patients with up to 3 cores each (559 samples), but statistics refer to 'samples' without clarifying whether cores are treated as independent. Pseudoreplication would inflate p-values."},{"response":"The referee is correct that Otsu's method assumes a bimodal intensity distribution and that this assumption may not hold for all markers in multiplexed fluorescence imaging. We did not validate the Otsu thresholds against expert annotation, and this is a genuine limitation. In the revised manuscript, we will address this in two ways. First, we will provide representative intensity histograms for each marker so that readers can assess the bimodality assumption directly. Second, we will perform a limited validation: a board-certified pathologist (co-author R.B.) will manually annotate cell types on a representative subset of tissue regions (at least 5 cores covering both GCB and ABC subtypes), and we will compute agreement metrics (Cohen's kappa or similar) between manual annotation and our Otsu-based classification. We will report these validation results and discuss markers for which classification agreement is poor. We acknowledge that for markers with genuinely continuous distributions, Otsu thresholding may introduce systematic misclassification, and we will discuss this as a limitation. We note that our cell classification scheme uses combinations of markers (not single markers alone) for most cell types, which provides some robustness to threshold errors on individual markers, but this does not eliminate the concern.","revision_made":"yes","referee_comment":"Otsu thresholding assumes bimodal intensity distributions; fluorescence signals are frequently continuous. No validation against expert annotation or ground truth provided. Misclassification affects all downstream analyses."}],"tokens_in":13509,"tokens_out":1066,"duration_ms":123898,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Two things matter here. First, the pipeline is a legitimate and useful piece of work — t-CyCiF imaging across 106 patients with 14 markers, Cellpose segmentation, marker-based cell classification, proximity graphs, and a 470-feature aggregation per sample. This is a real dataset and a real pipeline, and the specific combination applied to DLBCL TME is new. Second, the formal statistical comparison between GCB and ABC subtypes has two problems that together undermine the 'significantly different' language: no multiple-testing correction across 470 Mann-Whitney U tests, and possible pseudoreplication if multiple cores per patient are treated as independent samples. The stress-test note elevates these correctly; I agree they are more load-bearing than the Otsu thresholding concern, though that is also valid and unaddressed. At p<0.05 across 470 tests, roughly 23 features would pass by chance. The paper highlights the top 20 and draws biological conclusions from them without Bonferroni, Benjamini-Hochberg, or permutation correction. That is the central problem. The pseudoreplication question — 106 patients, up to 3 cores each, 559 total samples, with the unit of analysis described as 'samples' — needs explicit clarification. If cores from the same patient are treated as independent, all p-values are inflated. The descriptive observations in Figures 2C and 3 may hold regardless, because they are presented as proportions and enrichment heatmaps rather than formal tests. The biological findings — immune-rich ABC vs. immune-poor GCB, M2-macrophage interaction differences, CD4/CD8 ratio inversion — are plausible and consistent with prior literature. The discussion is appropriately cautious about several of these points. The Otsu thresholding without ground-truth validation is a real but secondary concern; even with perfect cell classification, the statistical framework cannot support robust significance claims as currently presented. The proximity threshold d* = 0 is chosen qualitatively, which the authors acknowledge. These are fixable issues. The pipeline, the dataset, and the descriptive characterization have value for researchers working on DLBCL microenvironment or building similar imaging pipelines. The paper deserves a serious referee who can push the authors to add multiple-testing correction, clarify the unit of analysis, and validate at least a subset of Otsu thresholds against expert annotation. If those are addressed, this becomes a solid methods-and-characterization paper.","headline":"A well-built multiplexed imaging pipeline for DLBCL TME characterization, but the statistical framework cannot support the significance claims as written.","tokens_in":14596,"tokens_out":582,"would_cite":false,"duration_ms":100741,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"ABC lymphoma tumors host distinct immune neighborhoods","keywords":[],"falsifier":"If the Otsu thresholds do not correspond to biologically meaningful positivity boundaries, the cell-type labels — and therefore all composition, interaction, and morphology-by-cell-type results — could be artifacts of thresholding rather than reflections of genuine TME differences between GCB and ABC tumors.","tokens_in":13665,"feed_emoji":"🔬","tokens_out":670,"duration_ms":156649,"temperature":0.7,"pith_summary":"This paper applies a deep-learning pipeline to multiplexed fluorescence images of 106 diffuse large B-cell lymphoma (DLBCL) patient samples to show that the two main DLBCL subtypes — germinal center B-cell-like (GCB) and activated B-cell-like (ABC), as defined by the standard Hans classifier — have systematically different tumor microenvironments. ABC tumors are enriched in immune cells, including M2-macrophages and CD8+ T-cells, and show preferential spatial interaction between M2-macrophages and tumor cells. GCB tumors, by contrast, are relatively immune-poor, with more B-cells and a different macrophage balance. The pipeline segments cell nuclei, classifies cells by protein-marker positivity, constructs proximity graphs to quantify which cell types sit next to each other, and extracts morphological features — yielding a 470-feature vector per tumor sample. A statistical comparison identifies the morphology of M2-macrophages and CD8+ T-cells, along with specific cell-cell interaction patterns, as the most significant discriminators between the two subtypes.","feed_headline":"ABC lymphoma tumors host distinct immune neighborhoods","feed_subtitle":"Deep-learning analysis of 106 patient samples shows the two DLBCL subtypes differ in immune cell makeup, macrophage interactions, and cell","key_machinery":"The pipeline combines Cellpose-based nuclear segmentation, Otsu-thresholded marker binarization for cell-type classification, centroid-distance proximity graphs for spatial interaction analysis, and distribution-summary morphological features (area, eccentricity, circularity, etc.) aggregated into per-sample feature vectors. The Hans classifier serves as the reference subtype assignment against which all TME features are compared.","core_discovery":"The central finding is that GCB and ABC DLBCL subtypes, classified by the Hans algorithm based on tumor-cell protein markers, also differ sharply in the composition, spatial organization, and morphology of surrounding non-tumor cells. ABC tumors carry a richer immune infiltrate with preferential M2-macrophage–tumor-cell contact, while GCB tumors are comparatively immune-sparse. The morphological features of M2-macrophages and CD8+ T-cells are among the strongest statistical discriminators between the two subtypes, suggesting that single-cell shape characteristics carry subtype-distinguishing information beyond what cell counts alone provide.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Deep learning maps immune cell differences across DLBCL subtypes","ABC and GCB lymphoma tumors show distinct microenvironment layouts","M2-macrophage interactions distinguish ABC from GCB lymphoma","DLBCL subtypes differ in immune infiltrate and cell morphology","Tumor microenvironment features separate DLBCL cell-of-origin subtypes"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The pipeline uses Otsu's thresholding method to binarize each cell's protein-expression signal into positive or negative for each marker. Otsu assumes the signal distribution splits into two clear groups, but fluorescence intensities in multiplexed imaging are often continuous rather than bimodal. If the threshold mislabels cells — for instance, calling a cell CD20-positive when it is not — every downstream cell-type proportion, interaction count, and morphological comparison","fun_headline_variants_meta":{"raw":{"variants":["Deep learning maps immune cell differences across DLBCL subtypes","ABC and GCB lymphoma tumors show distinct microenvironment layouts","M2-macrophage interactions distinguish ABC from GCB lymphoma","DLBCL subtypes differ in immune infiltrate and cell morphology","Tumor microenvironment features separate DLBCL cell-of-origin subtypes"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":746,"prompt_tokens":656,"completion_tokens":90,"prompt_tokens_details":null},"tokens_in":656,"tokens_out":90,"duration_ms":53588,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T15:42:47.712438+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If the Otsu thresholds do not correspond to biologically meaningful positivity boundaries, the cell-type labels — and therefore all composition, interaction, and morphology-by-cell-type results — could be artifacts of thresholding rather than reflections of genuine TME differences between GCB and ABC tumors.","supporting_citations":[],"review_version":1}