{"id":"c14b30bf-d38f-4b46-91cd-1f6ab829b771","arxiv_id":"1908.05667","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Across 23 lung nodules and 320 CT reconstruction settings, thicker slices reduced volumetric reproducibility but improved the reproducibility of histogram and texture features, with no single parameter set optimal for both.","lead":"This study reconstructed 23 lung nodules under 320 combinations of CT dose, reconstruction kernel, and slice thickness, then measured how nodule volume and texture features changed. It found that thicker slices made volume measurements less consistent but texture measurements more consistent, suggesting scan protocols must balance these two goals.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper equates p>0.05 with 'compatible'; without an equivalence margin or per-lesion agreement metric, the central thickness trade-off percentages are uncalibrated.","rationale":"The reader's CONDITIONAL verdict is appropriate. The data collection is a genuine contribution: 23 real nodules, 320 systematically varied reconstructions, expert ROIs, and a broad feature set. The volume results (means, SDs, p-value matrix) directly support the claim that slice thickness increases volumetric variability. However, the texture half of the central trade-off depends on binary labels derived from failure-to-reject t-tests. The specific problems are (1) compatibility is not equivalence, so no clinically meaningful tolerance is enforced; (2) the unpaired test ignores that measurements come from the same nodules; and (3) tens of millions of correlated tests are thresholded without multiplicity control. These do not necessarily overturn the directional conclusion, but they prevent the quantitative percentages from being used as evidence for the strength of the trade-off or for protocol design. I recommend keeping the CONDITIONAL verdict: the authors should replace the p>0.05 compatibility labels with an equivalence/effect-size analysis, or explicitly restrict the claims to the qualitative trends visible in the volume distributions. My concern overlaps with the reader's weakest_assumption but is broader: the root issue is not only multiple comparisons and pairing but the use of significance tests as equivalence tests.","tokens_in":16217,"tokens_out":9673,"duration_ms":109729,"concrete_test":"Re-run the compatibility analysis using two one-sided equivalence tests (TOST) with a pre-specified margin per feature (e.g., ±10% of the mean for volume and ±one within-condition standard deviation for texture features), computing paired within-nodule differences and applying Benjamini-Hochberg FDR at 5% before labeling a pair incompatible. Then regenerate Figs. 7-10 and the best/worst compatibility numbers. If the monotonic increase of texture compatibility with slice thickness disappears or reverses, the central trade-off claim fails; if it persists, the statistical critique affects the reported magnitudes and their interpretation, not the main direction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the definition in Section II.C/E: a reconstruction-condition pair is 'compatible' if a two-tailed t-test fails to reject equality (Eq. 2, with threshold |t|<1.96 / p>0.05). This inverts significance testing. Failure to reject is not evidence of reproducibility: it can mean the true difference is small, the sample (n=23) is too small, or an unpaired test on the same nodules is underpowered. No equivalence margin or effect-size threshold is specified. A feature that shifts by a large but highly variable amount across patients can be labeled compatible, while a small consistent shift can be labeled incompatible. The 65,945,600 tests are also run at unadjusted α=0.05 on highly correlated feature families, so the reported percentages (e.g., 87.45% density compatibility, 24.47% best average) are not calibrated probabilities of reproducibility. Because the abstract's claim that 'as thickness increases, texture reproducibility improves' is quantified solely through these compatibility ratios, the trade-off is not quantitatively established, even though the direction is plausible and the volume analysis is more direct.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a retrospective study of 23 lung nodules whose raw CT data were reconstructed under 320 combinations of radiation dose (4 levels), reconstruction kernel (10), and slice thickness (8). Volumes were measured from 100%-dose/B50f reference segmentations, and 28 histogram- and texture-based features were computed for each of the 320 reconstructions. The authors define two reconstruction conditions as 'compatible' for a feature if an unpaired two-tailed t-test fails to reject equality at the p<0.05 level, and they summarize compatibility averaged over 28 features and 23 patients. Their main claims are that volumetric reproducibility decreases with increasing slice thickness, while radiomic feature reproducibility improves with increasing thickness, and that no single universal CT protocol can simultaneously maximize volumetric and radiomic reproducibility. The paper therefore recommends balanced standardization of acquisition parameters.","tokens_in":16455,"tokens_out":4415,"duration_ms":46252,"significance":"If the quantitative claims are valid, the paper would provide a practically useful map of how common CT protocol variations affect lung-nodule volumetry and radiomics, with direct implications for multi-site study design. The main strengths are the carefully controlled reconstruction grid applied to the same 23 lesions and the direct presentation of raw volume trends in Fig. 5. The paper ships no code or machine-checked proofs, and the statistical framework used to quantify 'compatibility' is not reliable. The qualitative direction of the thickness effect is plausible, but the headline percentages and compatibility maps, which carry the central conclusion, are not currently supported.","major_comments":[{"comment":"The compatibility ratio is defined using the t-test criterion from Eq. (2), but failing to reject equality (p>0.05) is not evidence of reproducibility. With n=23, the test has low power, so a large but highly variable shift can be labeled compatible while a small consistent shift can be labeled incompatible. The reported percentages (e.g., 24.47% best average, 87.45% density under dose changes, 2.65% worst average) are therefore uncalibrated as measures of reproducibility. The analysis should be replaced with an equivalence test using a pre-specified margin (e.g., TOST), or with per-feature effect sizes and confidence intervals.","section":"II.E, Eq. (3)"},{"comment":"The t-test is applied as an unpaired two-sample test to measurements taken on the same 23 nodules. Because all reconstruction conditions are applied to the same lesions, the data are paired; ignoring this structure changes the standard errors and p-values. This affects the volumetric compatibility matrix in Fig. 6 and every radiomic compatibility result in Figs. 4, 7, 8, 9, and 10. A paired test or a mixed-effects model is needed for valid inference.","section":"II.C and II.E, Eq. (2)"},{"comment":"The manuscript reports 65,945,600 (320×320×28×23) hypothesis tests at an uncorrected alpha of 0.05, with no multiple-comparison correction and with highly correlated feature families. Even if the individual tests were valid, the expected number of false positives is enormous, and the reported compatibility percentages are not interpretable as probabilities of reproducibility. The authors should either correct for multiplicity or reframe the analysis as an estimation problem with effect-size summaries rather than significance thresholds.","section":"III.B"},{"comment":"The central claim that 'as thickness increases, volumetric reproducibility decreases, while reproducibility of histogram- and texture-based features ... improves' is quantitatively supported only by the flawed compatibility percentages. The volumetric direction is additionally supported by the raw means and standard deviations in Fig. 5, but the radiomic direction has no such independent support. A reanalysis with equivalence margins may confirm the qualitative direction, but the specific numbers in the abstract and Results should be revised to reflect the corrected analysis.","section":"Abstract and Discussion"}],"minor_comments":[{"comment":"The text states 'If t<1.96 (P<0.05)' but the intended condition is |t|<1.96, which corresponds to p>0.05, i.e., failure to reject equality. Please correct the wording and use p-values consistently.","section":"II.C and II.E"},{"comment":"Equation (2) is missing from the rendered manuscript; only the variables m1, m2, s1, s2, n1, n2 are defined. Equation (3) is also garbled by the rendering and should be rewritten with an explicit denominator (presumably 28×23).","section":"II.C, Eq. (2)"},{"comment":"Section II.D says NGTDM has 3 features, while the abstract says 2 and Table I lists 3 (Coarseness, Complexity, Texture Strength). Please reconcile the counts.","section":"II.D and Table I"},{"comment":"The compatibility map is difficult to read because the color legend and axis ordering are not fully labeled in the figure as reproduced. Please provide an explicit legend and clarify that the diagonal reflects the average over 28 features and 23 patients.","section":"Fig. 4"},{"comment":"The recommendations in the Discussion are derived by data-mining the same 23 cases without external validation. The authors should explicitly label these as exploratory hypotheses rather than validated guidelines.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":"The statistical flaw is serious and load-bearing, but the underlying dataset and reconstruction design are valuable and the qualitative direction of the thickness effect is plausible. I recommend major revision rather than rejection. The manuscript would benefit from collaboration with a statistician to implement paired equivalence testing with pre-specified margins and appropriate multiplicity control."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look for the data they generated, less for the statistical framing. The group reconstructed 23 real nodules across 4 dose levels, 10 kernels, and 8 thicknesses, with expert segmentations applied consistently across conditions. That is real work, and the descriptive volume finding—mean volume drifts downward and variability grows as slices thicken—is clear from the numbers alone. The texture story (thicker slices make histogram/texture features more stable across protocols) is directionally plausible and does appear in their compatibility tables, but that's where the trouble starts.\n\nThe paper defines compatibility as failing to reject a two-tailed t-test at p>0.05, then runs 320×320×28×23 comparisons without any multiple-comparison correction. Failure to reject is not evidence of reproducibility; with n=23 and an unpaired test on paired data, it just means the test was underpowered. A large but noisy shift can look 'compatible,' a small consistent shift 'incompatible.' They also misreport the threshold in the text ('t<1.96 (P<0.05)'), which doesn't help. So the headline percentages—87.45% for density, 24.47% best average—should not be read as calibrated measures of reproducibility. What's left is the qualitative ordering, which may well survive a proper equivalence analysis but isn't established here.\n\nThe paper is honest about its limitations: n=23, single vendor, no external validation, and the recommendations are data-mined from the same cases. That's fair as far as it goes. The main gap is methodological, not a lack of effort or a hidden agenda.\n\nWho gets value? Anyone designing multi-site CT radiomics protocols will want the descriptive volume curve and the general message that slice thickness is the dominant factor and that volume and texture reproducibility pull in opposite directions. They should not quote the compatibility percentages. The paper deserves a serious referee, not a desk reject; a good revision would add an equivalence margin (e.g., ±10% or a standardized effect size), account for the paired structure, and correct for multiple comparisons or report effect sizes instead of raw p-values.\n\nMy call: accept peer review, but flag the statistics as load-bearing. I wouldn't cite the percentages, though I'd cite the volume-thickness trend and the reconstruction grid.","headline":"A labor-intensive CT reconstruction grid produces a plausible thickness trade-off, but the paper's signature percentages rest on p>0.05-as-equivalence and 65 million uncorrected tests.","tokens_in":16968,"tokens_out":2530,"would_cite":false,"duration_ms":25936,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Slice thickness splits CT feature reproducibility: thin slices preserve volume, thick slices stabilize texture, so no single protocol maximizes both.","keywords":["computed tomography","lung nodule","radiomics","texture analysis","slice thickness","reconstruction kernel","radiation dose","reproducibility"],"falsifier":"Recompute the compatibility ratios on the same reconstructed images using a paired comparison that accounts for the repeated measurement of the same 23 nodules and that adjusts for the enormous number of tests being run. If texture features no longer show higher compatibility at thick slices, or if volumetric reproducibility no longer falls as thickness increases, the central trade-off claim fails. A physical phantom with known dimensions and densities, reconstructed through the same 320-condition grid, would separate true measurement error from segmentation variability.","tokens_in":16061,"feed_emoji":"🫁","tokens_out":12740,"duration_ms":116194,"temperature":0.7,"pith_summary":"This paper tests how much lung-nodule measurements change when CT scans are reconstructed at different radiation doses, reconstruction kernels, and slice thicknesses. It claims that slice thickness is the main factor: thinner slices keep nodule volume reproducible, while thicker slices make histogram- and texture-based radiomic features (quantitative statistics of image patterns) more stable across other parameter changes. Because the two trends oppose each other, no single CT protocol can maximize both kinds of reproducibility. If the claim is right, multi-site imaging studies must fix or carefully standardize acquisition parameters—especially slice thickness—before radiomic features can be compared quantitatively.","feed_headline":"Thicker CT slices stabilize radiomics but distort nodule volume","feed_subtitle":"Nodule volume and texture features cannot stay reproducible together when CT parameters change.","key_machinery":"The analysis is carried by a compatibility map. Raw CT data from each nodule were re-reconstructed into 320 conditions (4 doses × 10 kernels × 8 thicknesses); reference regions of interest were segmented once on 100%-dose B50f images at each thickness and then applied to the other 40 dose–kernel combinations at that same thickness. For every pair of reconstruction conditions, a two-tailed $t$-test compared each of 28 image features (histogram, GLCM, RLM, NGLDM, and NGTDM families), and a pair was called compatible when the test gave $t<1.96$ ($p<0.05$). A compatibility ratio (Equation 3) then aggregated these pairwise labels over the 28 features and 23 patients, producing the maps and tables from which the opposing thickness trends are read.","core_discovery":"The paper's central claim is that slice thickness is the dominant driver of quantitative-feature reproducibility in chest CT, and that it moves volume and texture reproducibility in opposite directions. Across 320 acquisition/reconstruction combinations (4 dose levels × 10 kernels × 8 thicknesses) applied to 23 nodules, volumetric reproducibility was best at 2 mm and degraded as slices thickened; 5-mm slices gave the lowest volumes, on average about 4% below the nodule's average volume. Histogram- and texture-based features, by contrast, became more reproducible at thicker slices: the highest average compatibility (24.47% of feature-patient pairs) occurred at 5 mm with the smoothest kernel at 100% dose, and the lowest (2.65%) at 0.6 mm with the sharpest kernel at 12.5% dose. The authors conclude that no universal parameter set keeps both volume and texture reproducible, so multi-study comparability requires balanced standardization of acquisition parameters.","pith_inferences":["If the trade-off generalizes to other scanner families and nodule types, radiomics models trained on heterogeneous clinical CT data carry a hidden confound: site-to-site differences in feature behavior may reflect slice-thickness-dependent reproducibility rather than tumor biology, so slice thickness should be a stratification or adjustment variable.","The paper's 'change only one parameter' advice predicts a directly testable pattern: pairs of reconstruction conditions that differ in exactly one parameter should show systematically higher compatibility ratios than pairs that differ in two or three parameters; that comparison can be quantified from the compatibility map.","A phantom-based rerun of the same 320-condition grid would separate scanner-physics effects from segmentation effects, since the reference ROIs themselves were defined at one dose–kernel setting and may carry some of the observed variability."],"forward_implications":["Serial measurements of the same nodule should keep slice thickness constant, because changing thickness is the strongest single disruptor of histogram- and texture-based feature compatibility.","If a protocol change is unavoidable, change only one parameter—dose, kernel, or thickness—and keep the change minimal; multiple simultaneous parameter changes produce the lowest compatibility.","For volumetric endpoints, thickness near 2 mm is the most reproducible setting, while increasing thickness biases volumes downward (5-mm slices averaged about 4% below the average volume).","A multi-site radiomics study needs a priori protocol harmonization rather than post hoc correction, because compatibility is very limited even across reconstructions of the same raw data."],"supporting_citations":[{"why":"earlier demonstration that radiomics features are not perfectly reproducible across imaging conditions, framing the question tested here","marker":"[5]"},{"why":"phantom pilot showing CT parameter changes alter texture features, motivating the texture-feature analysis","marker":"[6]"},{"why":"prior quantification of dose reduction and reconstruction effects on density and texture features of lung nodules, the direct comparator for radiomic reproducibility","marker":"[7]"},{"why":"characterization of iterative reconstruction performance used to justify including iterative kernels in the grid","marker":"[8]"},{"why":"prior finding that nodule volume is robust to dose and kernel, which the volumetric analysis extends across slice thicknesses","marker":"[20]"},{"why":"review of volumetric assessment variability in CT, supporting volume as a key reproducibility outcome","marker":"[28]"},{"why":"evidence that iterative reconstructions differ from filtered back projection, supporting kernel sharpness as a reproducibility factor","marker":"[40]"},{"why":"source for the over-smoothing mechanism invoked to explain why thick slices improve texture-feature compatibility","marker":"[41]"}],"fun_headline_variants":["Thick slices sharpen radiomics, skew nodule volume","CT slice thickness trades nodule volume for texture stability","Volume and texture CT features can't stay stable together","No CT parameter set preserves both nodule volume and texture","Slice thickness splits CT reproducibility: volume vs texture"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument depends on equating 'no statistically significant difference at the $p<0.05$ level in a two-tailed $t$-test' with 'reproducible,' even though the study performs tens of millions of such tests on measurements taken from the same 23 nodules; if that standard is too permissive, the compatibility percentages and the opposing thickness trends they produce are not a calibrated measure of real reproducibility.","fun_headline_variants_meta":{"raw":{"variants":["Thick slices sharpen radiomics, skew nodule volume","CT slice thickness trades nodule volume for texture stability","Volume and texture CT features can't stay stable together","No CT parameter set preserves both nodule volume and texture","Slice thickness splits CT reproducibility: volume vs texture"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000678,"raw_usage":{"total_tokens":3135,"prompt_tokens":1052,"completion_tokens":2083,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":668,"completion_tokens_details":{"reasoning_tokens":2006}},"tokens_in":668,"tokens_out":2083,"duration_ms":13172,"temperature":1.0,"reasoning_tokens":2006,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:16:59.736281+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the compatibility ratios on the same reconstructed images using a paired comparison that accounts for the repeated measurement of the same 23 nodules and that adjusts for the enormous number of tests being run. If texture features no longer show higher compatibility at thick slices, or if volumetric reproducibility no longer falls as thickness increases, the central trade-off claim fails. A physical phantom with known dimensions and densities, reconstructed through the same 320-condition grid, would separate true measurement error from segmentation variability.","supporting_citations":[{"cited_title":"Reproducibility of radiomics for deciphering tumor phenotype with imaging,","cited_arxiv_id":null,"evidence_quote":"earlier demonstration that radiomics features are not perfectly reproducible across imaging conditions, framing the question tested here"},{"cited_title":"Quantitative assessment of variation in CT parameters on texture features: Pilot study using a nonanatomic phantom,","cited_arxiv_id":null,"evidence_quote":"phantom pilot showing CT parameter changes alter texture features, motivating the texture-feature analysis"},{"cited_title":"Variability in CT lung-nodule quantification: Effects of dose reduction and reconstruction methods on density and texture based features,","cited_arxiv_id":null,"evidence_quote":"prior quantification of dose reduction and reconstruction effects on density and texture features of lung nodules, the direct comparator for radiomic reproducibility"},{"cited_title":"Evaluating iterative reconstruction performance in computed tomography,","cited_arxiv_id":null,"evidence_quote":"characterization of iterative reconstruction performance used to justify including iterative kernels in the grid"},{"cited_title":"Variability in CT lung-nodule volumetry: Effects of dose reduction and reconstruction methods,","cited_arxiv_id":null,"evidence_quote":"prior finding that nodule volume is robust to dose and kernel, which the volumetric analysis extends across slice thicknesses"},{"cited_title":"Noncalcified lung nodules: Volumetric assessment with thoracic CT,","cited_arxiv_id":null,"evidence_quote":"review of volumetric assessment variability in CT, supporting volume as a key reproducibility outcome"},{"cited_title":"Performance of iterative image reconstruction in CT of the paranasal sinuses: A phantom study,","cited_arxiv_id":null,"evidence_quote":"evidence that iterative reconstructions differ from filtered back projection, supporting kernel sharpness as a reproducibility factor"},{"cited_title":"Exploring variability in CT characterization of tumors: A preliminary phantom study,","cited_arxiv_id":null,"evidence_quote":"source for the over-smoothing mechanism invoked to explain why thick slices improve texture-feature compatibility"}],"review_version":1}