{"id":"3be4556e-c50b-43a7-b541-896b40fbfc3f","arxiv_id":"1908.08746","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"RatLesNet, a compact 3D fully convolutional network, segments rodent brain stroke lesions from T2-weighted MRI with an average Dice of 0.88 in cross-validation and 0.79 across studies.","lead":"This paper introduces RatLesNet, a deep-learning network that automatically outlines stroke lesions in rat brain MRI scans. It reports better accuracy than two existing 3D networks and suggests a practical way to automate a time-consuming step in pre-clinical research.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significance testing: RatLesNet's advantage over VoxResNet is small, overlaps reported variability, and reverses on one generalization study; the central superiority claim is not yet established.","rationale":"The reader's verdict is CONDITIONAL, and this stress-test supports that verdict rather than moving it. The concern raised here is about the statistical support for the central comparison with VoxResNet, which the reader mentioned in the rationale ('no statistical significance tests') but did not make the weakest assumption. The headline difference is small and within one standard deviation of the reported per-scan Dice, and one of three generalization studies favors VoxResNet, so the claim 'outperforms VoxResNet' is not yet established. However, this is an addressable evidence gap, not an internal inconsistency: the reported numbers are internally plausible, the lesioned-only and with-sham results are both reported transparently, and the cross-study failure on 03AUG2015 is acknowledged in the text. A paired per-animal significance analysis would settle whether the central claim survives. Until that is provided, CONDITIONAL remains the appropriate verdict; ACCEPT would be premature and REJECT would be too harsh given the consistent direction of the remaining comparisons.","tokens_in":6244,"tokens_out":12032,"duration_ms":122460,"concrete_test":"Obtain per-animal Dice predictions for the five cross-validation folds and the three generalization studies; run a paired Wilcoxon signed-rank test (or a 10,000-sample bootstrap) comparing RatLesNet with VoxResNet on the lesioned-only 02NOV2016 scans and on each generalization study, reporting 95% CIs and effect sizes. If the CV comparison is not significant (p>0.05) or the D35 study significantly favors VoxResNet, the paper should downgrade the claim of general superiority.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that RatLesNet outperforms 3D U-Net and VoxResNet for rodent lesion segmentation, with average Dice 0.88 in cross-validation and 0.79 in cross-study generalization. The load-bearing comparison is RatLesNet vs VoxResNet: Table 1 gives 0.88 vs 0.85 including sham animals and 0.76 vs 0.70 for lesioned-only scans; per-timepoint differences are 0.06-0.07, comparable to the reported standard deviations (e.g., 2h no-sham 0.67 vs 0.60; 24h no-sham 0.85 vs 0.79). No statistical test, confidence interval, or per-animal paired analysis is provided, so it is unknown whether these differences exceed chance. The generalization result is also mixed: VoxResNet achieves the highest Dice on 03AUG2015 (0.71 vs 0.68, Table 2), so the claim of overall superiority rests on averaging across studies. If the cross-validation advantage over VoxResNet is not significant on lesioned-only scans, the central claim of superiority is unsupported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RatLesNet, a compact 3D fully convolutional network (0.37M parameters) for automatic segmentation of ischemic lesions in T2-weighted rat brain MRI. It is trained and evaluated on 131 scans from four tMCAO studies. In a five-fold cross-validation experiment on one study (02NOV2016), the authors report an average Dice of 0.88, exceeding 3D U-Net (0.64) and VoxResNet (0.85) when all animals are included, and 0.76 vs 0.70 and 0.30 respectively when sham animals are excluded. In a cross-study generalization test, training on 02NOV2016 and testing on the other three studies yields average Dice of 0.79 for RatLesNet, 0.76 for VoxResNet, and 0.62 for 3D U-Net. The authors also claim that RatLesNet exceeds inter-operator variability (average Dice 0.87 including shams). The paper emphasizes that the method requires no skull-stripping, bias-field correction, or post-processing.","tokens_in":6423,"tokens_out":5834,"duration_ms":57746,"significance":"If the reported performance holds, RatLesNet would be a practically valuable tool for preclinical stroke research, where manual segmentation is time-consuming and subjective. The evaluation design is a strength: it includes a multi-study dataset, a realistic cross-study generalization test, and comparisons against two strong 3D FCN baselines. The architecture's small parameter count and ability to process entire volumes without preprocessing are also attractive for reproducibility and clinical translation. However, the central claim of superiority over VoxResNet rests on small average Dice differences with overlapping standard deviations and no statistical testing, and the cross-study result reverses on one of the three test studies. The evidence is promising but not yet conclusive.","major_comments":[{"comment":"No statistical significance testing or confidence intervals are provided for the comparison between RatLesNet and VoxResNet. The differences are small relative to the reported variability: for example, at 2h without shams, RatLesNet gives 0.67±0.12 vs VoxResNet 0.60±0.16; at 24h without shams, 0.85±0.11 vs 0.79±0.20; and on average without shams, 0.76±0.14 vs 0.70±0.20. Since the same test folds are used for all methods, a paired per-animal test (e.g., Wilcoxon signed-rank test) or bootstrap confidence intervals on the paired differences is needed to support the abstract's claim that RatLesNet is 'quantitatively better' than the compared architectures.","section":"Section 3, Table 1 (Cross-validation)"},{"comment":"The statement that RatLesNet 'achieves higher Dice coefficients than inter-operator variability' is not supported. The average Dice including sham animals is 0.88 for RatLesNet vs 0.87 for inter-operator agreement, a difference of 0.01; excluding shams it is 0.76 vs 0.73. Moreover, the inter-operator benchmark is based on only one additional manual segmentation of a single study, and RatLesNet is trained on the reference annotator's labels. Exceeding that benchmark may therefore reflect learning the reference annotator's segmentation style rather than improved biological accuracy. The authors should qualify this claim and, if possible, evaluate against multiple independent raters.","section":"Section 2 and Section 3, Table 1 (Inter-operator comparison)"},{"comment":"The cross-study generalization claim is not robust across studies. RatLesNet outperforms VoxResNet by large margins on 03MAY2016 (0.82 vs 0.77) and 02OCT2017 (0.84 vs 0.78), but on 03AUG2015 VoxResNet achieves the highest Dice (0.71±0.23 vs 0.68±0.26 for RatLesNet). Because no significance testing is provided, the overall claim that RatLesNet 'outperformed the other FCNs at generalizing' rests on averaging over studies and may depend on the particular composition of the test set. The authors should report per-study pairwise statistical comparisons and discuss the time-point difference (35 days for 03AUG2015 vs 2h/24h in the training data) as a likely source of the reversal.","section":"Section 3, Table 2 (Generalization capability)"},{"comment":"The sentence 'The models that provided the reported results were trained with the best performing learning rate found' is ambiguous and potentially consequential. If the learning rate was selected using the held-out test studies, the generalization results would be invalid. If it was selected using the validation folds inside the cross-validation procedure, this should be stated explicitly. Please clarify the model-selection protocol.","section":"Section 3, Cross-validation paragraph"}],"minor_comments":[{"comment":"The phrase 'between 3.7% and 38% higher' is ambiguous because Dice improvements can be reported as absolute percentage points or relative percentages. Please specify which convention is used.","section":"Abstract"},{"comment":"The inter-operator Dice of 0.73 reported in the text is inconsistent with the 'Inter-operator' column in Table 1, which shows an average of 0.87 when shams are included. Please clarify how sham animals (no lesion) enter the inter-operator Dice computation, e.g., whether they are assigned Dice 1 by convention.","section":"Section 2 and Table 1"},{"comment":"The text states that Adam was used 'with a starting learning rate of 10^-5', but later the models are said to be trained with the 'best performing learning rate found'. Please state the final learning rate(s) used and the search procedure.","section":"Section 2, Training"},{"comment":"The architecture diagram is difficult to interpret: the number of channels at each stage and the positions of the 1x1 convolution bottleneck layers are not annotated. Adding these details would improve reproducibility.","section":"Figure 2"},{"comment":"The phrase 'crispier' in the text describing VoxResNet's predictions is informal; consider using 'sharper' or 'more clearly defined'.","section":"Section 3, Figure 3 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript describes a useful and well-motivated application, and the architectural choices are sensible. The main obstacle to acceptance is the lack of statistical evidence for the claimed superiority over VoxResNet. Paired significance tests, confidence intervals, and a clearer model-selection statement would address the central concern. The inter-operator comparison should be framed as a benchmark against a single extra rater, not as evidence of biological accuracy. These are fixable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a genuinely useful practical paper, not a methodological breakthrough. The authors train a compact 3D fully convolutional network on 131 T2-weighted rat brain scans from four studies and compare it against 3D U-Net and VoxResNet. RatLesNet is small (0.37M parameters), needs no preprocessing or postprocessing, and runs in half a second per scan. That's a real contribution to preclinical stroke research, where manual tracing is the bottleneck.\n\nThe architecture is a sensible combination of dense blocks, 1x1 bottleneck convolutions, and max-pooling/unpooling with index reuse. Nothing conceptually new, but the design is well-motivated for small datasets. The evaluation is more thorough than most in this niche: five-fold cross-validation on one study plus a cross-study generalization test on three others, with both lesion-only and sham-inclusive Dice scores. The authors are also honest about the one generalization study where VoxResNet wins.\n\nThe soft spot is the statistics. No significance tests, no confidence intervals, no per-animal paired analysis. The average cross-validation advantage over VoxResNet is 0.88 vs 0.85 with shams and 0.76 vs 0.70 without, and the per-timepoint gaps (0.06-0.07) sit right on top of the standard deviations. So the central claim that RatLesNet outperforms the generic networks is plausible but not fully established. The comparison to inter-operator variability is also weaker than it looks: the network was trained on the same reference labels it is being judged against, and the 0.01 average gap over one independent segmentation is not compelling. The lack of code and data makes reproducibility checks harder, though training details are reported.\n\nNone of this sinks the paper. The core message—that a small 3D FCN can do practical rodent lesion segmentation and generalize across studies—holds up. It would benefit from a paired significance test and ideally a release of the code and trained model. I'd send it to peer review; it deserves careful refereeing, but reviewers should ask for the missing statistical evidence.","headline":"A compact 3D FCN for rat lesion segmentation that is genuinely useful; the main superiority claim needs significance testing.","tokens_in":6995,"tokens_out":1924,"would_cite":true,"duration_ms":19228,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"RatLesNet, a 3D fully convolutional network with 0.37M parameters, segments ischemic stroke lesions in rat brain MRI with an average Dice of 0.88 in five-fold cross-validation, outperforming 3D U-Net and VoxResNet, and reaches 0.79 in…","keywords":["RatLesNet","fully convolutional network","lesion segmentation","rat brain MRI","ischemic stroke","3D segmentation","Dice coefficient","cross-study generalization"],"falsifier":"Re-run the five-fold cross-validation using a second independent operator's manual segmentations as ground truth; if RatLesNet's average Dice falls to the measured inter-operator level of about 0.73, the reported 0.88 reflects the reference annotator's style rather than true lesion boundaries.","tokens_in":6028,"feed_emoji":"🧠","tokens_out":11750,"duration_ms":93831,"temperature":0.7,"pith_summary":"RatLesNet is a 3D fully convolutional network, a neural network that maps an entire 3D image directly to per-voxel labels, built specifically for segmenting stroke lesions from rat brain MRI. The paper claims that this small network—0.37 million parameters—beats two general-purpose 3D segmentation networks, 3D U-Net and VoxResNet, on the task: average Dice 0.88 versus 0.85 and 0.64 in five-fold cross-validation, and 0.79 versus 0.76 and 0.62 when trained on one study and tested on three others. A sympathetic reader would care because manual lesion tracing is slow, subjective, and a bottleneck in preclinical stroke research, and an automatic method that runs on unprocessed images in about half a second per scan could make large cohort studies routine. The paper also reports that RatLesNet's average Dice exceeds the measured inter-operator agreement between two human segmentations, and that its boundaries look qualitatively cleaner around the lesion.","feed_headline":"Tiny 3D network segments rat brain lesions at Dice 0.88","feed_subtitle":"RatLesNet beats larger 3D U-Net and VoxResNet using raw MRI and no post-processing.","key_machinery":"The argument is carried by RatLesNet's dense-block architecture: within each block, every $3\\times3\\times3$ convolution is concatenated with all preceding feature maps in that block (growth rate 18), $1\\times1\\times1$ convolutions reset the channel count between blocks, and max-pooling with index-preserving unpooling halves and restores spatial dimensions. This lets the network see the full 3D volume while keeping parameters at 0.37M, which the paper says reduces overfitting relative to 3D U-Net's 19M and VoxResNet's 1.5M parameters. Training uses raw standardized T2 volumes with a mini-batch of 1, cross-entropy loss, and early stopping on validation loss; inference takes about half a second per scan.","core_discovery":"The central discovery is that a compact dense-block architecture designed for the specific geometry of rodent stroke lesions outperforms larger networks built for anatomical segmentation, despite making no use of preprocessing, skull-stripping, or post-processing. RatLesNet processes each whole $256\\times256\\times18$ volume with $1\\times1\\times1$ convolutions for channel adjustment, dense blocks with growth rate 18, and max-pooling/unpooling that reuses pooling indices to restore spatial resolution. In five-fold cross-validation on the 02NOV2016 study, the paper reports an average Dice of 0.88 (0.76 excluding lesion-free sham animals), above the inter-operator Dice of 0.87 (0.73 excluding shams); on generalization to three other studies it averages 0.79, with the best scores on 24-hour post-stroke scans and a lower 0.68 on the 35-day study, where VoxResNet scores 0.71. The paper frames this as the first fully convolutional network specifically designed for rat brain lesion segmentation, with an architecture that accepts any input-channel count and therefore extends to multimodal MRI.","pith_inferences":["The paper suggests the parameter gap explains the accuracy gap, but it does not isolate capacity from architecture; a controlled comparison that matches parameter budgets across designs would test that explanation.","Because the network is trained against a single operator's labels and evaluated on that same reference, the 0.88 Dice may partly reflect learning that operator's delineation style; using an ensemble of expert segmentations as training targets would probe this.","The dense-block plus $1\\times1\\times1$ bottleneck pattern that works here for rare lesion voxels is a candidate for other small-lesion or small-animal segmentation tasks where datasets are too small for large 3D networks.","Adding chronic-stage training examples and a scanner-matched validation set would likely raise the 35-day study's Dice above VoxResNet's 0.71, since the training study contains only 2-hour and 24-hour lesions."],"forward_implications":["Preclinical stroke pipelines can replace manual tracing with fully automatic RatLesNet segmentation on raw T2-weighted MRI, eliminating the skull-stripping, bias correction, and hole-filling steps that earlier rodent lesion tools require.","A whole-volume 3D network with 0.37M parameters trains in about six hours on one GPU and segments a scan in about half a second, making segmentation of hundreds of animals practical.","Because the input channel count is configurable, the same architecture can be applied to multimodal or multi-sequence MRI without structural changes.","On 24-hour post-stroke lesions from unseen studies, RatLesNet's average Dice (0.82 and 0.84) exceeds the measured inter-operator agreement of 0.73, indicating automatic segmentations can be at least as reproducible as human ones for that stage.","The lower Dice of 0.68 on the 35-day study shows the generalization is lesion-stage dependent; adding chronic-stage lesions to training is the direct extension the paper acknowledges."],"supporting_citations":[{"why":"Supplies the original U-Net architecture that the compared 3D U-Net and the dense-block design build on.","marker":"[16]"},{"why":"One of the two comparison baselines; its 19M-parameter 3D U-Net is the architecture RatLesNet is compared against.","marker":"[5]"},{"why":"The other comparison baseline; its 1.5M-parameter VoxResNet defines the anatomical-segmentation approach the paper compares against.","marker":"[3]"},{"why":"Provides the dense connection pattern used inside RatLesNet's dense blocks.","marker":"[10]"},{"why":"Supplies the residual-style concatenation idea adapted by the blocks.","marker":"[9]"},{"why":"Supplies the max-pooling index reuse mechanism used for unpooling in RatLesNet.","marker":"[15]"},{"why":"Defines the Dice coefficient used to score every segmentation result.","marker":"[7]"},{"why":"Describes the transient middle cerebral artery occlusion model that generated the lesion data.","marker":"[12]"}],"fun_headline_variants":["RatLesNet: first FCN for rat brain lesion MRI","Compact net beats U-Net on rodent stroke lesions","Deep learning scores 0.88 Dice on rat MRI lesions","New FCN tops 3D U-Net for rat lesion segmentation","Rat brain lesion auto-segmentation hits 0.88 Dice"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The manual segmentations used for training and evaluation are treated as ground truth, and the four studies are assumed similar enough to transfer; if the reference labels carry one annotator's bias or the studies differ too much, the reported Dice gaps do not show true segmentation accuracy.","fun_headline_variants_meta":{"raw":{"variants":["RatLesNet: first FCN for rat brain lesion MRI","Compact net beats U-Net on rodent stroke lesions","Deep learning scores 0.88 Dice on rat MRI lesions","New FCN tops 3D U-Net for rat lesion segmentation","Rat brain lesion auto-segmentation hits 0.88 Dice"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000176,"raw_usage":{"total_tokens":1351,"prompt_tokens":1066,"completion_tokens":285,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":198}},"tokens_in":682,"tokens_out":285,"duration_ms":3390,"temperature":1.0,"reasoning_tokens":198,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:30:14.587349+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the five-fold cross-validation using a second independent operator's manual segmentations as ground truth; if RatLesNet's average Dice falls to the measured inter-operator level of about 0.73, the reported 0.88 reflects the reference annotator's style rather than true lesion boundaries.","supporting_citations":[{"cited_title":"In: International conference on medical image computing and computer-assisted intervention","cited_arxiv_id":null,"evidence_quote":"One of the two comparison baselines; its 19M-parameter 3D U-Net is the architecture RatLesNet is compared against."},{"cited_title":"NeuroImage 170, 446–455 (2018)","cited_arxiv_id":null,"evidence_quote":"The other comparison baseline; its 1.5M-parameter VoxResNet defines the anatomical-segmentation approach the paper compares against."},{"cited_title":"Ecology 26(3), 297–302 (1945)","cited_arxiv_id":null,"evidence_quote":"Defines the Dice coefficient used to score every segmentation result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the transient middle cerebral artery occlusion model that generated the lesion data."}],"review_version":1}