{"id":"87ad1282-4163-438e-8131-660f563cf616","arxiv_id":"2411.16961","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A single dynamic-head U-Net segments 14 glomerular tissue and lesion classes across human and rodent kidney slides, with rodent-to-human transfer learning adding about three Dice points on human lesions.","lead":"Glo-In-One-v2 trains one dynamic-head neural network to outline 14 kidney structures and lesions in human and mouse tissue, and reports a 76.5% average Dice score on rodent glomeruli. The authors also show that adding rodent lesion images to human training improves human lesion segmentation by about three Dice points, and they release the code as a Docker-based tool.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 3 and Table 4 rely on inconsistent species/class test sets: ME is scored in the rodent-only table despite having no rodent annotations, and ME is omitted from the human table that is claimed to cover all human lesion classes.","rationale":"Reading the paper in good faith, the contribution is a practical open-source toolkit, a large partially labeled glomerular dataset, and a dynamic-head architecture for multi-class segmentation. The release of code and trained weights is genuine independent support, and the method is a reasonable incremental application of prior dynamic-head ideas. The central numeric claims, however, are the 76.5% average Dice on rodent glomeruli and the +3.3-point human lesion improvement from rodent-augmented training, and these numbers cannot be audited from the tables because the class-to-species and test-set mappings are internally inconsistent. Table 1 is the canonical label distribution and gives ME exclusively to humans, yet Table 3, described as covering the entire rodent dataset, reports ME at 66.2 and folds it into the 76.5% average. Table 4, described as covering all human lesion classes, omits ME entirely, even though ME is the second-most-frequent human lesion annotation. Thus the two headline averages are computed over different, partly impossible, class sets, and no statement about ME's contribution to the transfer gain is possible. The 'more than 3%' claim is also confounded because R&H2H+T adds tissue data on top of rodent and human lesion data, so the gain over H2H mixes two interventions; the cleaner R&H2H comparison gives only +2.5 points. These are fixable reporting and experimental-design problems rather than evidence that the model is fraudulent or that the architecture is unsound. A corrected evaluation on explicitly defined per-class test splits could well confirm the qualitative conclusions, so the appropriate disposition is the same conditional verdict the reader reached: the paper should be accepted only after the reported averages are recomputed and the inconsistencies are resolved.","tokens_in":11938,"tokens_out":6682,"duration_ms":61082,"concrete_test":"Use the released GitHub evaluation code and trained weights, obtaining VUMC data under a data-use agreement if needed, to recompute two numbers: (i) the Table 3 average with ME removed and with ME evaluated only on rodent test patches; (ii) the Table 4 six-class average for H2H, R&H2H, and R&H2H+T after adding ME to the five reported classes. If ME cannot be scored on rodent patches, the 76.5% average is not rodent-only; if adding ME changes the H2H-to-R&H2H+T gap (currently +3.3) by more than 1 Dice point or reverses its sign, the abstract's 'more than 3%' transfer claim is not supported as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline numeric claims depend on two test-set definitions that are internally inconsistent. Table 1 assigns all 227 mesangial-expansion (ME) annotations to the human row and none to rodents, yet Table 3 is introduced as measuring performance 'over the entire rodent dataset' and reports an ME Dice of 66.2; that ME entry is included in the reported 76.5% average. If Table 3 actually used human ME test patches, the average is not rodent-only; if it did not, the ME column lacks a valid test set. Conversely, Table 4 is introduced as presenting 'performance metrics for all lesion classes in the complete human dataset,' but its columns are GS, HS, MA, NS, and SS only; ME, the second-largest human lesion class by annotation count (227), is absent. The abstract's 'more than 3%' transfer gain therefore compares H2H=67.9 with R&H2H+T=71.2 over a five-class subset, and the effect on ME is unmeasured. A second confound compounds the issue: R&H2H+T differs from H2H by adding both rodent lesion data and all tissue-segmentation data, so the +3.3 points cannot be attributed specifically to rodent-to-human transfer; R&H2H without tissue data is only +2.5 over H2H. The central claim is not falsified, but the reported averages are not well-defined for the claimed task until the class and test-set composition is fixed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Glo-In-One-v2, a dynamic-head segmentation network and Dockerized toolkit for holistic intraglomerular analysis. The model is trained on partially labeled human and rodent histopathology patches to segment 14 classes: five tissue/cell classes (Bowman's capsule, tuft, mesangium, mesangial cells, podocytes) and nine lesion classes (adhesion, capsular drop, global sclerosis, hyalinosis, mesangial lysis, microaneurysm, nodular sclerosis, mesangial expansion, segmental sclerosis). The authors report a 76.5% average Dice on rodent glomeruli, outperforming several CNN and transformer baselines, and report that adding rodent data plus tissue data to human lesion training (R&H2H+T) improves human lesion Dice by 3.3 points over human-only training (H2H). The paper also describes a web-mined unannotated image pool and a companion detection module. The central architecture and dataset contributions are plausible, but the headline numeric claims rest on inconsistent test-set definitions that need to be fixed.","tokens_in":12244,"tokens_out":3855,"duration_ms":35534,"significance":"If the reported results survive a corrected evaluation, the paper would make a useful contribution: it tackles a genuinely hard partially labeled, multi-class, cross-species segmentation setting; it provides a large annotated glomerular dataset (23,529 patches) and a public codebase; and its transfer-learning comparison is a genuine held-out evaluation rather than a circular one. The dynamic-head formulation is a reasonable way to handle overlapping and hierarchical classes such as Bowman's capsule containing tuft and mesangium. The manuscript is also honest about data-sharing restrictions. However, the current inconsistencies in which test set produced the reported averages prevent the paper from being accepted as is; the claims need to be re-anchored to a single, coherent class/test-set definition.","major_comments":[{"comment":"Table 1 reports no rodent mesangial-expansion (ME) annotations (rodent ME is listed as '—', while human ME is 227), yet Table 3, introduced as results 'over the entire rodent dataset,' reports an ME Dice of 66.2 and includes it in the 76.5% average. This is internally inconsistent: either a rodent ME test set exists and Table 1 is incomplete, or ME was evaluated on human patches inside a rodent-only table. The authors must state exactly which test patches produce the ME column and recompute the average (and the headline 76.5%) under that definition.","section":"§4.1, Table 1; §4.3.1, Table 3"},{"comment":"The paragraph introducing Table 4 says it presents 'performance metrics for all lesion classes in the complete human dataset,' but the table has only five lesion columns (GS, HS, MA, NS, SS) and omits ME, the second-largest human lesion class by annotation count (227, per Table 1). Consequently, the abstract's 'more than 3%' improvement from 67.9 (H2H) to 71.2 (R&H2H+T) is measured on five classes only, not on all nine human lesion classes, and the effect of transfer on ME is unmeasured. Please add ME with a clear test-set definition, or revise the claim to state explicitly that it applies to the five-class subset.","section":"§4.3.2, Table 4"},{"comment":"The R&H2H+T condition differs from H2H by adding both rodent lesion data and all tissue-segmentation data, so the +3.3 Dice gain cannot be attributed specifically to rodent-to-human transfer; the paper's own numbers show that R&H2H (lesion data only) gives +2.5 over H2H, implying the tissue data adds roughly 0.8 Dice points. This is a confounded comparison for the stated transfer-learning conclusion. Please add a controlled ablation and report uncertainty (multiple seeds with means and confidence intervals) before claiming a transfer-driven improvement.","section":"§4.2, §4.3.2"},{"comment":"Several baseline entries are at or near 50% Dice for lesion classes (e.g., U-Nets AH 50.2 and CD 50.8; Multi-class ME 49.1, MA 49.3, and ML 49.6; Segmenter MA 55.5 and SS 50.2), a pattern consistent with degenerate predictions that always predict background or a single class. With many classes near that level, the statement that the dynamic-head model 'outperforms all baselines' is not yet a meaningful comparison. Please report per-class precision/recall or the proportion of non-empty predictions, and characterize or exclude degenerate baseline runs.","section":"§4.3.1, Table 3"}],"minor_comments":[{"comment":"The task vector is called m-dimensional and the text refers to 'the i th class of lesion,' but the ordering of the 14 classes is never defined; please specify the class ordering used for the one-hot encoding.","section":"§3.1, Eq. (1)"},{"comment":"The image-pool and training descriptions say the pool size matches the number of classes and that the best model is chosen by average Dice over 200 epochs, but no hyperparameters (learning rate, optimizer, batch size, augmentation, epochs) are reported; a reproducibility appendix with these values is needed.","section":"§4.2"},{"comment":"The abstract emphasizes the large dataset, but the Data Availability section states that a portion of the data cannot be made public; please clarify how much of the 23,529-patch dataset is publicly accessible and under what terms.","section":"Data Availability"},{"comment":"In the Swinunetr row, '73.3-' appears instead of a Dice value for HS; also, Table 5 is referenced in §4.3.2 as being in the appendix but it appears after §6, so the numbering/placement should be adjusted.","section":"Table 5"},{"comment":"There are multiple typos and grammatical errors that should be corrected: 'glomeular' in the keywords, 'reprensents' in §3.1, 'Adhension' in Fig. 4, 'hybird' and 'stragety' in Fig. 6, and 'of nine require' in §2.","section":"Throughout"},{"comment":"The web-mining subsection describes collecting 10,000 compound figures and obtaining over 30,000 unannotated glomerular images, but the paper never reports an experiment using these images; please state whether they were used in pretraining or contrastive learning, or remove the subsection if they are not part of the pipeline.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a systems/dataset contribution, and the code release is valuable, but the main dataset is not publicly available despite the abstract's emphasis on scale. The inconsistencies between Tables 1, 3, and 4 are the key barrier to acceptance; they must be resolved before the headline numbers (76.5% average Dice and the +3.3-point transfer gain) can be considered reliable. I would also encourage the editor to ask for a version of the evaluation with confidence intervals or multiple seeds, given the large number of near-degenerate baseline entries."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline on this one is: the dataset and the containerized tool are the contribution, not the architecture. The dynamic head comes straight from the authors' own single-dynamic-network paper (ref 25) and DoDNet (ref 27), so if you are hoping for a new method, you will be disappointed. What is genuinely useful: a curated, partially labeled set of 23,529 glomerular patches covering 14 tissue/cell/lesion classes across human and mouse, and a Dockerized pipeline that runs end-to-end from WSI to multi-channel masks. The cross-species transfer experiment (rodent data as auxiliary for human lesion segmentation) is a reasonable idea and not something I have seen evaluated at this scale.\n\nNow the soft spots. The stress-test note is right: Tables 3 and 4 are internally inconsistent. Table 1 gives mesangial expansion (ME) zero rodent annotations, yet Table 3 is introduced as rodent-only and reports an ME Dice of 66.2, folding that number into the 76.5% average. Table 4 claims to cover all human lesion classes but lists only GS, HS, MA, NS, SS; ME, which has 227 human annotations, is absent. So the two headline numbers — 76.5% and the 'more than 3%' transfer gain — are not well-defined until the species/class composition is fixed. On top of that, R&H2H+T differs from H2H by adding both rodent lesions and all tissue data, so the +3.3 points cannot be chalked up to rodent-to-human transfer; the lesion-only R&H2H is +2.5, which is a weaker but more honest claim. And from a methodology standpoint, the comparisons have no confidence intervals, no repeated seeds, and several baselines sit at Dice ~50%, which suggests the evaluation protocol may be too forgiving or the classes are very small.\n\nNone of this kills the paper. The core idea — a single dynamic-head model trained on partially labeled multi-species data — is plausible, and the release of a working containerized tool with a large annotated dataset is a real service to the nephropathology imaging community.\n\nWho should read it: anyone building glomerular segmentation pipelines or looking for a ready-made tool; anyone doing cross-species transfer in pathology. It is not a methods breakthrough, and the numbers need to be re-reported cleanly before I would trust the headline claims. This deserves peer review, but with major revision: fix the table definitions, report uncertainty, and untangle the ablation.\n\nRecommendation: send to review, but the authors need to clean up the test-set definitions before it is acceptable.","headline":"Large useful glomerular dataset and a working containerized tool, but the headline numbers rest on inconsistent test-set definitions that need fixing.","tokens_in":2,"tokens_out":3180,"would_cite":true,"duration_ms":59119,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One model maps 14 glomerular classes in mice and humans","keywords":["glomerular segmentation","whole-slide image","renal pathology","lesion segmentation","partially labeled","dynamic head","transfer learning","open-source toolkit"],"falsifier":"Re-run the released Glo-In-One-v2 model on the test set while reassigning the 227 mesangial-expansion annotations from human to rodent (or vice versa, per Table 1) and recompute the average Dice; if the rodent average drops below 76.5% or the human transfer gain vanishes, the central claims are not reproducible.","tokens_in":11745,"feed_emoji":"🔬","tokens_out":8327,"duration_ms":67232,"temperature":0.7,"pith_summary":"This paper claims that a single deep-learning model with a class-aware dynamic head can segment all 14 intraglomerular tissue and lesion classes — five tissues (Bowman's capsule, tuft, mesangium, mesangial cells, podocytes) and nine lesions (adhesion, capsular drop, global sclerosis, hyalinosis, mesangial lysis, microaneurysm, nodular sclerosis, mesangial expansion, segmental sclerosis). On a curated set of 23,529 annotated glomeruli from 368 whole-slide images, the model reaches a 76.5% average Dice on rodent glomeruli, outperforming per-class U-Net, DeepLabv3, transformer, and multi-class baselines. For human lesions, the paper reports that adding rodent lesion data to training (R&H2H) raises average Dice from 67.9% to 70.4%, and further adding the tissue data (R&H2H+T) brings it to 71.2% — a 3.3-point gain over human-only training. If these results hold, the open-source tool would give nephropathologists a single-command, cross-species pipeline for fine-grained glomerular quantification.","feed_headline":"One model maps 14 glomerular classes in mice and humans","feed_subtitle":"Rodent training data lifts human lesion Dice from 67.9 to 71.2 percent, a 3.3-point gain.","key_machinery":"The central mechanism is the class-aware dynamic head: a residual U-Net backbone extracts image features, a global-average-pooled representation is concatenated with a one-hot task encoding for the target class, and a small convolutional controller generates the three convolution kernels of a lightweight head that predicts that class's mask from the decoder features. This lets a single model serve all 14 classes under partial labeling, and is what the paper claims lets it handle hierarchical relationships such as Bowman's capsule, tuft, and mesangium being nested regions.","core_discovery":"On its own terms, the paper demonstrates that a single residual U-Net with a dynamic, class-aware head can jointly handle 14 segmentation classes even when each training image is annotated for only one class (partially labeled data). The task is encoded as a one-hot vector that is concatenated with the global-pooled image features; a small controller produces the convolution kernels for the head, so one network can output a mask for any requested class. Across the rodent test set the model achieves a 76.5% average Dice, with per-class scores from 57.1% (mesangial lysis) to 96.3% (Bowman's capsule). In the rodent-to-human transfer experiments, the hybrid training strategy R&H2H (rodent + human lesion data) yields 70.4% average human lesion Dice, and R&H2H+T (adding all tissue data) yields 71.2%, compared with 67.9% for human-only H2H training. The paper attributes the gains to the dynamic head's ability to reason about overlapping and subset/superset classes, and releases the toolkit as a Docker container.","pith_inferences":["A natural extension would be to check whether adding rodent tissue data helps human tissue segmentation as much as it helps lesion segmentation, a comparison the paper does not report.","The transfer gain is reported only as an average; disaggregating by lesion class would show whether the +3.3 points is driven by one class (e.g., nodular or segmental sclerosis) or is uniform.","The same dynamic-head scheme could be applied to other partially labeled multi-class pathology problems, such as tubular, vascular, and interstitial compartments in the same kidney sections.","Because the paper releases the model and weights, the reproducibility of the headline numbers can be tested directly against the curated test sets, including the mesangial-expansion annotations whose species assignment is unclear from the tables."],"forward_implications":["Clinicians can run the released Docker toolkit on raw whole-slide images and get per-class masks for all 14 glomerular structures and lesions in a single command.","Rodent histopathology can serve as auxiliary training data to improve human lesion segmentation when human annotations are scarce, with a reported 3.3-point Dice gain.","The dynamic-head design outperforms multi-head baselines on classes with anatomical overlap, such as global sclerosis covering the Bowman's capsule region.","The curated dataset of 23,529 partially labeled glomeruli, split at patient level, provides a shared benchmark for cross-species intraglomerular segmentation."],"supporting_citations":[{"why":"Prior Glo-In-One toolkit; provides the glomerular detection module and the web-mined dataset that Glo-In-One-v2 builds on.","marker":"[3]"},{"why":"Single dynamic network for multi-label renal pathology segmentation; the source of the dynamic-head architecture used here.","marker":"[25]"},{"why":"DoDNet dynamic filter generation; the mechanism that produces class-specific convolution kernels from features and task encoding.","marker":"[27]"},{"why":"KPMP spatial segmentation dataset; supplies the human glomerular annotations that complement the manually curated rodent/human patches.","marker":"[30]"},{"why":"Multi-structure segmentation from partially labeled datasets; defines the partially-labeled training setting and serves as a key baseline.","marker":"[34]"}],"fun_headline_variants":["One dynamic head segments 14 glomerular labels across species","Rodent data lifts human lesion Dice by 3.3 points","Glo-In-One-v2: 14 classes from partially labeled WSIs","Single net, 14 classes, human and mouse: Glo-In-One-v2","Transfer learning boosts lesion segmentation across species"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline Dice averages assume the data tables correctly assign every annotation to human or rodent; if mesangial-expansion annotations are placed in the wrong species column, the 76.5% and 3.3-point figures would not describe the claimed 14-class model.","fun_headline_variants_meta":{"raw":{"variants":["One dynamic head segments 14 glomerular labels across species","Rodent data lifts human lesion Dice by 3.3 points","Glo-In-One-v2: 14 classes from partially labeled WSIs","Single net, 14 classes, human and mouse: Glo-In-One-v2","Transfer learning boosts lesion segmentation across species"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000378,"raw_usage":{"total_tokens":2104,"prompt_tokens":1131,"completion_tokens":973,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":747,"completion_tokens_details":{"reasoning_tokens":884}},"tokens_in":747,"tokens_out":973,"duration_ms":8302,"temperature":1.0,"reasoning_tokens":884,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:41:50.037609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the released Glo-In-One-v2 model on the test set while reassigning the 227 mesangial-expansion annotations from human to rodent (or vice versa, per Table 1) and recompute the average Dice; if the rodent average drops below 76.5% or the human transfer gain vanishes, the central claims are not reproducible.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior Glo-In-One toolkit; provides the glomerular detection module and the web-mined dataset that Glo-In-One-v2 builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Single dynamic network for multi-label renal pathology segmentation; the source of the dynamic-head architecture used here."},{"cited_title":"Zhang, Y","cited_arxiv_id":null,"evidence_quote":"DoDNet dynamic filter generation; the mechanism that produces class-specific convolution kernels from features and task encoding."},{"cited_title":"The results here are in whole or part based upon data generated by the Kidney Precision Medicine Project","cited_arxiv_id":null,"evidence_quote":"KPMP spatial segmentation dataset; supplies the human glomerular annotations that complement the manually curated rodent/human patches."},{"cited_title":"Gonz \\'a lez, G","cited_arxiv_id":null,"evidence_quote":"Multi-structure segmentation from partially labeled datasets; defines the partially-labeled training setting and serves as a key baseline."}],"review_version":1}