{"id":"d9c49c69-c063-42b2-a124-e441b3ea1e90","arxiv_id":"2508.21398","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Releases a region-annotated laparoscopy image dataset for endometriosis with 25,682 frames and 520 expert annotations across four pathology classes.","lead":"GLENDA is a new image dataset for endometriosis in gynecologic laparoscopy, with over 25,000 frames and 520 region annotations drawn by medical experts. It is an endometriosis-specific annotated resource aimed at training computer vision models for detection, classification, and localization.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annotation reliability is unvalidated: no inter-annotator agreement, adjudication, or histologic confirmation for the 520 regions, and the paper itself notes experts often disagree; keyframe-to-sequence propagation relies on unquantified minimal camera motion.","rationale":"The reader's weakest assumption—annotation correctness and the unverified keyframe-to-sequence validity—is the most load-bearing point because it directly affects the dataset's value for every stated purpose: classification, detection, localization, and tracking. The paper gives no quantitative evidence for either the accuracy of the regions or the temporal validity of keyframe annotations. I considered whether the 'first of its kind' novelty claim or the accessibility of the dataset URL should be the primary concern, but those are either literature claims or practical logistics rather than correctness risks. The annotation issue is both internally flagged and externally plausible, since visual endometriosis diagnosis is known to be difficult even among specialists, as the authors themselves concede. The proposed concrete test would settle the concern in one validation study. Since the reader already assigned a conditional verdict, my stress test does not change that verdict, but it sharpens the condition under which acceptance should be granted.","tokens_in":7298,"tokens_out":3687,"duration_ms":40319,"concrete_test":"Ask the authors to run and report an inter-annotator validation: sample 50 annotated frames stratified by class (e.g., 20 peritoneum, 10 ovary, 10 DIE, 10 uterus) and have two independent expert endometriosis surgeons re-annotate them with the same ECAT tool; report per-class pixel IoU and frame-level Cohen's/Fleiss' kappa. In the same pass, take 10 random sequences per positive class, have an expert annotate the final frame, and compute IoU between the keyframe mask and the final-frame mask after optical-flow/affine registration. If IoU is below 0.5 or kappa below 0.6 on any class, the ground truth and the 'visible throughout segment' assumption are not supported, and the paper should add confidence/uncertainty labels rather than binary masks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that GLENDA is the first region-annotated endometriosis laparoscopy dataset—stands or falls on the correctness of its 520 expert hand-drawn regions. Section 2 reports only that 'leading medical experts' created annotations; it gives no number of annotators, no inter-annotator agreement, no adjudication, and no histopathologic confirmation. The paper's own Conclusion weakens this assumption: the authors state that even specialists sometimes cannot classify endometriosis without further inspection and propose a future 'suspicion' class, acknowledging that the binary region labels may encode diagnostic uncertainty. Separately, Section 2 asserts that keyframe annotations remain valid throughout video segments because camera motion was 'kept at a minimal level,' but no motion metric, threshold, or per-sequence verification is provided. Since most of the 12K+ positive frames are unannotated frames in these segments, tracking-based augmentation inherits this unverified assumption. If the labels are noisy or keyframe masks drift, models trained on GLENDA inherit systematic label noise, undermining the dataset's claimed reuse for detection, localization, and tracking.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GLENDA, a dataset of gynecologic laparoscopy images with region-level annotations for endometriosis. It contains 25,682 frames from 138 video sequences, of which 520 hand-drawn region annotations are provided on 302 frames across four pathological classes (peritoneum, ovary, uterus, DIE) plus a no-pathology class. The authors describe the data collection process using the ECAT tool, the dataset structure, and intended uses for classification, detection, localization, and tracking. The main claim is that GLENDA is the first publicly released region-annotated endometriosis laparoscopy dataset.","tokens_in":7600,"tokens_out":2098,"duration_ms":25340,"significance":"If the annotations are accurate, GLENDA fills a genuine gap: no other public dataset offers region-level endometriosis annotations from laparoscopy. The dataset has a plausible scale (25K+ frames, 520 regions) and the authors are transparent about practical limitations such as class imbalance and temporal redundancy. The paper does not include any experiments or benchmarks, but for a dataset release that is not an intrinsic flaw. The decisive issue is whether the ground-truth labels can be trusted for training and evaluation; the manuscript currently provides no inter-annotator agreement, adjudication procedure, or histopathologic confirmation, and it relies on an unquantified keyframe-to-sequence propagation assumption. These are load-bearing because the contribution is precisely the correctness of the 520 expert regions and the usability of the video sequences for annotation augmentation.","major_comments":[{"comment":"The central claim that GLENDA is a usable ground-truth dataset is not supported by any quantitative annotation-reliability evidence. Section 2 states only that 'leading medical experts' created the annotations; it does not report the number of annotators, the number of cases annotated by more than one expert, inter-annotator agreement (e.g., Dice, IoU, Cohen's kappa), or an adjudication protocol. Section 4 and the Conclusion further acknowledge that even specialists sometimes cannot classify endometriosis without further inspection, and the authors propose a future 'suspicion' class. This self-admitted uncertainty is not reflected in the released binary region labels. Without any reliability measure, the reader cannot distinguish systematic label noise from accurate ground truth. This is the core weakness of the paper's central claim and should be addressed by adding annotation statistic","section":"Section 2 (Dataset Creation) and Section 4 (Limitations)"},{"comment":"The assertion that keyframe annotations remain valid throughout video segments because 'camera motion is kept at a minimal level' is not quantified. No motion metric, threshold, or per-sequence verification is provided. Since most of the 12K+ positive frames are unannotated frames that would be augmented by tracking keyframe masks, propagation error could substantially affect any downstream localization or tracking evaluation. The authors should either provide per-sequence camera-motion statistics (e.g., frame differencing, optical flow magnitude), or release only the annotated keyframes for evaluation and clearly mark propagated annotations as weak labels.","section":"Section 2 (Dataset Creation)"},{"comment":"The class distribution is extremely imbalanced, with peritoneum having 402 of 520 annotations and uterus only 14 annotations on 8 frames. The paper discusses this in Section 4, but it remains a significant limitation for any model trained on GLENDA as-is. The authors should provide explicit guidance on recommended evaluation protocols (e.g., class-wise AP, sequence-level splits, excluding or reweighting under-represented classes) and ideally include a benchmark or baseline experiment to demonstrate that the dataset is actually usable for the claimed detection/localization tasks. Without such a sanity check, the dataset's practical value is uncertain.","section":"Section 3, Table 1"}],"minor_comments":[{"comment":"Typo: 'a a complete dataset' should be 'a complete dataset'.","section":"Section 2"},{"comment":"Typo: 'addionally' should be 'additionally'.","section":"Section 3.2"},{"comment":"The acronym expansion 'Gynecologic Laparoscy ENdometriosis DAtaset' contains a typo ('Laparoscy') and does not match the spelled-out form used in the title.","section":"Section 5"},{"comment":"The 'T otal' row uses an unusual space; also 'max. cat.: 3' is not explained in the table caption or surrounding text. Please clarify whether the maximum of three categories per frame is a property of the annotation protocol or of the data.","section":"Table 1"},{"comment":"The figures show examples of each class, but the captions do not specify whether the green overlay corresponds to the keyframe annotation or the propagated mask. The caption in Figure 3 mentions 'keyframe annotations (green overlay)', while Figures 4–6 do not repeat this; please standardize.","section":"Section 3.1, Figure captions"},{"comment":"The recommendation to split training/validation/test by video sequence is sound, but the sentence 'yielding a perfect classification score' is colloquial and could be misread as a claim about the dataset's difficulty. Rephrase to make clear this is a warning about train/test leakage.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The dataset's existence and self-reported counts are internally consistent (annotated frames sum to 302, sequences to 138, and frames to 25,682; pathology frames sum to 12,244). The stress-test concern about annotation reliability is substantive and lands on the paper's central claim: for a dataset paper, unvalidated expert annotations are a major missing piece, especially when the authors themselves note that specialists disagree. The keyframe propagation assumption is also unquantified. Neither issue is fatal in principle—an IAA study, adjudication summary, and motion statistics could be added in a revision—so I lean major_revision rather than reject. I would also encourage the authors to provide a small benchmark or at least a baseline experiment to show that the dataset supports the claimed tasks, though I do not regard that as strictly mandatory for a dataset-release paper in this venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"GLENDA is likely the first public region-annotated endometriosis laparoscopy dataset, and that alone makes it worth a look. The paper describes the resource clearly: 25,682 frames, 520 hand-drawn regions on 302 frames across peritoneum, ovary, uterus, DIE, plus a no-pathology set, drawn from 138 sequences. The counts line up, the file structure is easy to follow, and the authors are upfront about the class imbalance and the danger of temporal leakage between video frames. They even flag that even specialists cannot always tell endometriosis apart without further inspection, which is why they plan a \"suspicion\" class. That honesty is a point in their favor.\n\nWhat is genuinely new is the gap: no prior laparoscopy dataset gives region-level endometriosis labels. Cholec80 and LapGyn4 exist for gallbladder and generic gynecologic content, but not for this condition. The dataset, if the annotations are reliable, would enable detection and localization work that otherwise has no public benchmark.\n\nNow the soft spots. The big one is that annotation quality is asserted, not demonstrated. There is no inter-annotator agreement, no adjudication, no histopathologic confirmation, and no baseline detection numbers. The paper's own conclusion undercuts the reliability assumption: if specialists themselves disagree or need biopsy to decide, then the binary region labels encode uncertainty that users cannot see. Most of the 12K+ positive frames come from keyframe annotations propagated across sequences; the claim that camera motion is \"minimal\" is not quantified. That propagation is a reasonable design choice, but without a motion metric or verification, the extra annotations are an attractive idea rather than a validated fact. This is addressable—a small IAA study and one detection baseline would largely answer the concern.\n\nThere are other minor issues: the uterus class is tiny (8 annotated images, 14 regions), so training on it directly is not practical; and the dataset URL could not be checked from the manuscript, which is a practical problem if the resource is meant to be public. None of these are fatal. The paper is honest, the resource fills a real gap, and the structural choices are sensible.\n\nWho is this for? People working on surgical video analysis, especially endometriosis detection, and anyone building benchmarks in niche medical domains. It deserves peer review—not a desk reject—but the referee should push for annotation reliability evidence and baseline results before publication.\n\nRecommendation: engage with it, but ask for the missing validation.","headline":"First public region-annotated endometriosis laparoscopy dataset; a solid resource paper, but annotation quality is asserted, not demonstrated.","tokens_in":8006,"tokens_out":2500,"would_cite":true,"duration_ms":27126,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper releases GLENDA, a 25,682-frame laparoscopy dataset annotated with 520 expert-drawn endometriosis lesion regions, described as the first public region-based endometriosis dataset.","keywords":["endometriosis","laparoscopy","surgical video dataset","region-based annotation","lesion detection","medical image analysis","machine learning","deep learning"],"falsifier":"Collect a sample of GLENDA's video segments and have independent endometriosis specialists draw regions on non-keyframe frames, then measure overlap with the keyframe annotation propagated by tracking; alternatively, biopsy-confirm a subset of annotated lesions. If propagated regions consistently miss the lesion (low IoU) or if confirmed false-positive annotations are frequent, the ground-truth and tracking assumptions are falsified.","tokens_in":7274,"feed_emoji":"🩺","tokens_out":5963,"duration_ms":57572,"temperature":0.7,"pith_summary":"GLENDA is introduced as the first publicly available image dataset with region-based annotations of endometriosis in gynecologic laparoscopy. It contains 25,682 video frames, roughly half showing endometriosis and half showing no visible pathology, together with 520 hand-drawn expert regions on 302 frames that mark four pathological locations: peritoneum, ovary, uterus, and deep infiltrating endometriosis. The dataset is designed so the same images support binary pathology classification, multi-label classification, lesion detection and localization, and, through 138 annotated video sequences with minimal camera motion, tracking-based annotation augmentation. The paper's claim matters because medical imaging datasets, especially in gynecologic laparoscopy, are scarce; releasing one with expert region labels gives computer vision research a supervised benchmark for a painful condition that currently requires time-consuming manual review of surgical recordings.","feed_headline":"First region-annotated endometriosis laparoscopy dataset debuts","feed_subtitle":"25,682 frames plus 520 expert-drawn regions give vision models four endometriosis classes to learn.","key_machinery":"The central object is the GLENDA dataset itself, structured as paired frame images and binary annotation images whose file names encode video ID, sequence range, frame ID, class, and annotation ID, allowing frames to be mapped to their regions by partial path matching. The annotation scheme carries the argument: every endometriosis region is a closed freehand drawing, polygon, or rectangle drawn in the ECAT tool, tied either to a single frame or to a keyframe of a video segment. The load-bearing design choice is that keyframes are selected only where camera motion is minimal, so a region drawn on one frame is claimed to remain visible for all frames in the segment, enabling tracking-based au","core_discovery":"The authors' central claim is that GLENDA is 'the first of its kind': the first public endometriosis dataset whose annotations are regions, not just image-level or video-level labels. To establish this, they introduce a dataset of 25,682 frames from over 400 gynecologic laparoscopy videos, including more than 12,000 positive frames with visible endometriosis and more than 13,000 negative frames without visible endometriosis. A total of 520 annotations were hand-drawn by leading endometriosis experts on 302 frames, distributed across four location-based categories—peritoneum (402 annotations), ovary (51), uterus (14), and deep infiltrating endometriosis (53)—using the ECAT annotation tool wit","pith_inferences":["Beyond the paper's stated plans, a validation study comparing tracked keyframe regions against fresh expert annotations on non-keyframe frames would test whether the minimal-camera-motion assumption actually holds; without such a check, tracking-based augmentation could silently propagate annotation errors.","The uterus class, with only 14 annotations on 8 frames, is likely too small for standalone deep-learning training, so the dataset's practical utility for rare classes depends on augmentation, class fusion, or future expansion.","Because no-pathology segments carry no region annotations, negative regions cannot serve as hard negatives in detection training; detectors trained on GLENDA can only learn positive regions and frame-level negatives.","The paper's warning about near-duplicate sequential frames implies that an independent benchmark built on GLENDA should split by video or sequence rather than by random frame, which would give a more realistic estimate of model performance."],"forward_implications":["Binary pathology-versus-no-pathology classification can be trained directly on the 25K+ frames without using region labels.","Multi-label classification and localization can be trained on the 520 expert regions across four endometriosis locations.","Because 138 sequences are annotated at keyframes with minimal camera motion, tracking algorithms can propagate regions to neighboring frames, multiplying the effective number of labeled samples.","The file-naming convention lets users map every frame to its binary annotation masks and convert them to bounding boxes or polygons as needed.","Planned extensions—additional categories, lesion severities, and an 'endometriosis suspicion' class—aim to improve classifier robustness on visually ambiguous cases."],"supporting_citations":[{"why":"Supplies the revised rASRM classification that defines the peritoneum and ovary endometriosis locations used as GLENDA categories.","marker":"[2]"},{"why":"Supplies the Enzian classification system used for the deep infiltrating endometriosis category.","marker":"[3]"},{"why":"LapGyn4 is the closest prior gynecologic laparoscopy dataset, providing the comparison that supports GLENDA's novelty.","marker":"[4]"},{"why":"ECAT is the annotation tool used to create all 520 hand-drawn regions, including the keyframe and sequence mechanism.","marker":"[6]"},{"why":"The TUM LapChole dataset is cited as another public laparoscopy dataset, illustrating that prior releases focus on other procedures.","marker":"[9]"},{"why":"The Cholec80 dataset is cited as a prominent public laparoscopy dataset, used as evidence that such datasets are few and mostly cholecystectomy-based.","marker":"[11]"},{"why":"The GI dataset is cited among the sparse public endoscopy datasets, reinforcing the scarcity that GLENDA addresses.","marker":"[12]"}],"fun_headline_variants":["First region-annotated endometriosis laparoscopy dataset","GLENDA: first region-level endometriosis dataset for AI","Expert-drawn regions debut in endometriosis laparoscopy data","Endometriosis region annotations: GLENDA sets first standard"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The value of GLENDA rests on the assumption that an expert's hand-drawn region on a keyframe is an accurate label for endometriosis presence and location, and that this label stays correct for every frame of the video segment because camera motion is minimal.","fun_headline_variants_meta":{"raw":{"variants":["First region-annotated endometriosis laparoscopy dataset","GLENDA: first region-level endometriosis dataset for AI","Expert-drawn regions debut in endometriosis laparoscopy data","Endometriosis region annotations: GLENDA sets first standard"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1243,"prompt_tokens":735,"completion_tokens":508,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":479,"completion_tokens_details":{"reasoning_tokens":440}},"tokens_in":479,"tokens_out":508,"duration_ms":5223,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:18:26.336476+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a sample of GLENDA's video segments and have independent endometriosis specialists draw regions on non-keyframe frames, then measure overlap with the keyframe annotation propagated by tracking; alternatively, biopsy-confirm a subset of annotated lesions. If propagated regions consistently miss the lesion (low IoU) or if confirmed false-positive annotations are frequent, the ground-truth and tracking assumptions are falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the revised rASRM classification that defines the peritoneum and ovary endometriosis locations used as GLENDA categories."},{"cited_title":"coloproctology 39(2), 121–133 (mar 2017)","cited_arxiv_id":null,"evidence_quote":"Supplies the Enzian classification system used for the deep infiltrating endometriosis category."},{"cited_title":"In: Proc","cited_arxiv_id":null,"evidence_quote":"LapGyn4 is the closest prior gynecologic laparoscopy dataset, providing the comparison that supports GLENDA's novelty."},{"cited_title":"In: Intl","cited_arxiv_id":null,"evidence_quote":"ECAT is the annotation tool used to create all 520 hand-drawn regions, including the keyframe and sequence mechanism."},{"cited_title":"IEEE Transactions on Medical Imaging 36(1), 86–97 (jan 2017)","cited_arxiv_id":null,"evidence_quote":"The Cholec80 dataset is cited as a prominent public laparoscopy dataset, used as evidence that such datasets are few and mostly cholecystectomy-based."},{"cited_title":"Medical image analysis 30, 144–157 (2016)","cited_arxiv_id":null,"evidence_quote":"The GI dataset is cited among the sparse public endoscopy datasets, reinforcing the scarcity that GLENDA addresses."}],"review_version":1}