{"id":"4ad17515-8f25-49f5-bb5d-e053e7e6dd31","arxiv_id":"2606.15937","paper_version":2,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An adapted Mask2Former with 200 queries, a Feature Refinement Module, and auxiliary rare-class supervision reaches 70.08% mIoU on the GOOSE 2D FGSS benchmark.","lead":"GOOSE-M2F adapts Mask2Former for fine-grained semantic segmentation on unstructured outdoor terrain with 64 classes and severe long-tailed imbalance. The targeted changes plus public code deliver a reported 70.08% composite mIoU and third place on the GOOSE challenge.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Attribution of mIoU gains to the three listed modules lacks supporting ablations","rationale":"The reader's weakest assumption directly identifies the same attribution gap. Because the contribution is empirical and code is public, the proposed ablation is a single, decisive check. No other internal inconsistency or unsupported assumption appears load-bearing for the leaderboard claim.","tokens_in":1863,"tokens_out":314,"duration_ms":20740,"concrete_test":"Using the released code, train the unmodified Swin-Large Mask2Former baseline with the identical multi-stage schedule, loss, augmentations, and inference engine but without the 200-query change, FRM, and auxiliary head; report validation composite mIoU. If the gap to 70.08% is <3 points, the headline attribution does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim states that extending Mask2Former with 200 queries, FRM (ASPP-lite + CBAM), and an Auxiliary Supervision Head, plus the multi-stage training recipe, produces the reported 70.08% composite mIoU. No ablation results are referenced that isolate these additions from the Swin-Large backbone, Distribution-Balanced loss, Rare-Class Copy-Paste, dynamic re-weighting, EMA, or the sliding-window + TTA inference pipeline. In a long-tailed 64-class setting, it remains possible that the bulk of the gain arises from hyper-parameter tuning or the base architecture rather than the three targeted contributions.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper presents GOOSE-M2F, a task-specific adaptation of Mask2Former (Swin-Large backbone) for the GOOSE 2D FGSS benchmark involving 64 fine-grained classes in long-tailed unstructured outdoor terrain. It introduces three targeted extensions—200 object queries, a Feature Refinement Module (FRM) combining ASPP-lite and CBAM, and an Auxiliary Supervision Head—paired with a multi-stage training recipe (Distribution-Balanced loss, Rare-Class Copy-Paste, dynamic re-weighting, EMA) and inference pipeline (sliding-window with Gaussian blending and 4-scale TTA). The model reports 70.08% Official Composite mIoU (63.55% fine, 76.61% coarse), placing 3rd on the leaderboard, with public code and models released.","tokens_in":1957,"tokens_out":457,"duration_ms":34366,"significance":"If the reported gains hold and are attributable to the proposed modules, the work supplies a competitive, reproducible baseline for long-tailed fine-grained semantic segmentation in challenging outdoor settings. The public GitHub and Hugging Face releases constitute a clear strength for verification and extension by the community.","major_comments":[{"comment":"Abstract and §3 (contributions and method): The central claim attributes the 70.08% composite mIoU primarily to the three listed additions (200 queries, FRM, Auxiliary Supervision Head) plus the described training/inference pipeline. No ablation tables or quantitative isolation of these components from the Swin-Large backbone, Distribution-Balanced loss, Rare-Class Copy-Paste, or TTA are referenced, leaving open the possibility that gains derive mainly from hyper-parameter tuning or the base architecture rather than the targeted modules.","section":"Abstract, §3"}],"minor_comments":[{"comment":"Abstract: The Hugging Face link contains the placeholder 'XYZ9843'; replace with the actual model identifier for reproducibility.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads primarily as a competition entry; confirm whether the journal scope favors such reports or requires stronger methodological novelty beyond leaderboard placement."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback and for recognizing the significance and reproducibility of the work through the public code and model releases. We address the single major comment below and will incorporate the requested changes in the revised manuscript.","responses":[{"response":"We agree that the current version of the manuscript does not provide ablation tables that isolate the contributions of the three proposed modules (200 queries, FRM, Auxiliary Supervision Head) from the Swin-Large Mask2Former baseline, the Distribution-Balanced loss, Rare-Class Copy-Paste augmentation, or the TTA/inference pipeline. This leaves the attribution of the reported gains open to the interpretation raised by the referee. In the revised manuscript we will add a dedicated ablation study section that reports incremental performance when each component is added in turn, together with controls that hold the training recipe and inference pipeline fixed. These tables will be referenced from both the abstract and §3.","revision_made":"yes","referee_comment":"[Abstract, §3] Abstract and §3 (contributions and method): The central claim attributes the 70.08% composite mIoU primarily to the three listed additions (200 queries, FRM, Auxiliary Supervision Head) plus the described training/inference pipeline. No ablation tables or quantitative isolation of these components from the Swin-Large backbone, Distribution-Balanced loss, Rare-Class Copy-Paste, or TTA are referenced, leaving open the possibility that gains derive mainly from hyper-parameter tuning or the base architecture rather than the targeted modules."}],"tokens_in":1467,"tokens_out":335,"duration_ms":27452,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"GOOSE-M2F reaches third on the GOOSE 2D FGSS leaderboard with 70.08% composite mIoU by extending Mask2Former. The authors increase object queries to 200, add a Feature Refinement Module with ASPP-lite and CBAM, insert an auxiliary supervision head, and apply a multi-stage training recipe that includes distribution-balanced loss, rare-class copy-paste, dynamic re-weighting, and EMA. At inference they use sliding-window processing with Gaussian blending and 4-scale TTA.\n\nThe paper does one thing clearly: it ships a working system for a 64-class long-tailed outdoor segmentation task and releases both code and trained models. That makes the result immediately usable for anyone who needs a strong baseline on the GOOSE benchmark or similar unstructured terrain data.\n\nThe soft spot is the missing evidence on what actually drove the score. The abstract and the stress-test note both indicate no ablation tables that isolate the three listed additions from the Swin-Large backbone, the loss functions, the augmentation choices, or the inference pipeline. In a severely long-tailed setting, it is possible the bulk of the lift comes from tuning or the base architecture rather than the targeted modules. Without those breakdowns the central claim about the three contributions stays under-supported.\n\nThis paper is for people already working on the GOOSE challenge or on fine-grained segmentation for robotics. A reader who wants the leaderboard number and the public implementation will find it useful. It does not reorganize the field or introduce new theory.\n\nI would bring it to a reading group as maybe. I would not cite it in my own work in the next year unless I needed the exact baseline. It deserves peer review because the empirical result is verifiable with the released artifacts, even if the analysis of contributions needs tightening.","headline":"GOOSE-M2F reaches third on the GOOSE leaderboard via Mask2Former adaptations, but attribution of gains to the new modules is not supported by ablations.","tokens_in":2496,"tokens_out":445,"would_cite":false,"duration_ms":33440,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"GOOSE-M2F adapts Mask2Former with 200 queries, a refinement module, and auxiliary supervision to reach 70.08% mIoU on long-tailed 64-class outdoor segmentation.","keywords":["Mask2Former","semantic segmentation","long-tailed distribution","fine-grained segmentation","outdoor terrain","object queries","feature refinement"],"falsifier":"An ablation that trains the plain Mask2Former baseline with the identical multi-stage recipe and inference engine and then checks whether composite mIoU stays near 70% or falls well below it on the GOOSE test set.","tokens_in":2741,"feed_emoji":"🗺️","tokens_out":761,"duration_ms":42533,"temperature":0.7,"pith_summary":"The paper demonstrates an adaptation of Mask2Former for the GOOSE 2D Fine-Grained Semantic Segmentation challenge involving 64 classes in unstructured outdoor terrain with severe long-tailed imbalance where rare classes occupy fewer than 50 pixels per image. It extends the baseline by raising object queries to 200, inserting a Feature Refinement Module that pairs ASPP-lite with CBAM dual-attention, and adding an Auxiliary Supervision Head for direct per-pixel gradients on rare classes. These elements combine with a multi-stage training regimen using distribution-balanced loss, rare-class copy-paste, dynamic re-weighting, and EMA, followed by sliding-window inference with Gaussian blending and 4-scale test-time augmentation. A sympathetic reader would care because the approach targets a practical robotics setting where accurate recognition of infrequent terrain elements matters for navigation safety.","feed_headline":"Mask2Former adaptation scores 70% mIoU on long-tailed terrain segmentation","feed_subtitle":"Extra queries, refinement module, and auxiliary head handle 64 classes where rare ones span under 50 pixels","key_machinery":"The Feature Refinement Module (FRM) that merges ASPP-lite with CBAM dual-attention, paired with 200 object queries and an auxiliary per-pixel supervision head, to improve representation and gradients for rare classes under long-tailed distributions.","core_discovery":"Extending Swin-Large Mask2Former with 200 object queries to avoid saturation, a Feature Refinement Module combining ASPP-lite and CBAM, and an Auxiliary Supervision Head, together with multi-stage training and dense inference, produces 70.08% Official Composite mIoU (63.55% fine, 76.61% coarse) and third place on the GOOSE 2D FGSS leaderboard.","pith_inferences":["The same query increase and auxiliary head could be tested on other long-tailed segmentation datasets outside outdoor robotics.","The sliding-window blending technique may transfer to high-resolution tasks such as aerial or medical image analysis.","The FRM pattern could be swapped into other transformer-based segmentors to check for similar gains on imbalanced data."],"forward_implications":["Raising object queries to 200 removes representational saturation when modeling 64 classes.","The FRM supplies refined features that help distinguish fine-grained terrain details.","The auxiliary head supplies direct gradients that improve learning of classes with very few pixels.","Distribution-balanced loss combined with rare-class copy-paste reduces the impact of long-tailed imbalance.","Sliding-window inference with Gaussian blending and multi-scale TTA contributes an additional 10.57% to the final score."],"fun_headline_variants":["GOOSE-M2F scores 70.08% mIoU on long-tailed fine-grained terrain","Mask2Former variant scores 70.08% mIoU third on GOOSE leaderboard","Mask2Former with queries and FRM scores 70.08% mIoU on GOOSE","GOOSE-M2F adaptation scores 70.08% mIoU on 64-class outdoor terrain"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The listed architectural changes and training steps, rather than hyperparameter search or the Swin-Large backbone alone, are the main sources of the mIoU improvement.","fun_headline_variants_meta":{"raw":{"variants":["GOOSE-M2F scores 70.08% mIoU on long-tailed fine-grained terrain","Mask2Former variant scores 70.08% mIoU third on GOOSE leaderboard","Mask2Former with queries and FRM scores 70.08% mIoU on GOOSE","GOOSE-M2F adaptation scores 70.08% mIoU on 64-class outdoor terrain"]},"model":"grok-4.3","cost_usd":0.007897,"raw_usage":{"total_tokens":3641,"prompt_tokens":749,"num_sources_used":0,"completion_tokens":100,"cost_in_usd_ticks":78974500,"prompt_tokens_details":{"text_tokens":749,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2792,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":749,"tokens_out":100,"duration_ms":27686,"temperature":1.0,"reasoning_tokens":2792,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T03:25:42.113889+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An ablation that trains the plain Mask2Former baseline with the identical multi-stage recipe and inference engine and then checks whether composite mIoU stays near 70% or falls well below it on the GOOSE test set.","supporting_citations":[],"review_version":1}