Pith. sign in

REVIEW 3 cited by

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.01343 v2 pith:M3Y37W2J submitted 2024-08-02 cs.CV cs.AIcs.LG

StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation

classification cs.CV cs.AIcs.LG
keywords informationduringencodingfeaturefusionmulti-modalsegmentationstitchfusion
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Multimodal semantic segmentation shows significant potential for enhancing segmentation accuracy in complex scenes. However, current methods often incorporate specialized feature fusion modules tailored to specific modalities, thereby restricting input flexibility and increasing the number of training parameters. To address these challenges, we propose StitchFusion, a straightforward yet effective modal fusion framework that integrates large-scale pre-trained models directly as encoders and feature fusers. This approach facilitates comprehensive multi-modal and multi-scale feature fusion, accommodating any visual modal inputs. Specifically, Our framework achieves modal integration during encoding by sharing multi-modal visual information. To enhance information exchange across modalities, we introduce a multi-directional adapter module (MultiAdapter) to enable cross-modal information transfer during encoding. By leveraging MultiAdapter to propagate multi-scale information across pre-trained encoders during the encoding process, StitchFusion achieves multi-modal visual information integration during encoding. Extensive comparative experiments demonstrate that our model achieves state-of-the-art performance on four multi-modal segmentation datasets with minimal additional parameters. Furthermore, the experimental integration of MultiAdapter with existing Feature Fusion Modules (FFMs) highlights their complementary nature. Our code is available at StitchFusion_repo.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

    cs.CV 2026-04 unverdicted novelty 7.0

    OVRSISBenchV2 is a realistic benchmark expanding scene and category coverage for open-vocabulary remote sensing segmentation, with Pi-Seg baseline showing strong transfer via positive-incentive noise perturbations.

  2. Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline

    cs.CV 2026-04 unverdicted novelty 5.5

    OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.

  3. Multimodal SAM-adapter for Semantic Segmentation

    cs.CV 2025-09 conditional novelty 5.0

    A side-tuning adapter injects RGB-plus-auxiliary-sensor fused features into SAM's encoder, reaching state-of-the-art semantic segmentation on DeLiVER, FMB, and MUSES.